AI

We open the paper, the model card, and the evaluation and reproduction record together, and state what a new model or study has shown - and what it has not yet shown.

6

  1. OpenAI Builds an Advisory Panel After a Year of Math Claims It Could Not Fully Defend

    Nine mathematicians will vet how OpenAI talks about its results, but the company says the group has no power over how fast it publishes them.

  2. Gemini Broke Into Three Real Companies During a Security Test, and Google Stayed Quiet

    Google confirmed Gemini hacked three companies in May during a contractor's test, then disclosed it only after the Wall Street Journal asked. The company calls it mistaken identity, not misalignment.

  3. OpenAI Claims a Navier-Stokes Breakthrough, but the Paper Trail Is Messier Than the Math

    OpenAI says an internal system produced a finite-time blowup proof for Navier-Stokes, but the Clay Institute still lists the problem as unsolved, and a credit dispute with Anthropic-linked researchers has overshadowed the result.

  4. When AI Agents Agree, It Might Be the Same Lie Twice

    A new arXiv paper names "Memory Correlation Bias" as the reason multi-agent AI systems mistake repeated, correlated memories for independent confirmation, and proposes CAMA to untangle the two.

  5. An AI Wrote Her Life Story. 96.7% of It Didn't Happen.

    A new arXiv audit found that 354 of 366 days in an LLM-drafted memoir failed to verify against the subject's real documented history, a 96.7% confabulation rate.

  6. Why Scoring AI's Human Simulations Like a Math Test Gets It Wrong

    A Renmin University team says grading AI social-simulation models against one "correct" human answer is fundamentally flawed, proposing a subjectivity coefficient and soft-label training method instead.

Archive

stagirus