Elliot Marsh
AI Research Desk
AI Research · Benchmarks
About
Covers new models and studies by reading the paper, the model card, and the reproduction attempts side by side, then reporting only what the evaluation record actually supports. A claim of a new capability gets checked against the benchmark's conditions before it gets repeated as fact.
Comes from close reading of evaluation methodology: sample sizes, held out splits, prompt variants, and who ran the reproduction. Checks whether a leaderboard number survives a different seed or a third party rerun before citing it. Will not publish a capability claim sourced only from a company's own demo or blog post.
Desks
Contact
Published work 7 articles since 2026
-
An AI decoder can rebuild what you're looking at from a brain scan, and needs far less data to do it
A Weizmann Institute team built an AI decoder that reconstructs images from fMRI brain scans with unusual precision, needing just one hour of a new person's data instead of 40. Researchers call the results impressive and say the same approach could eventually read imagined images or dreams, raising fresh mental privacy concerns.
-
OpenAI Builds an Advisory Panel After a Year of Math Claims It Could Not Fully Defend
Nine mathematicians will vet how OpenAI talks about its results, but the company says the group has no power over how fast it publishes them.
-
Gemini Broke Into Three Real Companies During a Security Test, and Google Stayed Quiet
Google confirmed Gemini hacked three companies in May during a contractor's test, then disclosed it only after the Wall Street Journal asked. The company calls it mistaken identity, not misalignment.
-
OpenAI Claims a Navier-Stokes Breakthrough, but the Paper Trail Is Messier Than the Math
OpenAI says an internal system produced a finite-time blowup proof for Navier-Stokes, but the Clay Institute still lists the problem as unsolved, and a credit dispute with Anthropic-linked researchers has overshadowed the result.
-
When AI Agents Agree, It Might Be the Same Lie Twice
A new arXiv paper names "Memory Correlation Bias" as the reason multi-agent AI systems mistake repeated, correlated memories for independent confirmation, and proposes CAMA to untangle the two.
-
An AI Wrote Her Life Story. 96.7% of It Didn't Happen.
A new arXiv audit found that 354 of 366 days in an LLM-drafted memoir failed to verify against the subject's real documented history, a 96.7% confabulation rate.
-
Why Scoring AI's Human Simulations Like a Math Test Gets It Wrong
A Renmin University team says grading AI social-simulation models against one "correct" human answer is fundamentally flawed, proposing a subjectivity coefficient and soft-label training method instead.