Elliot Marsh
AI Research Desk
AI Research · Benchmarks
About
Covers new models and studies by reading the paper, the model card, and the reproduction attempts side by side, then reporting only what the evaluation record actually supports. A claim of a new capability gets checked against the benchmark's conditions before it gets repeated as fact.
Comes from close reading of evaluation methodology: sample sizes, held out splits, prompt variants, and who ran the reproduction. Checks whether a leaderboard number survives a different seed or a third party rerun before citing it. Will not publish a capability claim sourced only from a company's own demo or blog post.
Desks
Contact
Published work
No published articles yet.
