Tag
#benchmarks
Reports tagged benchmarks.
Reports
2 reportsJul 2026
07-27 Frontier Research AI Coding Agents in Mid-2026: The Score Went Up, the Checking Did Not Get Cheaper One independently run leaderboard cleared ninety percent this July. Every headline number measures what an agent can produce; almost none measures what it costs to check. 07-21 Flagship Contamination Is the Least of It — Why Harder and More Private Benchmarks Won't Fix AI Measurement OpenAI audited the tasks its own model kept failing and found that most of the tests were wrong. Contamination cannot explain that, and neither can the remedies the field has been buying.