Tag

#evaluation

Reports tagged evaluation.

Reports

5 reports
Aug 2026
08-02 Frontier Research The Audit Join Is Still a Hypothesis A replayable table can still encode the wrong experiment. Across five domains, the useful audit object is the local decision joined to evidence, execution, versions, and status. 19 cited · of 265 sources
Jul 2026
07-29 Frontier Research The Open-Weight Gap Is Four Points Wide and 1.56 Terabytes Deep: What It Costs to Run the Best Open Models Yourself Kimi K3 pulled open weights to within four points of the closed frontier, then shipped as a 1.56-terabyte download with a 64-accelerator floor. The gap did not close. It moved. 36 cited · of 402 sources 07-29 Deep Dive The Description Is Always Loaded: How Agent Skills Break Tasks Without Being Invoked Skill libraries lift an agent's average and quietly break tasks it already solved, and one channel of that damage runs through a line of metadata that sits in context whether or not the skill is opened. 12 cited · of 362 sources 07-22 Flagship Agent Harnesses Multiply What the Model Already Has — They Don't Supply It In the one six-model study that measured it, swapping only the scaffolding cut the bill by 41 percent — while the quality gain moved with the model underneath. 14 cited · of 42 sources 07-21 Flagship Contamination Is the Least of It — Why Harder and More Private Benchmarks Won't Fix AI Measurement OpenAI audited the tasks its own model kept failing and found that most of the tests were wrong. Contamination cannot explain that, and neither can the remedies the field has been buying. 30 cited · of 36 sources