Tag
#evaluation
Reports tagged evaluation.
Reports
5 reportsAug 2026
08-02 Frontier Research The Audit Join Is Still a Hypothesis A replayable table can still encode the wrong experiment. Across five domains, the useful audit object is the local decision joined to evidence, execution, versions, and status. Jul 2026
07-29 Frontier Research The Open-Weight Gap Is Four Points Wide and 1.56 Terabytes Deep: What It Costs to Run the Best Open Models Yourself Kimi K3 pulled open weights to within four points of the closed frontier, then shipped as a 1.56-terabyte download with a 64-accelerator floor. The gap did not close. It moved. 07-29 Deep Dive The Description Is Always Loaded: How Agent Skills Break Tasks Without Being Invoked Skill libraries lift an agent's average and quietly break tasks it already solved, and one channel of that damage runs through a line of metadata that sits in context whether or not the skill is opened. 07-22 Flagship Agent Harnesses Multiply What the Model Already Has — They Don't Supply It In the one six-model study that measured it, swapping only the scaffolding cut the bill by 41 percent — while the quality gain moved with the model underneath. 07-21 Flagship Contamination Is the Least of It — Why Harder and More Private Benchmarks Won't Fix AI Measurement OpenAI audited the tasks its own model kept failing and found that most of the tests were wrong. Contamination cannot explain that, and neither can the remedies the field has been buying.