Archive

Every report we have published

17 reports, 4,274 sources read, 469 receipts attached. Filter by column, or search a title.

All 17 reports · page 2 of 2
Jul 2026
07-29 Deep Dive The Description Is Always Loaded: How Agent Skills Break Tasks Without Being Invoked Skill libraries lift an agent's average and quietly break tasks it already solved, and one channel of that damage runs through a line of metadata that sits in context whether or not the skill is opened. 12 cited · of 362 sources 07-29 Frontier Research Test-Time Compute Buys Candidates, Not Judgment: More Samples Don't Tell You Which Answer Is Right More inference compute reliably buys more chances at a right answer. It does not buy the ability to tell which one — and that gap sits under the flattening returns, the reversals at long chain lengths, and the gains that do not survive re-evaluation. 18 cited · of 252 sources 07-29 Flagship The AI Buildout Isn't a Demand Bet — It's a Duration Bet: The Risk Is Asset Lifetimes and Refinancing, Not Missing Demand The measured evidence has already settled the demand question. What it has not settled is whether the revenue arrives before the silicon depreciates and the debt raised against it comes due. 40 cited · of 214 sources 07-27 Frontier Research AI Coding Agents in Mid-2026: The Score Went Up, the Checking Did Not Get Cheaper One independently run leaderboard cleared ninety percent this July. Every headline number measures what an agent can produce; almost none measures what it costs to check. 33 cited · of 195 sources 07-27 Deep Dive Ternary LLMs Remove the Multiplier, Not the Cost Ternary weights really do remove the multiplier and shrink a model eightfold. What the compression ratio hides is where that cost comes back: in tokens, in kernels, in silicon. 22 cited · of 85 sources 07-22 Flagship Agent Harnesses Multiply What the Model Already Has — They Don't Supply It In the one six-model study that measured it, swapping only the scaffolding cut the bill by 41 percent — while the quality gain moved with the model underneath. 14 cited · of 42 sources 07-21 Flagship Contamination Is the Least of It — Why Harder and More Private Benchmarks Won't Fix AI Measurement OpenAI audited the tasks its own model kept failing and found that most of the tests were wrong. Contamination cannot explain that, and neither can the remedies the field has been buying. 30 cited · of 36 sources