Column
Frontier Research
What does the evidence actually show? Briefings that weigh the freshest research on one fast-moving question.
Reports
4 reports · 1,114 sources read · 106 receiptsAug 2026
08-02 The Audit Join Is Still a Hypothesis A replayable table can still encode the wrong experiment. Across five domains, the useful audit object is the local decision joined to evidence, execution, versions, and status. Jul 2026
07-29 The Open-Weight Gap Is Four Points Wide and 1.56 Terabytes Deep: What It Costs to Run the Best Open Models Yourself Kimi K3 pulled open weights to within four points of the closed frontier, then shipped as a 1.56-terabyte download with a 64-accelerator floor. The gap did not close. It moved. 07-29 Test-Time Compute Buys Candidates, Not Judgment: More Samples Don't Tell You Which Answer Is Right More inference compute reliably buys more chances at a right answer. It does not buy the ability to tell which one — and that gap sits under the flattening returns, the reversals at long chain lengths, and the gains that do not survive re-evaluation. 07-27 AI Coding Agents in Mid-2026: The Score Went Up, the Checking Did Not Get Cheaper One independently run leaderboard cleared ninety percent this July. Every headline number measures what an agent can produce; almost none measures what it costs to check.