Archive

Every report we have published

17 reports, 4,274 sources read, 469 receipts attached. Filter by column, or search a title.

All 17 reports · page 1 of 2
Aug 2026
08-04 Deep Dive The Cache Miss After the Gap Coding agents serve 95.7% of their input tokens from cache. The steps that miss cluster in one place — the first call after an idle gap — and that is where the serving work went. 31 cited · of 298 sources 08-02 Frontier Research The Audit Join Is Still a Hypothesis A replayable table can still encode the wrong experiment. Across five domains, the useful audit object is the local decision joined to evidence, execution, versions, and status. 19 cited · of 265 sources 08-01 Deep Dive Distillation Mostly Reweights — Until the Teacher Adds Something New A weaker teacher can beat a stronger one when its token distribution overlaps the student's. In one 2026 study that overlap tracked successful distillation — but did not guarantee it. 24 cited · of 513 sources
Jul 2026
07-31 Deep Dive Distillation Copies the Transcript, Not the Teacher A student model learns from a transcript, not from a mind — which is why capability transfers cheaply, why refusal quietly does not, and why the policy fight is aimed at the wrong part of the machine. 31 cited · of 377 sources 07-29 Flagship The Blast Radius Was the Bug: Agent Incidents Are an Authority Problem, Not an Intelligence One Most of the enterprise agent damage on public record arrived with no attacker attached. What made it unrecoverable was not the model's mistake but the standing authority already sitting behind it. 29 cited · of 147 sources 07-29 Deep Dive Containment Moved Downstream: How Deployed Agents Fight Prompt Injection by Giving Up Reach In-model defenses give much of their gain back once attackers study them. The cuts that hold sit on what the agent may do — they cost reach, and no test has yet matched the scale that broke the rest. 28 cited · of 256 sources 07-29 Deep Dive Two Workloads, One Wire: What Splitting Prefill and Decode Actually Buys Splitting prefill from decode does not remove the contention between them. It relocates that contention onto a wire, and every layer built since exists to pay that bill down. 31 cited · of 296 sources 07-29 Deep Dive Every Token Is a Read: How Memory Bandwidth Prices AI Inference, and Who Collects the Rent An H100 is sold on its arithmetic, but emitting a token is mostly a memory read — and the firms that sell those bytes now book a richer margin than the company whose name is on the accelerator. 30 cited · of 323 sources 07-29 Deep Dive The KV Cache Became Inventory — and Almost Nobody Charges Rent Yet Serving stacks now give the KV cache its own flash tier, its own router and its own rack. The bill a customer sees still discounts recomputation avoided, not memory held. 41 cited · of 211 sources 07-29 Frontier Research The Open-Weight Gap Is Four Points Wide and 1.56 Terabytes Deep: What It Costs to Run the Best Open Models Yourself Kimi K3 pulled open weights to within four points of the closed frontier, then shipped as a 1.56-terabyte download with a 64-accelerator floor. The gap did not close. It moved. 36 cited · of 402 sources