We show our work. Every claim, back to its source.

10 reports · 962 sources read · 367 receipts published

All reports

10 reports
Jul 2026
07-25 Flagship Agents Don't Fail Smart An anatomy of the year's agent disasters — the database dropped during a code freeze, the runaway cloud bill, the leaked repository — and the case that the fix is least authority, not a smarter model. 32 cited · of 85 read 07-24 Deep Dive Prompt injection in production: how agent containment actually works When you give a language model tools and let it read the open web, the real question stops being whether it can be fooled and starts being what it can reach once it is. 27 cited · of 106 read 07-24 Frontier Research The KV-Cache Economy Long context increasingly behaves like a managed memory product: KV-cache mechanics shape price, latency, and capacity — how they set the rate, what vendor cache discounts actually buy, and where hit rate, tiering, and prefix discipline decide your bill. 45 cited · of 116 read 07-24 Flagship The Buildout Bet The money is real, the compute is scarce, and the revenue is finally arriving — so why does the safest wager in technology suddenly feel like the riskiest one on the board? 27 cited · of 92 read 07-22 Deep Dive The Great Unbundling of the GPU: How Prefill/Decode Disaggregation Became the Production Default In two years, splitting LLM inference into two specialized fleets went from an OSDI paper to the assumed architecture at NVIDIA, AWS, Google, Moonshot, and DeepSeek — and the interesting fights now are about the wire between the fleets, the scheduler above them, and whether research will quietly glue the halves back together. 50 cited · of 109 read 07-22 Flagship The New Unit of AI Capability Is the Configuration The harness — the loop of tools, context discipline, verification, and budgets around a model — has become a capability layer in its own right. But what performs, and what fails, is never the harness alone: it is a model, a harness, and an environment, operating together under a budget. 21 cited · of 81 read 07-22 Frontier Research How Close Are Open Weights to the Frontier, Really? Nine months of trillion-parameter open releases have held the capability gap near a measured four months — but "open" is now as much a legal, operational, and geopolitical fact as a benchmark one, and each of those layers tells a different story. 41 cited · of 117 read 07-21 Flagship The Benchmark Collapse Why public LLM scores stopped meaning what we think they mean 36 cited · of 82 read 07-21 Deep Dive The Memory Bill Comes Due: HBM and the Economics of Inference Why the price of a token is set by moving bytes, not math — and who collects along the way 40 cited · of 81 read 07-21 Frontier Research Test-Time Compute in Mid-2026: Buy the Knob, Not the Mode Eight months of effort knobs, harness records, parallel agents, capacity-bound GPUs, and one unbeaten benchmark: a buyer's review of which layer of the test-time compute stack deserves the money. 48 cited · of 93 read