Tag

#inference-economics

Reports tagged inference-economics.

Reports

6 reports
Jul 2026
07-29 Deep Dive Two Workloads, One Wire: What Splitting Prefill and Decode Actually Buys Splitting prefill from decode does not remove the contention between them. It relocates that contention onto a wire, and every layer built since exists to pay that bill down. 31 cited · of 296 sources 07-29 Deep Dive Every Token Is a Read: How Memory Bandwidth Prices AI Inference, and Who Collects the Rent An H100 is sold on its arithmetic, but emitting a token is mostly a memory read — and the firms that sell those bytes now book a richer margin than the company whose name is on the accelerator. 30 cited · of 323 sources 07-29 Deep Dive The KV Cache Became Inventory — and Almost Nobody Charges Rent Yet Serving stacks now give the KV cache its own flash tier, its own router and its own rack. The bill a customer sees still discounts recomputation avoided, not memory held. 41 cited · of 211 sources 07-29 Frontier Research Test-Time Compute Buys Candidates, Not Judgment: More Samples Don't Tell You Which Answer Is Right More inference compute reliably buys more chances at a right answer. It does not buy the ability to tell which one — and that gap sits under the flattening returns, the reversals at long chain lengths, and the gains that do not survive re-evaluation. 18 cited · of 252 sources 07-29 Flagship The AI Buildout Isn't a Demand Bet — It's a Duration Bet: The Risk Is Asset Lifetimes and Refinancing, Not Missing Demand The measured evidence has already settled the demand question. What it has not settled is whether the revenue arrives before the silicon depreciates and the debt raised against it comes due. 40 cited · of 214 sources 07-27 Deep Dive Ternary LLMs Remove the Multiplier, Not the Cost Ternary weights really do remove the multiplier and shrink a model eightfold. What the compression ratio hides is where that cost comes back: in tokens, in kernels, in silicon. 22 cited · of 85 sources