Test-Time Compute in Mid-2026: Buy the Knob, Not the Mode
Eight months of effort knobs, harness records, parallel agents, capacity-bound GPUs, and one unbeaten benchmark: a buyer's review of which layer of the test-time compute stack deserves the money.
The heavy-mode premium is temporary: Gemini 3.1 Pro shipped 77.1% on ARC-AGI-2 at its predecessor's unchanged price one week after the vendor's expensive thinking mode set its record [, ]; treat vendor modes as previews of the next base model, and put durable engineering into effort controls, stopping policy, and harnesses instead.
Spend now lives in the knob, the harness, and the router, not the rate card: an open-source refinement harness bought a leaderboard-topping ARC-AGI-2 score for about $31 per task on a commodity base model [], and per-token deflation is being outrun by reasoning-token volume [].
Very little of it transfers to interaction: on ARC-AGI-3's turn-based environments every frontier system at launch scored under 1% while humans solve essentially every environment [].
Should you pay for the vendor's heavy-thinking mode?
Only as a rental. The pattern of the last eight months is that what a heavy mode demonstrates, the next base-model revision tends to absorb, so the decision is about timing, not capability. Named precisely, the evidence is one verified cycle plus a suggestive parallel: Gemini 3 Deep Think's ARC-AGI-2 record was absorbed in bulk by the base Gemini 3.1 Pro one week later [, , ], and the base Opus 4.6 scored 68.8% on the same benchmark [] against 37.6% for Opus 4.5's best verified thinking configuration []. It is a pattern worth betting on, not a law; a counterexample follows, and the closing section carries its falsification markers.
The period opened with the effort knob, not a bigger mode. Anthropic's Claude Opus 4.5 (November 24, 2025) made an API-level effort parameter the headline: at medium effort it matches Sonnet 4.5's best SWE-bench Verified score while using 76% fewer output tokens, and at high effort it exceeds that score by 4.3 points while still using 48% fewer []. The same launch cut Opus pricing 67%, from $15/$75 to $5/$25 per million tokens, with the model scoring 80.9% on SWE-bench Verified []. Both sources are secondary launch briefs, evidence tier C, but they agree with each other and with Anthropic's framing: token efficiency at fixed quality became the product. Practitioner measurement filled in the shape of the curve: one single-author experiment (configuration-level evidence, tier B, one task family) priced low, medium, and high effort at $0.021, $0.059, and $0.122 per call and found medium effort delivered "89% of the quality at 52% of the cost" of high []. That diminishing-returns shape, not any single score, is the argument for defaulting to the middle of the knob.
OpenAI made the opposite bet two weeks later and charged for it. GPT-5.2 (December 11, 2025) was the first model above 90% on ARC-AGI-1, with a 400,000-token context and 70.9% on GDPval, outputs a third-party infrastructure analysis describes as arriving at more than 11x the speed and under 1% of the cost of human experts across 44 occupations []. Its price rose 1.4x over GPT-5.1, to $1.75/$14 per million tokens, framed explicitly as the compute intensity of extended reasoning []. A price increase justified by thinking depth was new; it did not become a trend.
Then February 2026 compressed the whole thesis into fourteen days. Anthropic's Opus 4.6 (February 5) replaced the manual thinking-budget parameter with adaptive thinking (the model calibrates its own depth under a four-tier effort control) and scored 68.8% on ARC-AGI-2 at Max effort with a 120K thinking-token ceiling []. Google's upgraded Gemini 3 Deep Think (February 12) hit 84.6% on ARC-AGI-2, verified by the ARC Prize Foundation, with a 3455 Codeforces Elo [], gated behind the $250/month AI Ultra subscription []. And one week later, Gemini 3.1 Pro (February 19) posted 77.1% on ARC-AGI-2 (against 31.1% for Gemini 3 Pro, its immediate predecessor) as a base model, at pricing unchanged from that predecessor [, ]. For calibration: secondary coverage of the benchmark puts the human baseline near 60%, at roughly $17 per task in labor-cost terms [], though other secondary coverage claims 84% []; neither source is the benchmark operator. The bulk of what the $250/month mode demonstrated in week one was selling at commodity prices by week two. The same coverage, though, calls the base model "a different architectural approach" rather than the mode's machinery folded down []: the verified absorption is of capability, not mechanism.
The live counterexample is OpenAI's own price ladder: the depreciation thesis predicts frontier capability at unchanged base prices, yet OpenAI's flagship rate rose from $1.75/$14 per million tokens at GPT-5.2 [] to $5/$30 for GPT-5.6 Sol [, ]. That trajectory fits the thesis only if each generation's capability keeps reappearing cheaper one rung down, exactly what the closing section's absorption marker tests. A premium that holds across two base-model cycles would mean the rental advice is wrong for that line.
The decision rule follows directly. Rent heavy modes for problems that cannot wait a quarter; do not architect systems around them. The durable interface is the knob: effort controls now exist across the Anthropic and OpenAI lines [, ], and DeepSeek V4 ships the same idea as three reasoning depths: Non-think, Think High, and Think Max []. The knobs, not the mode SKUs, are where allocation policy lives.
Where is scaling actually happening — deeper chains, harnesses, or clocks?
Above the model, mostly. The clearest verified evidence of the period is that, on the December 2025 leaderboard, a refinement harness on a cheap base model out-scored everything below the frontier's most expensive mode. The newest vendor products now productize parallelism and wall-clock budgets rather than longer chains.
The ARC Prize Foundation's own 2025 analysis (a primary source, since it operates the benchmark) names refinement loops the central theme of the year, in a competition that drew 1,455 teams and 15,154 submissions []. The concrete exhibit is Poetiq's open-source harness: running on Gemini 3 Pro, it lifted ARC-AGI-2 performance from the model's 31% baseline at $0.81 per task to 54% at $31 per task, topping the verified leaderboard at the time of publication; the best verified commercial model entry, Opus 4.5 (Thinking, 64k), scored 37.6% at $2.20 per task []. Because the harness comparison holds the underlying model constant, this is the closest thing the period offers to tier-A evidence for where scaling gains live. Poetiq announced preliminary results on November 20, 2025, and swapped in Gemini 3 and GPT-5.1 within hours of their releases; the harness is model-agnostic by construction []. The same report shows why the layer matters: on one ARC-AGI task, Gemini 3 Pro needed 96 reasoning tokens where Deep Think spent 138,000 []. A harness that decides when refinement is worth another iteration operates on exactly the allocation margin that vendor modes handle crudely. The vendors' July releases point the same direction. GPT-5.6 (previewed July 9, 2026 to selected partners; tiered as Luna, Terra, and Sol at $1/$6, $2.50/$15, and $5/$30 per million tokens) ships two explicit test-time products: a max effort tier for deeper single-chain reasoning, and an ultra mode that OpenAI describes as "coordinating four agents in parallel by default, trading higher token use for stronger results and faster time-to-result on demanding tasks" [, , ]. OpenAI's launch post (a vendor-authored source, so its comparative numbers deserve the standard discount) reports the period's most interesting scaling curve in hours rather than tokens: on ExploitGym, GPT-5.6 reaches a 24.9% pass rate under a two-hour cap against GPT-5.5's 15.1%, and 33.7% when given six hours []. The same post claims a new state of the art of 80 on the Artificial Analysis Coding Agent Index at the time of publication, 2.8 points above Fable 5, using less than half the output tokens and costing about one-third less []. Note what the flagship pitch is: not a higher score at any price, but the same frontier at half the tokens. Grok 4.5 (July 8, 2026, at $2/$6 per million tokens) likewise launched on the vendor-stated claim of roughly 2x the token efficiency of comparable models, solving tasks in under half the steps [].
So the stack now has four scaling layers: the model, the effort knob, the harness, and the clock. The harness result is the one that should change roadmaps: it is verified by the benchmark operator, it is open source, and it beat every commercial entry below the top vendor mode at a price teams can actually pay. Wall-clock scaling is the one to watch skeptically: the ExploitGym curve is a single vendor-reported result on a security benchmark, and nothing in the saved evidence yet shows it replicated elsewhere.
When should the model stop thinking?
Earlier than it wants to. The field spent this period turning that observation into control machinery on three distinct layers: training-free stopping rules, self-estimated budgets, and trained termination.
The saturation evidence starts at the product layer. A review of Opus 4.6 notes that on reasoning-heavy ARC-style tasks the model saturates its available thinking tokens at every effort level, producing similar scores regardless of the setting []; the knob stops mattering precisely where marketing implies it matters most. The waste it is masking can be extreme: one preprint documents QwQ-32B generating 100 times more tokens than a comparable non-reasoning model to answer what 2+3 is [].
The training-free layer stops generation mid-chain. REFRAIN (ACL 2026, peer-reviewed; tier A for its controlled comparisons) cuts token usage 20–55% while maintaining or improving accuracy by detecting reflective-but-redundant reasoning, and frames "when-to-stop as a new and practical axis of test-time scaling" — reasoning "not just more, but just enough" []. A dynamic early-exit method at ICLR 2026 self-truncates chains of thought at confidence transitions, with no additional training, shortening reasoning by 19.1% to 80.1% across 11 reasoning models on 10 benchmarks while improving accuracy by 0.3% to 5.0% []. A companion ACL 2026 result, SAT, makes stopping step-wise rather than chain-wise: it treats reasoning as a finite-state machine that compresses easy steps and preserves depth for hard ones, cutting tokens 25.1% on average while adding 1.5 accuracy points across 9 models and 7 benchmarks [].
The budget layer moves the decision before generation. SelfBudgeter (a controlled comparison, but a single-team preprint revised April 2026 and unreplicated, so tier B) trains a model to estimate its own minimal token budget from the query, compressing responses 61% on average on math tasks; its 1.5B variant scores 84.10% on GSM8K at an average of 1,231.79 tokens per response where its baseline scored 73.09% at 2,865.08 []. And the trained-termination layer bakes stopping into the policy itself: JET, a single-team preprint revised March 2026, teaches models to terminate proactively; its 1.5B distilled model gains 4.6% accuracy while producing 46.3% shorter outputs on an Olympiad benchmark []. The consistent mechanism across all of these: overlong reasoning is not just waste, it actively steers models toward wrong conclusions [, ].
Latency belongs in the same ledger. Thinking tokens are billed in seconds as well as dollars: reasoning requests occupy hardware memory longer, reducing total system concurrency, and reasoning modes on low-complexity tasks produce timeout cascades with no measurable accuracy gain []. A stopping policy that saves tokens is also buying back user-facing seconds.
Adaptive thinking is this literature productized. When Anthropic deprecated manual thinking budgets in favor of model-calibrated depth [], the vendor position became: the model decides when to stop, you set the ceiling. That is the right default. The operational rule for teams: run adaptive or medium-effort settings as baseline, escalate effort only on verified failure, and treat any workload where accuracy does not move across effort tiers as a signal you are paying for tokens the model itself would discard.
Can the serving stack keep up with long reasoning?
Barely, and the constraint is memory, not arithmetic. The engineering literature of the period converges on a diagnosis: long reasoning changed the shape of the inference workload, and the hardware roadmap is being redrawn around it.
A measurement study on H200 clusters (tier B, systems benchmarking rather than controlled model comparison) describes the shift precisely: chat workloads emit output sequences around 500 tokens, while reasoning traces run past 10,000, pushing inference out of the compute-bound prefill regime into what the authors call a capacity-bound regime, where linearly growing KV caches exhaust GPU memory and leave "stranded capacity": requests throttled while compute sits idle []. A survey of inference I/O quantifies why the problem is structural: GPU compute grew roughly 18x from V100 to B200 while memory bandwidth grew about 9x, and autoregressive decode sits at an arithmetic intensity near 1, squarely bandwidth-bound, loading on the order of 140 GB from memory per token for a 70B model in FP16 []. Practitioner numbers make it concrete: a single Llama 3.1 70B request at 128K context consumes about 42.9 GB of GPU memory for KV cache alone, and prefix caching on a 128K-token prompt cuts time-to-first-token from roughly 11 seconds to 1.5 on an H100 []. Every reasoning token a stopping policy saves is also memory-residency time returned to the serving fleet: the overthinking literature and the capacity wall are the same problem seen from two floors of the stack.
The hardware response is real but comes wrapped in vendor arithmetic. NVIDIA's own benchmark post (vendor-authored, tier C for its comparative claims) reports 60,000 tokens per second per GPU on a 120B open-weight model on B200, a 15x reduction in cost per million tokens versus the prior generation, and 10x throughput per megawatt on mixture-of-experts models, alongside a headline claim that a $5 million rack system generates $75 million in token revenue []. Third-party rental economics are more modest but directionally consistent: at March 2026 on-demand prices ($4.54, $6.03, and $9.08 per hour for H200, B200, and GB200 respectively), a B200 serving Llama 3.3 70B in FP4 works out to roughly $0.17 per million tokens against $0.50 on H200, contingent on FP4 quantization []. One market analysis argues the spending mix has already flipped, with inference at 85% of enterprise AI budgets and fleet utilization sitting at 15–30% []. Those are directional tier-C figures, but they are consistent with every vendor's July pitch being token efficiency rather than token count.
The deepest response is architectural, and it came from the open-weight side: DeepSeek V4's hybrid sparse attention runs 1M-token context at 27% of the FLOPs and 10% of the KV cache of the prior generation []. Read as a serving story rather than a benchmark story, that is an attack on exactly the capacity wall the measurement literature describes. For buyers the implication is indirect but budgetary: serving economics now improve on two independent clocks, hardware generations and attention architectures, and any cost model that extrapolates today's per-token price more than a quarter forward is guessing.
What do open weights change about reasoning economics?
They set the floor. In this period the floor moved close enough to the frontier to change routing math, without resolving the question of who should self-host.
The releases came quarterly. Alibaba's Qwen3.5-397B-A17B (February 15, 2026, open weight) claims decoding throughput 8.6x and 19.0x that of its own closed Qwen3-Max at 32K and 256K context respectively, vendor-authored figures against the vendor's own prior model []. DeepSeek V4 (April 24, 2026, MIT license) spans V4-Pro at 1.6T total/49B active parameters and V4-Flash at 284B/13B, with Flash priced at $0.14/$0.28 per million tokens []. And Moonshot's Kimi K3 (a 2.8-trillion-parameter mixture-of-experts model with about 50B active parameters) launched hosted in July with weights announced for July 27, 2026, priced at $3/$15 per million tokens against $5/$30 for GPT-5.6 Sol and roughly $10/$50 for Claude Fable 5, and scoring 57.1 on the Artificial Analysis Intelligence Index, fourth among 189 models at the time of publication [, , ]. Demand ran ahead of supply: Moonshot paused new subscriptions on July 20 after launch traffic exceeded expectations [].
Two cautions keep this from being a simple arbitrage. First, cheap tokens plus an unmanaged appetite for them is not a cheap model: in third-party measurement, V4-Flash emitted 230M output tokens on an evaluation suite where the median model emitted 88M; as the analysis puts it, "The number that decides whether Flash is actually cheap is output tokens per task, not price per token" []. Second, the self-hosting math is genuinely unsettled. One enterprise analysis puts the break-even against frontier APIs at 100 to 256 million tokens per month and estimates full self-hosting costs at 3 to 5x the raw GPU rental price once labor and utilization waste are included []; another puts the crossover at 5 to 10 million tokens per month []; a hardware-focused total-cost analysis lands at roughly 2 to 3 million tokens per day for consumer hardware, warning that "Comparing local LLMs against cloud APIs on per-token price alone is a trap" []. A spread that wide is itself the finding: break-even depends more on a team's utilization and operations cost than on any property of the models.
The adoption data carries the same tension. By one enterprise-spend estimate, open-weight models held just 11% of enterprise market share in 2026, down from 19% in 2024, against $8.4 billion in 2025 API spend concentrated in three closed vendors []. That is a tier-C directional statistic measuring dollars rather than tokens, and it suggests premium reasoning dollars stay closed. For most teams the practical posture is the hybrid one the decision-framework literature converges on: frontier APIs for the hardest reasoning, open-weight endpoints for volume, and self-hosting only where sovereignty rules force it [].
What does a completed task cost, and who decides which model thinks?
Whatever the meter says, not what the price list says. Increasingly a router decides, because the price spread between capable models has grown too wide to ignore.
A practitioner analysis of what it calls the LLM cost paradox puts the mechanism plainly: per-token prices keep falling while token consumption for reasoning models grows around 5x annually, so cost per task climbs []. Its sharpest exhibit: one test suite completed for roughly $9.30 on one model and $95 on another, a 10x gap for identical results, driven entirely by token bloat []. This is undated web analysis, the weakest evidence tier cited here, but the direction is corroborated by a pricing-tool vendor's measurements: reasoning models routinely emit 10–50x more output tokens than standard models, which is why that source pushes accuracy-per-dollar as the only honest comparison metric [].
Routing is the operational answer, and it now has both peer-reviewed and production evidence. RouteLLM (ICLR 2025, peer-reviewed) held quality at 95% of GPT-4's while cutting costs 85% on MT-Bench, routing only 14% of queries to the strong model []. An agent-cost benchmark reports that routing simple steps to cheap models and escalating only complex reasoning cuts per-task costs 75–85% against running a single high-capability model []. One engineering team's coding assistant (a tier-C single case) cut daily spend from $3,000 to $970, about 68%, by sending 70% of tasks to cheaper models, an annualized difference over $740,000 []. At production scale the pattern holds: Agoda's centralized gateway with routing and cost attribution underpins more than 200 production applications []. The economic driver is simple: by mid-June 2026 the spread between the cheapest usable model and the most expensive was roughly 100x [].
Two costs of routing itself belong in the ledger. The routing literature's own warning is "silent quality regression": cheaper models producing subtle errors that dashboards miss, which makes quality gates a precondition rather than an afterthought []. And routing infrastructure is not free: one analysis of multi-signal routing estimates a three-year total cost of ownership between $1.19M and $1.69M at a million queries per month, mostly labor []. Combined with the effort knob, the 2026 budget playbook is four lines: route by difficulty, cap effort by default, gate quality on the cheap path, and meter cost-per-completed-task, because a model's real price is now a behavioral property, measured, not quoted.
Does any of it transfer — and what would change these judgments?
Almost none of it, yet. This is the finding that disciplines everything above. The ARC Prize Foundation released ARC-AGI-3 in March 2026: turn-based interactive environments in which an agent must explore, infer goals, and plan with no instructions []. Its technical report (posted April 2026) describes 25 public, 55 semi-private, and 55 private environments, scored by action efficiency relative to human play []. The launch results, from the benchmark operator: Gemini 3.1 Pro, the best frontier model, scored 0.37%; GPT-5.4 scored 0.26%; Opus 4.6 scored 0.25%; humans solve 100% of environments []. The best preview-phase score, 12.58%, came from a purpose-built reinforcement-learning-plus-graph-search agent, not a frontier LLM [].
Read this against the rest of the period. Effort knobs, refinement harnesses, parallel subagents, and six-hour budgets all moved static, verifiable benchmarks — ARC-AGI-2, released March 2025 [], went from a 31.1% frontier baseline [] to 84.6% [] in under a year. The same machinery, pointed at environments that must be explored rather than solved, has so far produced single digits at best. Test-time compute, in every form this review covers, amplifies competence on checkable, stateless problems; the saved evidence contains no demonstration that it buys exploration efficiency at any token price. That also carries a warning about the benchmarks it does move: ARC-AGI-2's headroom was consumed in months, and any static benchmark's remaining headroom should be treated as a consumable, not a moat.
Because this column's judgments should be falsifiable, here is what to watch, with markers. First, absorption: the depreciation thesis predicts the capabilities of GPT-5.6's ultra mode and Gemini's Deep Think appear in base-model successors at unchanged prices within roughly a quarter; a premium mode holding its edge past two base-model cycles, or OpenAI's price ladder holding its premium over the same span, would break the pattern. Second, interaction: the ARC Prize 2026 competition on Kaggle opened March 25, 2026 and closes November 2, 2026, with $850,000 at stake: $150,000 in progress prizes plus a $700,000 bonus unlocked only at 100%, and 9,902 entrants with 16,083 submissions at the page's snapshot []. The markers: whether any entrant beats the 12.58% preview score on hidden environments [], and whether any frontier LLM (not a purpose-built agent) reaches double digits. Third, the clock: whether the hours-versus-accuracy curve OpenAI reported on ExploitGym [] is replicated by a non-vendor party on a non-security benchmark. Fourth, the floor: whether Kimi K3's weights ship on July 27 as announced [, ] and what third-party hosts charge to serve a 2.8T-parameter mixture-of-experts model. Fifth, the meter: cost paradox claims [] predict cost per completed task keeps rising through 2026; two consecutive quarters of falling per-task costs on stable workloads would falsify them. The three-layer verdict stands until then: the knob is a solved product problem, so use it; stopping is a maturing research problem, so adopt its defaults; exploration is not yet a problem anyone has shown they can spend their way into.
Every key figure in this report is individually traced to a source extract.
These checks establish citation traceability and internal consistency. They do not independently reproduce the underlying experiments, guarantee that third-party figures are correct, or ensure that volatile values — prices, model versions, benchmark results — have not changed since retrieval. Evidence reflects sources as of publication (2026-07-21); citations last re-verified 2026-07-25.