LUMIERE

How Close Are Open Weights to the Frontier, Really?

Nine months of trillion-parameter open releases have held the capability gap near a measured four months — but "open" is now as much a legal, operational, and geopolitical fact as a benchmark one, and each of those layers tells a different story.

Evidence density · by section 117 sources → 41 cited
TL;DR
CONCLUSION

Open-weight models sit roughly 4 months behind the closed frontier on aggregate capability, and the newest releases hold genuine top-3 to top-5 positions on independent indices — but that single number hides four separate gaps (private benchmarks, licensing, deployment cost, safety posture) that each widen or narrow the story on their own terms [, ].

COSTS

For an operator the case is arithmetic — open weights buy a six-to-seven-times lower output-token price and a permissive-license exit from lock-in — but self-hosting the biggest models means holding roughly 744 GB of weights in memory before the first token, so "open" in practice usually means a cheap hosted API with self-hosting as an option rather than a default [, , ].

LIMITS

On contamination-resistant evaluations the parity picture inverts, and the open frontier's other weaknesses are structural rather than temporary — absent system cards, safeguards that fine-tuning can strip, and a supply base concentrated in a handful of Chinese labs whose continued openness is a business decision, not a law [, , ].

How much of the frontier is open now?

As of late July 2026, the strongest open-weight models occupy positions no open release has held before — top-3 or top-5 on the major independent indices — yet the measured distance to the closed frontier has not shrunk on aggregate. Epoch AI's May 29, 2026 analysis, covering January 1 through May 28, found the best open-weight models lagging the closed state of the art by an average of 4 months, or 8 points on its Capabilities Index (90% CI: 7–11) — slightly wider than the 3-month average it measured for January 2023 through October 2025 []. Both frontiers are moving fast; the offset between them is roughly constant. Nathan Lambert's read after the July releases is more generous — the debated 6–9 month gap now looks more like 3–5 months [] — but no serious tracker puts it at zero.

What changed in the last nine months is who supplies the open frontier and at what scale. Four releases define the period, all from Chinese labs. Moonshot's Kimi K2.6 arrived April 20 as a 1-trillion-parameter mixture-of-experts model with 32B active parameters, under a Modified MIT license at $0.60 input / $2.50 output per million tokens []; its vendor-reported 58.6% on SWE-Bench Pro edged GPT-5.4 (57.7%), Claude Opus 4.6 (53.4%), and Gemini 3.1 Pro (54.2%) on a contamination-resistant coding suite []. DeepSeek V4 followed on April 24: a 1.6T-total / 49B-active MoE plus a 284B "Flash" variant, both MIT-licensed with 1M-token context, priced at $0.435 / $0.87 per million tokens after DeepSeek made its launch discount permanent on May 22 [, ]. Z.ai's GLM-5.2, released in mid-June 2026 under MIT with 744B total and roughly 40B active parameters, became the first open model to score 51 on the Artificial Analysis Intelligence Index v4.1 — the highest open-weights result at the time of publication — while beating GPT-5.5 on its publisher-reported SWE-bench Pro score (62.1 vs 58.6) at $1.40 / $4.40 per million tokens against GPT-5.5's $5.00 / $30.00 [, , , ]. And on July 16 Moonshot announced Kimi K3, a 2.8-trillion-parameter MoE with weights promised for July 27; in Arena's blind testing, developers preferred K3 over Claude Fable 5 and GPT-5.6 Sol for front-end coding, and it outranked Opus 4.8 in the broader text ranking []. K3 sits third on the Artificial Analysis Intelligence Index, behind only the closed flagships [].

The Western side of the ledger is thinner and moving the other way. Meta never shipped Llama 4 Behemoth (~2T total / 288B active); it served as a distillation teacher, and Meta's frontier effort re-emerged in April 2026 as the closed-weight Muse Spark [] — a real capability recovery, jumping from Llama 4 Maverick's Intelligence Index score of 18 to Muse Spark's 52, but a closed one []. Alibaba runs a two-track strategy: its strongest models, Qwen3.7-Max (May 2026, scoring 56.6 on the Artificial Analysis index at the time []) and the multi-trillion-parameter Qwen3.8-Max previewed July 19, are closed at launch, with open weights promised "soon" and the flagship's benchmark claims so far unverified by any published model card [, ]. The genuinely open Western entries are smaller or newer: Mistral Large 3 (675B total / 41B active, Apache 2.0) with a larger family entering early access before the end of July [], Thinking Machines' Inkling (975B-total / 41B-active Apache 2.0, trained on 45T tokens) [], and — as of July 22 — Poolside's Laguna S 2.1, a 118B / 8B-active coding model reported to reach 70.2% on Terminal-Bench 2.1 []. Open-weight AI in 2026 is, by token volume and by capability ceiling, predominantly a Chinese export with a European and startup fringe. That is a strategic fact, not a footnote — and the rest of this piece is largely about its consequences.

Where is parity real — and where is it benchmark-only?

The benchmark on which open models most visibly "caught up" stopped measuring capability. On February 23, 2026, OpenAI published an audit of SWE-bench Verified and stopped reporting it: of 138 audited problems its models failed, 59.4% had material flaws in tests or task descriptions, and frontier progress on the suite had slowed to a crawl (74.9% to 80.9% over six months) while models reproduced memorized patches — evidence of training-set contamination across every major lab, closed ones included []. The scale of the inflation shows up when the same models face held-out code: Claude Opus 4.5 scored 80.9% on Verified but roughly 23% on SWE-bench Pro in March 2026 []. SWE-bench Pro exists precisely to resist this — 1,865 tasks across 41 repositories, including 18 commercial codebases legally unavailable to training crawlers — and on it, models drop further on the private split than the public one: GPT-5's early resolve rate fell from 23.3% public to 14.9% held-out, Claude Opus 4.1 from 22.7% to 17.8% [, ]. Any parity claim built on SWE-bench Verified should be treated as unmeasured.

Held to that stricter standard, open models split into genuine wins and exposed gaps — sometimes within the same model. The genuine wins: Kimi K2.6's 58.6% and GLM-5.2's 62.1% on SWE-Bench Pro — both publisher-reported — beat or matched GPT-5.5's reported 58.6 on a suite designed to punish memorization [, ]. GLM-5.2's 1524 Elo on GDPval-AA — Artificial Analysis's long-horizon, multi-turn knowledge-work benchmark — puts it third overall, above GPT-5.5's 1509 []. The GDPval-AA result is an independent measurement and the SWE-Bench Pro figures are self-reported on a contamination-resistant suite; together they support a narrow but real claim: for agentic coding and structured knowledge work, the best open models now sit inside the closed pack rather than behind it.

The exposed gaps are just as concrete. DeepSeek V4-Pro's headline 80.6% on SWE-bench Verified collapses to 8% pass@1 on DeepSWE, a written-from-scratch benchmark, against 70% for GPT-5.5 and 54% for Opus 4.7 — a spread that the aggregator reporting it attributes partly to harness effects, but which is far too large to explain away []. The widest measured gap is long-horizon autonomy: the best open computer-use agent, OpenCUA-72B, reaches a 45.0% success rate on OSWorld-Verified at 100 steps against Claude Sonnet 4.5's 61.4% [], and METR's March 2026 analysis found roughly half of test-passing agent pull requests would not actually be merged by maintainers []. Epoch flags the structural reason to expect this pattern: open models hill-climb public benchmarks aggressively and tend to perform worse on private ones, which means the 4-month aggregate gap is more likely understated than overstated [].

A second, quieter caution comes from how the rankings are built. One tracker that scores Chinese models on multiple contracts reports Kimi K3 as the overall leader (81) but MiniMax M3 (69.8) as the strongest model with downloadable weights, and notes that the confidence intervals for the top downloadable models overlap — MiMo-V2.5-Pro at 63.24–77.19 against MiniMax M3 at 65.84–73.75 — so the best open model title is not statistically decisive on that surface []. My reading: parity is real on cost-adjusted agentic coding and mid-tier reasoning; it is benchmark-only wherever a public test set has had a year to leak into pretraining; and the honest error bars on "how far behind" run from about 3 months to considerably more, depending on how private your evaluation is and which leaderboard contract you trust.

What does "open" actually license you to do?

The word "open" is doing more work than it can bear, and the licensing layer is where that shows most. Start with the distinction the industry mostly elides: the Open Source Initiative's Open Source AI Definition (OSAID 1.0, published October 28, 2024) requires four things — source code, model weights, training-data information sufficient to recreate the system, and a license granting use, study, modification, and sharing without field-of-use restrictions []. Almost none of the models in this review clear that bar, because none release their training data; DeepSeek, Qwen, Kimi, and GLM ship weights and a permissive license but not the dataset, which by OSAID's own logic makes them "open weights," not "open source" [, ]. The definition itself is contested — the Software Freedom Conservancy argues OSAID erodes the term by accepting "data information" instead of the data, while OSI counters that a full-data requirement would cede the field to vendors who release nothing []. An analysis presented at All Things Open 2025 found that nearly 60% of models labeled "open" on Hugging Face carry no license at all, and that Apache 2.0 (around 23%) and MIT are the most common OSI-approved terms among those that do [, ]. OSI has signaled a 1.1 or 2.0 revision through Q4 2026 []; until then, "open" on a model card is a marketing word, not a legal one.

Read as legal instruments, the top-end licenses have quietly become the most permissive they have ever been. DeepSeek V4 and GLM-5.2 ship under plain MIT — no usage caps, no acceptable-use policy, no disclosure obligations, and in DeepSeek's case explicitly no monthly-active-user clause [, , ]. Mistral's family is Apache 2.0, which its CEO markets as the compliance-friendly path for European enterprises that need on-premise deployment under EU jurisdiction []. The caveats are specific rather than sweeping. Kimi K2.6 uses a "Modified MIT" license whose attribution clause, on one license-tracker's reading of the Kimi Modified-MIT terms, activates only at very large deployment scale — permissive in substance but a legal-review item rather than a rubber stamp [, ]. The genuinely restrictive terms now attach to the models that no longer lead: the Llama Community License caps commercial use at 700M monthly active users [], and the standard Qwen license carries its own monthly-active-user ceiling [, ]. This is a reversal from 2023–2025, when the bespoke Llama-style license was the default template for "open."

The legal layer also determines who a regulator can reach. Under the EU AI Act, general-purpose-model obligations (Articles 53–55) have applied since August 2, 2025, with high-risk-system requirements landing August 2026 []. The Act carves out a partial exemption for models under a free and open license that also publish weights, architecture, and usage information — but that exemption is conditional and, on the reading of several compliance analysts, does not cover copyright-policy or training-data-summary duties, and evaporates entirely for any model classified as systemic-risk, defined by a training-compute threshold above 10^25 FLOPs []. Two consequences follow that operators tend to miss. First, a restrictive license can forfeit the exemption: an external advisory group to the EU AI Office concluded in January 2026 that the Llama Community License, because of its user-count cap, does not qualify as free and open for these purposes — a B-tier reading from compliance explainers rather than a court ruling, and it suggests rather than settles the question []. Second, deployers who substantially modify a model, or place it on the market under their own name, can themselves become "providers" with their own obligations — and a lower 3×10^24 FLOPs threshold governs when a downstream fine-tune is treated as a new model [, ]. For a European enterprise, the practical upshot is that plain MIT or Apache weights are not just cheaper; they are the configuration most likely to keep the lighter regulatory treatment.

What does it actually cost to run one?

Mostly money is what "open" buys, and the market has already voted. On OpenRouter, the share of routed tokens handled by US closed-model providers fell from roughly 70% in June 2025 to roughly 30% by July 2026; Chinese open-weight models peaked at 46% of US enterprise token volume in April, DeepSeek alone is the platform's largest provider at 16–17% of traffic, and Airbnb and Uber have publicly confirmed running Chinese open-weight models in production []. The driver is arithmetic: open-weight models cost six to seven times less per output token than frontier closed models, while on production coding workloads independent measurements put the capability gap as low as 2–3 percentage points []. Agentic workflows sharpen this, consuming up to 1,000 times more tokens than chat-shaped interactions — at that multiplier, a sixfold price difference decides architectures [].

The "just self-host" instinct runs into physics before it runs into economics. GLM-5.2 is a 744B-parameter MoE that activates only about 40B parameters per token, but every expert has to sit in memory at once — so at FP8 the weights alone need roughly 744 GB of VRAM, and about 1,488 GB at full BF16 precision, before any KV cache for its 1M-token context: one byte and two bytes per parameter [, ]. Aggressive quantization changes the shape of the problem rather than removing it: an Unsloth 1-bit build fits in about 223 GB, and a 512 GB M3 Ultra Mac Studio can run the model at all — but at 6.5 to 9.5 tokens per second at zero context, degrading to 5–5.5 as context fills, which is a demo speed, not a production one []. DeepSeek's designers clearly anticipated this: the 284B "Flash" variant activates only 13B parameters and is pitched at local deployment, and V4-Pro reports a 90% reduction in KV-cache consumption in long-context settings — architecture choices aimed squarely at making self-hosting tractable, though those two figures come from a lower-confidence secondary source and should be read with wider error bars [].

Put the physics and the economics together and the crossover math is unforgiving. H200 instances run $4.50–6.00 per hour on the hyperscalers ($3.50–4.50 on Lambda or CoreWeave), and break-even against frontier closed APIs lands around 2M–5M tokens per day — but against the optimized open-model API providers, who already run thin-margin infrastructure, break-even shifts to 50M+ tokens per day []. Realistic total-cost models add $1,800 to $18,000 per month of operations labor depending on tier [], and the biggest open models are not small deployments: DeepSeek V4-Pro needs a minimum of 8× H100 or 4× B200 for production inference []. For most teams, "open" in practice means a cheap hosted API for open weights, with self-hosting as a credible exit option rather than a day-one plan. What makes the pricing durable is that the exit is real — but it is an option, not a default.

The exception is the operator for whom control, not price, is the point. The one architectural advantage that no API can match is that downloaded weights have no network endpoint: for an air-gapped, data-residency-bound, or connectivity-limited deployment — a hospital diagnostic system, a bank's internal tooling, an industrial controller — a model that runs entirely inside the perimeter is qualitatively different from one gated behind a vendor's API, and the Chinese open releases are increasingly designed with that edge case in mind []. That argument is where the licensing and deployment layers meet: full weight custody plus a permissive license is a procurement story that a closed API cannot tell at any price. (One caveat worth flagging: that framing comes from a C-tier source and is stated here as an operator's rationale, not a measured outcome.)

How do the safety and security postures differ?

The clearest difference is documentation, and it is not close. Frontier closed releases ship with system cards and pre-deployment safety evaluations; Kimi K2.6 shipped with neither, inheriting K2.5's documented pattern of fewer refusals on CBRNE-adjacent prompts and politically skewed Chinese-language outputs, with no new system card published []. That gap matters more than it sounds, because the safety properties of an open-weight release are not fixed at publication. Two 2025–2026 studies make the point with controlled experiments rather than assertion. In one, researchers took a near-final checkpoint of an open-weight model and applied malicious fine-tuning — reinforcement learning with web-browsing tools — and found it substantially raised the model's measured biological-risk capability, though it left cybersecurity capability below the "High" threshold, with every model tested scoring 0% in cyber-range environments absent hints []. In a second, adversarial fine-tuning of an open genomic model on publicly available viral sequences circumvented the data-exclusion safeguard the model shipped with, restoring the ability to predict immune-escape mutations for a virus that had been held out of both the original and the fine-tuning data — which the authors read as direct evidence that data filtering is not an irreversible mitigation []. These are A-tier findings by design: each isolates the safeguard-removal variable rather than inferring it.

The security posture inverts the usual open-source intuition in one respect and confirms it in another. On the "many eyes" side, published weights and architecture do let external researchers run exactly the red-team evaluations above — which is why the marginal-risk evidence base for open models is, ironically, better than for closed ones. On the misuse side, the same openness means a released capability cannot be recalled: a closed lab can throttle or withdraw a model that turns out to be dangerous, while an open one has, in the phrase the policy literature keeps returning to, let the horse out of the barn. Both the misuse research and the standards commentary land on the same practical asymmetry — an open-weight release is a one-way door, so its safety case has to hold not for the model as shipped but for the model a determined fine-tuner can produce from it [, ]. On present evidence that case holds for the frontier open models today on the highest-consequence axes (no tested model crossed the "High" cybersecurity bar) — albeit on prior-generation open releases rather than the newest flagships — but it holds by margin, not by design, and the margin is a moving quantity.

Who controls the open frontier, and can policy reach it?

The concentration is the headline: the open frontier is, by token volume and capability ceiling, a Chinese export. That is a capability fact with a governance tail, and the governance tail is where evidence quality drops sharply, so the tiers matter. The strongest anchor is a policy document rather than a benchmark. The US NTIA's 2024 report on dual-use models with widely available weights concluded, in its own words, that "the government should not restrict the wide availability of model weights for dual-use foundation models at this time," recommending instead a monitoring program built on marginal-risk analysis — the risks unique to open release relative to closed — and specifically flagging the lag between a capability appearing in a proprietary model and reaching an open one as the metric to watch []. That lag time is exactly the 4-month figure Epoch now measures [, ]. The NTIA framing is A-tier for what it is — an official recommendation — but it is from 2024, and the policy weather has since shifted; treat it as the baseline that later moves are departing from, not as current US posture.

Where the story gets louder, the sourcing gets thinner, and I am flagging that rather than smoothing it over. Reporting that the US moved in April 2026 toward export-style controls on model weights and cloud access — "know your customer" rules for compute providers, aimed at closing the loophole of training on US servers — comes from a C-tier outlet and should be read as a directional signal, not a confirmed rule []. The recurring analytical claim in this space, that open-weight distribution structurally defeats export controls because weights are downloadable and fine-tunable without a procurement trail, appears across multiple commentaries but rests on argument rather than documented enforcement failure — explicitly C-tier; it suggests a real problem without proving one []. What can be said with more confidence sits at the intersection of the layers already established: because a downloaded model has no endpoint to regulate, the sovereignty argument has flipped polarity in an awkward way — the configuration that gives an operator full weight custody and license freedom is overwhelmingly Chinese-origin, which several sources note is itself becoming a procurement and policy question [, ]. Epoch adds the structural asymmetry from the other direction: closed labs sometimes withhold their most capable models for safety or commercial reasons, so the visible closed frontier is a floor, not a ceiling []. The honest summary is that the capability facts here are well-grounded and the geopolitical inferences drawn from them are mostly B/C-tier opinion; I have kept the two clearly separated.

What should an operator watch next?

The open frontier's current strength depends on a supply decision that a handful of labs re-make every quarter, and the leading indicators are about incentives more than capabilities. The bullish signal is that open releases keep multiplying: one close observer of the Chinese labs frames open weights as a strategic hedge against power concentration []. The bearish signal is the enclosure pattern among the scale players — Meta went closed with Muse Spark [, ], Alibaba keeps its flagships closed at launch [, ], and even Moonshot now opens weights on a delay rather than day one []. Both trends are real and they point opposite ways; which dominates is the single most consequential open question for anyone building on these weights.

Three concrete things to track. First, the private-benchmark gap, not the public one: watch SWE-bench Pro's held-out split, GDPval-AA, and any evaluation whose test set is legally uncrawlable, because that is where the 4-month figure is measured honestly and where it will move first [, , ]. Second, the Western fringe, which is where license-clean alternatives to Chinese weights would have to come from — Mistral's larger family, Thinking Machines' Inkling, and coding-specialist entrants like Poolside's July 22 release are the models to watch for whether open and not Chinese-origin becomes a viable procurement category rather than a smaller-and-older one [, , ]. Third, the regulatory treatment of downstream fine-tunes: if the EU's provider-status and compute-threshold rules bite in practice, the compliance calculus that currently favors plain-MIT weights could shift again []. Four months is short in calendar time; at the current release cadence it is a full model generation. Operators should use the four-month frontier where its private-benchmark case is real, and treat its continued existence as a policy and business outcome rather than a settled fact.

How we verified

Every key figure in this report is individually traced to a source extract.

Figures traced to source
100% 75 of 75 figures
Automated checks
Citation markers reconciled against the reference list Every cited URL verified against the evidence store Fact-to-citation attribution overlap checked Every named system grounded in a retrieved source Section citations confined to their pre-bound evidence set All quantitative figures traced to source extracts
117 sources retrieved 104 passed relevance screening 98 in the writer's working set 41 directly cited

These checks establish citation traceability and internal consistency. They do not independently reproduce the underlying experiments, guarantee that third-party figures are correct, or ensure that volatile values — prices, model versions, benchmark results — have not changed since retrieval. Evidence reflects sources as of publication (2026-07-22); citations last re-verified 2026-07-25.