SDSignal Desk

Exclusive: Paying for frontier AI models buys 4-month head start at 5x the cost

Sep 15, 2026, 5:00 AM · Ars Technica

Image: Ars Technica

Mozilla’s new open-source AI report says the closed-frontier lead is about 4.4 months — and often five times the per-task price — so most routine work should default to open weights.

Why it matters

Mozilla’s State of Open Source AI report, shared with Ars ahead of a September 15 publication, finds the gap between top US closed models and leading Chinese open-weights models has closed to roughly 4.4 months. Moonshot’s Kimi K3 trails Anthropic’s Fable 5 by about three points on Artificial Analysis’s composite index while costing about 30% as much.

CTO Raffi Krikorian argues closed models earn their premium mainly for expert professional work, high-intensity retrieval, and long context — a workload decision, not an org-wide religion. DoorDash already routes routine work to Kimi and harder tasks to Fable.

We’re reading this as the end of “frontier for everything” as a default architecture.

From the desk

We’re watching buyers get sharper. METR-style time horizons make the premium concrete: if open models reliably handle a seven-hour job, closed models handle about twelve — a 1.7× length edge that open models absorb in roughly four months. Tasks under eight hours are contested ground for cheaper open models; beyond twelve hours, nobody is generally reliable yet.

Vals AI’s neutral-harness work on Terminal-Bench 2.1 sharpens the cost story further: Z.ai’s GLM 5.2 scored within a point of Claude Opus 4.7/4.8 at about one-fifth the per-task cost. Harnesses still muddy comparisons when labs ship custom tool stacks, which is why neutral evals matter.

OpenRouter token volumes already skew open — eight of the top ten models by August 2026 volume were open weights — even as older revenue studies showed closed models taking most dollars. Krikorian’s concentration warning is the strategic bite: most of the open models the world runs are Chinese, running an Android-like playbook. Useful open models deserve adoption for routine work. The downside if the West stays out of the open lane is default infrastructure set elsewhere. I’m watching public-compute reference models, foundation-held evals, and whether US/EU labs compete openly instead of only behind APIs.

Context

Open-weights still usually omit training data, pipelines, and training code. Closed vendors sell compliance packaging, support, and accountability that many orgs still need staff to replicate. Krikorian cites Switzerland’s Apertus as an example of national compute producing fully open reference models.

Who feels it

Engineering leaders
Default open for sub-eight-hour and routine workloads; pay closed only when a deadline buys the four-month edge.
Procurement
Price-to-capability spreads of ~5× invite portfolio routing, not single-vendor lock-in.
Policy
Concentration of open-weight leadership in China reframes industrial strategy beyond closed-frontier races.

What to watch

  1. Whether open-model revenue share rises with token share.
  2. More neutral-harness benches that separate model quality from lab tooling.
  3. Public compute and foundation efforts that publish fully open reference stacks.

Read the original

Continue at the source.

Ars Technica