SDSignal Desk

Cognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on FrontierCode at 64% Lower Cost

Sep 12, 2026, 4:56 PM · MarkTechPost

Image: MarkTechPost

Cognition’s SWE-2, RL-post-trained from Moonshot’s 2.8T Kimi K3, hits 50% on FrontierCode—within a point of Fable 5.1—at 64% lower cost, but only inside Devin.

Why it matters

Cognition, the company behind Devin, released SWE-2 as its most capable coding model yet. It’s post-trained with reinforcement learning from Kimi K3, Moonshot AI’s 2.8-trillion-parameter open model. Cognition reports 50.0% on FrontierCode 1.1 Main, within one point of Fable 5.1, at 64% lower cost—and it’s the first Cognition model with selectable reasoning-effort levels trained in a single RL run.

Deployability has a hard asterisk: no open weights, no standalone API. SWE-2 runs only inside Devin (Desktop and CLI now; Web and Fusion rolling out), free for paid tiers through October 10, 2026.

From the desk

We’re reading a capability story wrapped in a distribution lock.

On the numbers Cognition publishes, SWE-2 beats its K3 base across the board and leads Terminal-Bench 2.1 at 92.8%. It sits within a few points of GPT-6 Astra on several rows at a claimed quarter of the cost—while trailing badly on Terminal-Bench 4 (27.3% vs Fable 5.1’s 55.8% and Astra’s 57.9%). FrontierCode is Cognition’s own benchmark; rival scores are from Cognition’s harness. That’s not disqualifying, but it means we treat the leaderboard as vendor-reported until others replicate.

Behaviorally, the interesting claim is focus. SWE-1.7 over-explored; SWE-2 medium scores higher with 58% fewer turns and 81% less cost on FrontierCode, with mean steps collapsing from 127 to 53 at medium effort. First real edit arrives earlier. Stronger test coverage, resourcefulness when tools block, verification discipline when challenged—those are the agent habits that matter in production coding.

Training detail that earns attention: Pareto-informed cost penalties so all effort levels move the frontier in one RL run; length-weighted baselines; speculative decoding and low-precision kernels to keep train-serve mismatch down. Trust checks on politically sensitive China questions and framing-dependent vulnerability tests are disclosed—useful transparency even when the model stays closed.

We’re for cheaper, sharper coding agents when the evidence holds. The downside of Devin-only distribution is clear: capability gains that can’t be independently served, audited, or composed into other IDEs become platform gravity, not ecosystem progress. If this pattern scales, “best coding model” means “best inside one agent product.”

I’m watching independent FrontierCode and Terminal-Bench 4 replications—and whether an API ever appears.

Context

SWE-2 builds on the SWE-1.7 recipe (then from Kimi K2.7), now scaled to a base with nearly 3× the parameters. Cognition says RL still adds 5–6 points on many benchmarks atop K3. Public results are used where available; otherwise models run in native harnesses at best effort.

Who feels it

Devin customers
A stronger, cheaper in-product model with effort dials—timed with a free window through mid-October 2026 for paid tiers.
Competing coding-agent vendors
Pressure on cost/performance narratives, especially if Cognition’s harness numbers hold under outside eval.
Open-weights and API buyers
No path to run SWE-2 outside Devin; Kimi K3 remains the open base they can actually touch.
Benchmark maintainers
Vendor-owned benches plus harness-specific rival scores will draw scrutiny until third parties rerun.

What to watch

  1. Third-party reruns of FrontierCode and Terminal-Bench 4 comparing SWE-2’s Devin harness to rivals.
  2. Whether Cognition ever offers weights or a standalone API.
  3. How the October 10, 2026 free-access window converts into paid retention.
  4. Whether Terminal-Bench 4 remains the glaring weak spot in later checkpoints.

Read the original

Continue at the source.

MarkTechPost