JetBrains Releases Mellum2.1: A 12B MoE Open Model for Coding Agents
Oct 8, 2026, 9:17 AM · MarkTechPost
JetBrains shows how far reinforcement learning in real codebases can push a small open model. The benchmark jump is striking, and the numbers are still JetBrains' own.
Why it matters
JetBrains released Mellum2.1, an open coding model under Apache 2.0 that has 12 billion total parameters but activates only 2.5 billion per token. The architecture is the same as Mellum2. The gains, according to the release, come almost entirely from reinforcement learning in real software environments, where the model edits files in actual repositories and is rewarded when tests pass.
The headline result is agentic coding: SWE-bench Verified rose from 2.0 to 47.0 in JetBrains' own testing. Quantized builds start around 7 GB, which puts a capable coding sub-agent within reach of a single workstation.
From the desk
We like this release for what it says about where useful AI is heading. Not every coding task needs a frontier model in someone else's cloud. A small, fast, permissively licensed model that can explore a repo, make edits and check its own work is exactly what companies with private code and compliance worries have been asking for. Self-hosting keeps source code in-house, and that matters.
The training story is the more interesting part. Moving reinforcement learning from a short final step to the core of post-training, with millions of sandboxes across thousands of environments, is a recipe others can learn from. It suggests that environment design and verifiable rewards, not just parameter count, are where small models get their agentic skills.
Now the caution. Every score here is self-reported, and the release itself shows why that matters: JetBrains measured Qwen3.5-9B at 75.4 on LiveCodeBench v6, while Qwen's own card lists 65.6. Different test setups give different answers. Mellum2.1 also still trails Qwen3.5-9B on the hardest agentic tasks, including SWE-bench Pro and Terminal-Bench 2.1, and on knowledge-heavy reasoning.
The downside we keep returning to with coding agents is trust at scale. A cheap model that edits files and runs tests can be deployed as dozens of parallel workers, which is the point. It also multiplies the surface for subtle, test-passing bugs. I'm watching for independent evaluations and for how teams gate these agents before their changes merge.
Context
Mellum2.1 succeeds Mellum2 Thinking, which JetBrains open-sourced in June 2026. It is a reasoning model with a 131,072-token context window, aimed at agent workers, general reasoning and private self-hosted deployment. JetBrains also reports a lower HarmBench score than its predecessor, where lower is better.
Who feels it
- Developers
- A locally runnable coding sub-agent with permissive licensing, useful for private repos and offline work.
- Enterprises
- Self-hosting keeps code in-house, but teams will need their own evaluations rather than relying on vendor benchmarks.
- Open-model ecosystem
- Another data point that reinforcement learning in real environments can lift small models sharply on agentic tasks.
What to watch
- Independent SWE-bench and Terminal-Bench results from third parties
- Whether JetBrains integrates Mellum2.1 as a default agent worker in its IDEs
- Release of the multi-token prediction head for faster vLLM serving
- Competing small open coding models trained with similar recipes