Nace AI Open-Sources Drex 1.5: A 9B Decision Model That Scores Options, Not Text
Oct 9, 2026, 9:51 PM · MarkTechPost
Nace AI's Drex 1.5 puts an open-weight decision model on one GPU and roughly ties a closed leader. That is good news for builders, with a license and a benchmark caveat attached.
Why it matters
Nace.AI has open-sourced Drex 1.5, a 9-billion-parameter model that does not write text. It takes a state and a set of typed questions and returns a probability for each option, in one forward pass, according to MarkTechPost. Weights are on Hugging Face, and a hosted version is on OpenRouter at $0.04 per million input tokens with free output.
The headline result is parity. Nace reports 58.08 on the public Decision Index 0.3.1, against 57.96 for TypeSafe's closed Jev 1.13.0, the model that started this category. The two sit inside the board's tie band. Drex also speaks Jev's API format, so Nace says existing Jev clients can switch by changing a few environment variables. An open model that can drop into a closed model's slot changes the bargaining power in a young market.
From the desk
We are glad to see this. Decision models, the small systems that answer yes, no, pick one or rate it inside agents and back-end workflows, are the kind of AI that quietly does a lot of useful work. Having a credible open option matters, because these calls often touch sensitive internal data. Being able to run one locally, on a single 24 GB GPU or as a roughly 9.5 GB quantized file on a Mac, means teams do not have to ship every routing or approval decision to someone else's server.
The design also has a nice safety property. Drex can only answer with options the developer supplied, and no tokens are sampled, so there is nothing to hallucinate in the usual sense. For an agent gatekeeper, a model that cannot improvise is a feature. The long-document results are strong too: Nace reports 93.4% accuracy on 32K to 128K token inputs, and a clear drop when the same requests are cut to 8K, which suggests the long context is doing real work.
Now the asterisks. Nace says it trained on the official training splits of the index benchmarks and evaluated only on held-out splits. That is legitimate, but it means the model has seen the shape of these exact tasks. A top score on a board built from familiar task types tells us less about how it handles a company's own messy decisions. MarkTechPost says as much: results on new domains may differ. The Decision Index score also comes from Nace's own run of the official kit.
The weak spots are telling. On knowledge-heavy tests, Drex trails Jev badly: 45.4% on GPQA Diamond against 78.6%, and 58.7% on MMLU-Pro against 82.7%. Fine-grained sentiment is worse, at 7.4% per-review F1 on one aspect-sentiment test versus 29.5% for Jev. So this is a strong judge when the answer is in the document or the tool output, and a weak one when the answer depends on what the model knows. Builders should route accordingly.
Two practical points. Open weights here means a RAIL-M license with use restrictions, not a permissive license, so commercial teams need to read the terms. And local deployment through Ollama and llama.cpp requires Nace's own forks, not mainline builds, which adds maintenance risk.
Our read: a real step for open decision models and a useful check on closed pricing. The harm to watch is quiet over-trust. A cheap model that returns a confident probability with no explanation is easy to drop into approval flows where a human should still be looking. I'm watching for independent evaluations on unfamiliar tasks and whether Nace's changes land in mainline tooling.
Context
Decision models score a closed set of options rather than generating prose. Drex 1.5 is built on MiMo-V2.6-Distill-Qwen-9B, a distilled Qwen 3.5 9B backbone, with a separate pointer head that scores each option. Bespoke Labs' Nimble 9B v3 also sits within the Decision Index tie band, and Nace offers an agent skill for tools including Claude Code, Codex, Cursor and GitHub Copilot.
Who feels it
- Developers
- Jev-compatible requests and a local GGUF make it easy to test an open decision layer, though local runtimes need Nace's forks.
- Enterprises
- Self-hosting keeps routing and approval decisions on internal hardware, but the RAIL-M license needs legal review before commercial use.
- Closed decision-model vendors
- An open model roughly tied on the public index puts pressure on API pricing and lock-in.
- Agent builders
- Strong on tool and long-document decisions, weak on knowledge-heavy and nuanced sentiment calls, so routing by task type matters.
What to watch
- Independent Decision Index runs of Drex 1.5 beyond Nace's own
- Performance on decision tasks outside the benchmark training splits
- Whether Nace's llama.cpp and Ollama changes are merged upstream
- How TypeSafe responds on Jev pricing or capability
- Clarification of how the RAIL-M terms apply to commercial deployments