Liquid AI Releases Open-Weight d1-3B and d1-omni-600M: Multimodal Decision Models With Zero Output Tokens
Oct 7, 2026, 11:23 AM · MarkTechPost
Liquid AI's new models do not chat at all. They answer fixed questions with probabilities in milliseconds, and that unglamorous design may be what makes AI safe enough to run everywhere.
Why it matters
Liquid AI released Open d1, two open-weight decision models. d1-3B reads text and images. d1-omni-600M reads text paired with an image or with audio. Neither writes text. Instead, each takes a situation and a set of named questions and returns calibrated probabilities for the allowed answers in a single pass, with zero output tokens.
The pitch is speed and predictability. Liquid AI reports d1-3B answering one question in 8 milliseconds on an RTX 4090 and reading an image in 35 milliseconds on an Nvidia Jetson AGX Thor edge board. Target uses include routing, moderation, ticket triage, reranking, scoring and agent guardrails.
From the desk
We like this idea a lot, maybe more than the benchmark tables suggest. A huge share of real AI work is not conversation. It is yes or no, which team, how urgent, is this image acceptable. Using a chatty model for those jobs means paying for words nobody reads, then writing fragile code to parse them. A model that returns a typed, calibrated answer is easier to test, easier to audit and harder to talk into nonsense.
That makes d1 interesting as a guardrail layer. A fast checker that sits beside a larger agent and answers whether an action is allowed, with a probability attached, is the kind of boring infrastructure that makes autonomous systems safer.
Now the caveats. Liquid AI ran the Decision Index scorer itself, so its leading score under 10 billion parameters is not a leaderboard submission. The smaller omni model is an early research release with no published latency figures, and its audio training covered English requests only. That matters if it gets used for voice routing across languages and accents.
The downside of cheap, fast classifiers is scale. Moderation at the edge, frame-by-frame camera inspection and gesture control are useful. The same capability makes real-time automated judgment about people cheap to deploy everywhere, from content filters to surveillance cameras, and a confident probability can hide a biased decision. I'm watching how these models are evaluated on fairness, not just speed.
Context
d1-3B has 3.12 billion parameters and builds on Liquid's LFM2.5 vision-language model; d1-omni-600M has 587 million and builds on an encoder. Both are on Hugging Face with llama.cpp support. The LFM Open License allows free commercial use for companies under $10 million in annual revenue, which is more restrictive than Apache-licensed rivals.
Who feels it
- Developers
- Structured, millisecond answers for routing and classification remove a layer of output parsing from many apps.
- Edge and robotics teams
- Jetson-class latency makes on-device visual inspection and voice-command routing practical.
- Larger companies
- The license is free only below $10 million in revenue, so bigger firms need commercial terms.
What to watch
- Independent Decision Index results and leaderboard submissions
- Latency and multilingual audio results for d1-omni-600M
- Adoption of decision models as guardrails inside agent frameworks