Decision AI Models Explained: TypeSafe Jev vs Fastino GLiDE, GLiNER2.5-Decide and Open-Source Competitors
Oct 2, 2026, 8:31 PM · MarkTechPost

Decision models return choices and calibrated probabilities, not paragraphs — TypeSafe’s Jev set the category, and Fastino plus open-source forks arrived within weeks.
Why it matters
Asif Razzaq’s MarkTechPost explainer covers Decision AI: models that take text plus typed questions and return choices, scores or yes/no probabilities code can branch on. TypeSafe AI launched Jev after two years in stealth — a “System One model” for fast judgments. Within three weeks, Fastino Labs shipped rivals and open-source developers published Jev-style reproductions.
Jev’s primitives are Choice (up to 255 options), Score and Noul (Bernoulli-style probability). Questions run in parallel against one state; TypeSafe says it cannot return a type error because it never generates strings. Price: $0.042 per million input tokens, output free, 70–500 ms responses, 32K context on OpenRouter. Training uses Reinforcement Learning for Calibrated Decisions (RLCD) rather than preference-only RLHF.
On TypeSafe’s workflow evals, Jev hit 67.8% mean accuracy at $0.0004 per case and 0.4 seconds — matching Claude Sonnet 5 on accuracy at a tiny fraction of cost/latency, trailing a top frontier “sol” config at 74.1%.
From the desk
We’re watching a category split we wanted: stop asking a chat model to emit JSON you’ll regex, start asking a decision model for a typed judgment. Agent control flow, ticket triage, guardrails, routing and LLM-as-judge replacement are exactly where prose models are expensive and brittle.
We’re for this. Calibrated probabilities and parallel questions are how you build reliable agent gates without burning Sonnet-seconds on “should I interrupt?” Jev matching Sonnet accuracy on TypeSafe’s own workflows at roughly 300x lower cost per case (on their numbers) is the kind of efficiency that makes always-on agents economically sane.
The downside if decision models scale carelessly is silent misrouting: a wrong Choice with high confidence looks like certainty to code. Calibration helps only if it’s real out of distribution. These models also judge only the context they’re given — they won’t learn that one user hates morning pings unless you encode that state. And vendor evals that average GPT-6 Astra and Claude as labels inherit those models’ biases.
I’m watching Fastino’s benchmarks, GLiNER2.5-Decide’s joint decoding for non-contradictory safety labels, and whether open reproductions stay good enough that the category doesn’t collapse into one API.
Context
MarkTechPost maps use cases across agent control flow, classification, verification, eval/observability and search reranking — with Vercel, Arize, Langfuse, Buddy and OpenRouter already in the ecosystem narrative.
Who feels it
- Agent platform builders
- A cheap pre-LLM gate for continue/retry/escalate and tool choice without fragile JSON parsing.
- Eval & observability vendors
- Path to replace slow LLM-as-judge loops with calibrated decision calls on traces.
- Frontier LLM providers
- More traffic shifts to specialized decision models for branching; LLMs keep the prose and hard reasoning.
What to watch
- Independent evals outside TypeSafe’s workflow suite
- Fastino GLiDE and open-source Jev reproductions on shared benchmarks
- Production mute/act metrics when decision models gate proactive agents