.jpg)
Decision models are a new category of AI that returns bounded choices with confidence scores instead of generating text. TypeSafe AI's Jev is the first production-ready example, and community benchmarks show it cuts false positives in AI agent guardrails from 14.5% to 1.8% while improving recall from 9% to 81%. For teams building AI agents, this fills a gap language models alone cannot close.
In short
TypeSafe AI's Jev, released in September 2026, returns a bounded choice from a defined set of options, along with a confidence score and probability distribution. It does not generate text. The model is priced at $0.042 per million input tokens with 70 to 500ms response times, making it practical for real-time agent pipelines.
Language models conflate reasoning with generation. An LLM can write a convincing argument for any decision, whether or not the evidence supports it. A decision model separates the judgment from the narrative, which is what AI agent safety systems actually need.
An LLM is a debater and a decision model is a judge. Agent systems need both, serving different roles.
When AI agents run autonomously, they make decisions at every step: which tool to call, whether to proceed, whether to escalate to a human. Language models are poor at this because they generate text, not calibrated judgments. They say "yes" with the same confidence whether they are 99% sure or 51% sure.
In a simulation where 12 AI models ran food truck businesses, eight went bankrupt. The failure mode was consistent: models made confident decisions based on incomplete reasoning, then doubled down instead of escalating. This is the same pattern we see in production agent deployments.
The September AI lab breaches showed the same problem from the security side. When four AI labs were breached through nine intrusions, the root cause was agents executing actions without a decision gate checking whether the action was safe.
A decision model sits between the LLM's proposed action and the actual execution. It evaluates the action against defined criteria and returns a bounded decision (proceed, block, or escalate) with calibrated confidence.
Without a decision gate, agents execute first and ask questions never. The food truck simulation and the September breaches both trace back to the same missing piece: no judgment layer between proposal and execution.
In a standard agent architecture, the LLM proposes an action and the harness executes it. With Jev in the loop, the process adds a decision gate:
This is what AI harnesses are designed to support. The harness provides the plumbing for tool registration, memory, and orchestration. Jev adds the judgment layer that most harnesses have been missing.
The integration story is already moving. Vercel, Cloudflare, and LangChain have shipped connectors that let teams drop Jev into existing agent pipelines. You define the decision criteria, pass the context, and Jev returns a structured response.
The community harness benchmark tested Jev against standard LLM guardrails across three categories:
| Benchmark | Standard LLM Guardrails | Jev Decision Model |
|---|---|---|
| False positives | 14.5% | 1.8% |
| Recall@1 | 9% | 81% |
| InjecAgent coverage | Not tested | 41.1% at 100% accuracy |
The false positive drop is significant. Standard LLM guardrails flag roughly 1 in 7 safe actions as dangerous. Jev flags roughly 1 in 55. For teams running agents in production, that difference determines whether the agent is useful or constantly blocked.
The recall improvement matters just as much. Standard guardrails catch only 9% of actual violations. Jev catches 81%. That means 9 out of 10 dangerous actions slip through standard guardrails, while Jev catches 4 out of 5.
InjecAgent tests prompt injection resistance. Jev covers 41.1% of InjecAgent scenarios at 100% accuracy, meaning it blocks every injection it recognizes but does not yet recognize all of them.
Jev is early. The current version (1.13) has documented limitations:
| Limitation | What it means |
|---|---|
| Multi-hop questions | Struggles with decisions requiring chained reasoning steps |
| 3-SAT problems | Complex logical satisfiability exceeds current capacity |
| Prompt injection | Susceptible to injection attacks, same as LLMs |
| Context management | Long context windows degrade decision quality |
The prompt injection weakness is the most concerning. If an attacker can inject instructions into the context Jev evaluates, the decision model can be manipulated. This is the same attack surface that Claude Cowork and other agent platforms face. TypeSafe AI has proposed a context management approach that strips injected instructions before passing context to Jev, but this is still in development.
For teams considering Jev today, the assessment is straightforward. It improves AI agent guardrails accuracy but does not replace human review for high-stakes decisions. It serves as a decision gate. Human review still handles edge cases.
For teams already running AI agents in production, the path to adding decision models is incremental:
| Step | Action |
|---|---|
| Audit | Measure your current guardrail false positive rate and recall. The community benchmark gives you a baseline. |
| Map | Identify where your agent makes choices between tool calls. Each is a candidate for a decision gate. |
| Test | Pick one decision point, define criteria, and test Jev against existing guardrails using an LLM evaluation framework. |
| Threshold | Set confidence thresholds. Low-confidence decisions escalate to humans automatically. |
The cost profile supports experimentation. At $0.042 per million input tokens, running Jev as a decision gate on every agent action adds negligible cost compared to the LLM generating the action.
At NineTwoThree, we have shipped 150+ AI projects with a 97% success rate. 24 of 27 recent projects were ROI-positive. Our work with Consumer Reports delivered a $20M revenue lift through AI-powered search. K&L Wines saves 14+ hours per day at 99% SKU accuracy with our inventory AI. These results come from the same principle Jev applies: separating judgment from generation, and keeping humans in the loop where confidence is low.
If you are building AI agents and want to talk through how decision models fit into your architecture, reach out to our team. We build custom AI agents with the safety guardrails and evaluation frameworks that production systems require.
You can also explore our guide on effective guardrails for GenAI apps or our framework for tackling hallucinations with RAG.