Schema-Valid Isn't Correct: What 'Cannot Hallucinate' Really Means for Decision Models
You wired a decision model in front of your agent to gate tool calls, and it returned a perfectly valid "allow" for a request that should have been denied. The value was in your option set. The type was right. The schema validated.
You wired a decision model in front of your agent to gate tool calls, and it returned a perfectly valid "allow" for a request that should have been denied. The value was in your option set. The type was right. The schema validated. And the decision was wrong. That is the trap in the phrase "cannot hallucinate": it promises far less than it sounds like, and teams are shipping automated actions on the strength of a guarantee that was never about correctness.
TypeSafe's Jev, released on 15 September 2026 as the first "System One model," makes this concrete. It answers typed questions you define in code, returns a winning option with a probability for every choice, and writes no prose. Its answer is always one of the options you defined. Whether that option is the right one is a separate question the schema cannot touch.
This article separates the two. Schema-valid means well-formed; correct means true, and only reviewed labels and honest calibration can tell you which you have. That distinction is a content-modelling problem before it is an AI one, which is why it belongs here. Sanity is the AI Content Operating System whose structured, queryable content lets you hand a typed model the exact field a question is about, instead of a wall of markup that buries the deciding fact.
What does "cannot hallucinate" actually guarantee?
"Cannot hallucinate" guarantees format, not correctness. When you call a System One model like Jev, the set of valid answers is fixed by the schema before the request goes out: a Choice returns one of the options you enumerated, a Score lands on your labelled scale, and a Noul returns a number between 0 and 1. An off-schema value or a type error is structurally impossible, because the model cannot emit anything the schema did not define. That is a real property, and it removes a whole class of parsing failures that plague free-text generation.
What it does not remove is the possibility of the wrong valid answer. Jev can return "allow" when the correct call was "deny," and both are schema-valid. It can place a support ticket at "low priority" when it should be "urgent," and the scale accepted both. The guarantee is about the shape of the output, not its truth. Confusing the two is how teams end up automating on a promise that was never made.
The distinction matters most exactly where the stakes are highest. A prompt-injection screen, a résumé filter, or an SEO keep-merge-kill decision all return clean, typed values whether or not the judgement behind them is sound. A malformed answer at least announces that something broke. A wrong-but-valid answer looks identical to a correct one, which means you cannot catch it by validating the response. You can only catch it by comparing the answer to a label you trust, which is a measurement problem the schema does not solve for you.
Isn't constrained decoding already doing this for regular LLMs?
Yes, and the critics are right to say so. Constrained decoding is the technique behind structured-output modes and libraries such as Outlines: at each generation step, the sampler masks out any token that would break the target grammar or JSON schema, so an ordinary generative model can only emit output that validates. In practice this drives schema violations close to zero for models that were never designed as decision-makers. So the bare format guarantee, "the output will match the schema," is not new, and any launch messaging that leans on it as the headline is overselling a solved problem.
Conceding that point is what earns the right to state what is actually different. A System One model is not a generative model with a mask bolted on. Three things distinguish it. First, it returns a calibrated probability over every option, not just the winner, so "allow" at 0.51 and "allow" at 0.98 are different answers you can act on differently. Second, calibration is the training objective: Jev is trained with Reinforcement Learning for Calibrated Decisions rather than the human-preference tuning behind ordinary chat models, so the probabilities are meant to mean what they say. Third, it runs in a single parallel pass instead of generating token by token, which is where the speed and the flat cost of adding more questions come from.
The honest framing is that the format guarantee is table stakes, and the probability distribution is the product. If you only wanted schema-valid output, you already had it. What you did not have cheaply was a well-calibrated confidence number over a fixed option set, arriving fast enough to sit in front of every request rather than a sampled few.

Which failure modes survive a perfect schema?
A perfect schema stops malformed output and nothing else. Four failure modes survive it, and each one produces a clean, typed answer that looks trustworthy.
The first is the wrong valid option. The model picks a legitimate member of your set that happens to be incorrect for this state. No validator will flag it, because it broke no rule. Only a reviewed label reveals it.
The second is an option set that no longer fits the content. You defined the choices in code, then the world moved. A new content type, a renamed workflow state, or a product category that did not exist when you wrote the enum means the state now describes something none of your options names. The model still must return one of the options it has, so it returns the closest wrong one. This is why an explicit "other" option is best practice: it gives the model a way to say nothing fits instead of forcing a false pick. A schema that never offered "other" guarantees the model can never admit the gap.
The third is state that is missing the deciding fact. A typed judgement is only as good as what you hand it. If the field that determines the answer never made it into the 32,000-token state, the model is guessing over an incomplete picture, confidently, in the correct format. This is the failure that content modelling directly addresses, and we return to it below.
The fourth is untrusted text inside the state. When attacker-controlled content sits in what you pass, injected instructions can flip an "allow" into a "deny" or the reverse. A decision model shrinks what an attacker can make your system do to the options you defined, which is a genuine containment win, but it does not make hostile state safe. The judgement is still computed over text an adversary wrote.
How do you make "correct" measurable instead of assumed?
You make correctness measurable by comparing answers to labels you trust, before you trust the model. The schema tells you an answer is well-formed. Only a reviewed label tells you it is right, so the first move for any decision workflow is to assemble a set of states with known-good answers and check the model against them.
This is also where a subtle benchmark trap lives. TypeSafe reports that on its own four-workflow benchmark Jev scores 67.8%, behind GPT-5.6 Sol at 74.1% and Claude Opus 5 at 73.1%. Read the fine print: that "accuracy" is agreement with two frontier models used as consensus labels, not agreement with ground truth. Two large models can be confidently wrong together. So a headline accuracy number, even a self-reported one, is not a substitute for labels reviewed against your actual outcomes. Every performance and pricing figure here is self-reported by TypeSafe and has not been independently reproduced.
Calibration is the second measurable property, and it is the one that justifies the model existing. A well-calibrated model that says 0.9 should be right about ninety percent of the time at that confidence. You can check this directly: bucket your reviewed decisions by the confidence the model returned, then see whether the observed hit rate in each bucket matches the stated probability. If it does, the confidence number is doing real work and you can gate on it. If the buckets are miscalibrated, a high-confidence auto-approve threshold is unsafe no matter how clean the schema is. The point is that "correct" is not a property you assert; it is a rate you measure against reviewed labels, per option and per confidence band.
How should confidence gate what runs automatically?
Confidence should decide whether an answer is safe to act on without a human, and the threshold should rise with the cost and irreversibility of a mistake. The pattern that has emerged is confidence-gated routing: the answer says what to do, the confidence says whether to do it automatically, and anything below the bar escalates. A chat-moderation call that is cheap to reverse can auto-run at a lower threshold than a decision that deletes content or denies a person something.
The companion pattern is the escalation cascade. Run the decision model at volume across every request, then send only the uncertain minority to a frontier model or a human reviewer. This is what makes the economics work: Jev is priced at $0.042 per million input tokens with output unmetered, and TypeSafe reports 70 to 500 millisecond responses, so it is cheap enough to sit in front of everything while the expensive path handles the fraction that is genuinely hard. The confidence number is what sorts the two.
Three operational rules keep this honest. Evaluate your thresholds against reviewed labels before you trust them, because a threshold that looks safe in a demo can be miscalibrated on real traffic. Never convert a timeout or a rate limit into a high-confidence default: a failed call is not a confident "allow," and treating it as one turns an outage into a silent policy change. And pin the versioned model ID, log the model reported in each response, and version your question definitions alongside application code, so that when behaviour shifts you can tell whether the model changed, the questions changed, or the content did.
Why is this a content-modelling problem before it is an AI problem?
Because the quality of a typed judgement depends entirely on the state it is handed, and the state comes from your content. A decision model does not read your database or crawl your site; it judges exactly the text or JSON you pass it, inside a bounded budget. Get the state right and a small, fast model makes a sound call. Get it wrong, by omitting the deciding field or drowning it in markup, and no amount of schema rigor rescues the answer.
Structured content is what lets you pass the deciding fact and only the deciding fact. With GROQ you can project the exact field, block, or section a question concerns into a compact payload that fits comfortably inside a 32,000-token state budget. An HTML blob, by contrast, can only hand over a wall of markup and hope the model finds the fact inside it. Portable Text keeps rich text as annotated blocks rather than a flat string, so the structure that tells you which claim to judge survives all the way into the state. This is the difference between asking a question about a price and asking it about a whole page that mentions the price somewhere.
The option-set problem is the same problem from the other side. A content schema already defines answer spaces: enumerated fields, references to taxonomy documents, content types, and workflow states are Choice sets that exist before anyone writes a prompt. Sanity, the AI-native content platform, is built so those answer spaces, the deciding fields, and the machinery around them live together. Content Lake real-time subscriptions and Functions let you re-judge the moment a field changes so the answer never goes stale against the content. You store the returned label, per-option probabilities, and confidence as typed fields queryable beside the content they judge. Content Releases route a low-confidence answer to human review instead of publishing. And document history plus Audit logs record the state at decision time, which is the closest thing you get to an audit trail for a model that returns no rationale of its own.
How content platforms support a decision model's need for exact, judgeable state
| Feature | Sanity | Contentful | WordPress | Strapi |
|---|---|---|---|---|
| Pass the exact field as state | GROQ projects the precise field, block, or section into a compact payload, so a typed question sees only what it concerns inside a 32,000-token budget. | REST and GraphQL APIs return structured entries; field-level projection is possible but shaping a minimal payload takes more query assembly on the client. | Content is largely HTML in post_content; passing one clean fact usually means shipping a wall of markup or parsing it out yourself first. | REST and GraphQL expose structured fields; you can select fields per request, with payload shaping handled application-side. |
| Derive Choice option sets from the model | Enumerated fields, references to taxonomy documents, and content types are answer spaces that already exist in the schema before anyone writes a prompt. | Field validations and reference fields define enumerations and taxonomies you can read back to build option lists. | Taxonomies and categories exist but live across tables; assembling a clean option set means joining post meta and terms. | Enumeration fields and relations model taxonomies you can query to build option lists. |
| Re-judge on change | Content Lake real-time subscriptions and Functions fire on publish, so a decision can be recomputed the moment the deciding field changes. | Webhooks fire on publish and can trigger an external re-judgement pipeline you host and maintain. | Post save hooks or plugins can trigger outbound calls; you build and host the pipeline. | Lifecycle hooks on create and update can call out to a re-judgement service you run. |
| Store the returned label and probabilities | Add typed fields for the winning option, per-option probabilities, and confidence, then query them with GROQ alongside the content they judge. | Add fields to the content type to hold the label and probabilities; querying them beside source content is straightforward. | Store results in post meta; retrieval works but meta is loosely typed and easy to drift. | Add typed columns or a relation to hold results, queryable through the API. |
| Route low confidence to review | Content Releases and workflow states let a low-confidence answer stage for human review instead of publishing, governed in the Studio. | Workflows in higher tiers can gate publishing; a confidence threshold maps to a review step you configure. | Editorial workflow needs a plugin or custom post status to hold low-confidence items for review. | Draft and publish states plus custom review logic can hold uncertain items. |
| Record what the model saw | Document history and Audit logs capture the state at decision time, so an answer with no rationale still has a reconstructable input trail. | Entry versioning records content changes; correlating a stored decision to the exact version it read is doable with your own logging. | Revisions cover post content; capturing the exact judged state means adding your own logging. | Version history depends on configuration; capturing decision-time state is application work. |