AI Governance & Risk7 min read•

Confidence Thresholds and Review Queues: Shipping Model Decisions Safely

A model in a prototype does exactly one thing well: it works on the inputs you tried.

A model in a prototype does exactly one thing well: it works on the inputs you tried. Then it ships, a timeout returns nothing at 2am, your code reads that as a green light, and a refund gets auto-approved on a request the model never actually judged. Or a global 0.7 threshold that felt fine for tagging blog posts silently auto-closes fraud reviews at the same bar. The failure is rarely the model being wrong. It is the operations around the model treating a probability as a verdict, and treating an absent answer as a confident one.

This article reframes shipping a System One model like TypeSafe's Jev, released 15 September 2026, as an operations problem rather than a modelling one. Jev returns a typed decision plus a confidence signal and nothing else, so the work is deciding when code may act on that decision unattended, when a human sees it first, and what happens when no answer comes back at all.

That work is content-modelling before it is AI. A typed model is only as good as the state you hand it, and the review queue for low-confidence results is itself a content state. Sanity, the AI-native content platform, is relevant here for an architectural reason: structured, queryable content lets you pass a typed model the exact field a question is about, then model the review stage next to that content rather than in a separate system.

What does confidence actually mean on a typed decision?

Confidence is a calibration signal, not a correctness certificate. It tells you how peaked or spread the model's belief was, not whether the answer is right. Read it that way and you will build the wrong gate.

The three Jev question types report certainty differently, and the difference matters operationally. A Choice picks one option from a set you define (up to 255, and best practice is to include an explicit 'other' so the model can say nothing fits rather than pick the closest wrong answer). It returns the winning option, a probability for every option, and a separate confidence value that summarizes the shape of that whole distribution. A Score places the state on a worded ordered scale of 2 to 10 levels and returns a probability-weighted position that can land between levels, for example 1.035, plus probabilities and a confidence value. A Noul is a yes/no that returns a single number from 0 to 1: the probability the answer is yes. There is no separate confidence field on a Noul, because the probability already is the certainty measure. Asking for one is a category error.

So a high confidence value means the distribution was sharp, not that reality agrees with it. On TypeSafe's own four-workflow benchmark Jev reports 67.8% accuracy, level with GPT-5.6 Terra and behind GPT-5.6 Sol at 74.1%, but that 'accuracy' measures agreement with GPT-6 Astra and Claude Fable 5.1 as consensus labels, not agreement with ground truth. A confident wrong answer is entirely possible. Confidence earns you the right to skip a human on the easy cases; it never earns you the right to skip evaluation.

How does confidence-gated routing work in production?

Confidence-gated routing splits one decision into two: the answer says what to do, and the confidence decides whether code may do it unattended. Above the threshold, the action runs automatically. Below it, the case goes to a slower, more expensive path: a frontier LLM, a second model, or a human reviewer. This is the central pattern for running a decision model safely, and it is what turns a probability into an operational policy.

The hard part is that thresholds should rise with the cost and irreversibility of the mistake, which means one global threshold is almost always wrong. Auto-tagging a draft article at 0.75 is fine because the cost of a wrong tag is a two-second editorial fix. Auto-closing a security incident or auto-issuing a refund at the same 0.75 is reckless, because the mistake is expensive and hard to undo. Different questions in the same application will earn different bars, and reversible actions can sit far lower than destructive ones.

The economics are what make the cascade worth building. TypeSafe reports per-decision cost around $0.0004 against $0.0304 for GPT-5.6 Terra, and end-to-end response times of 70 to 500ms, both self-reported and not independently reproduced. That lets you classify cheaply and at high volume, let code handle everything above threshold, and reserve the expensive frontier model or the human only for the ambiguous remainder. Hassan El Mghari's 1kpapers project is the clearest public illustration: 1,018 papers classified across 24 topics for $0.08 total, after paying $3.99 to a generative model for the summaries the classifier judged. Different models for different parts of one workflow, gated on confidence.

Illustration for Confidence Thresholds and Review Queues: Shipping Model Decisions Safely
Illustration for Confidence Thresholds and Review Queues: Shipping Model Decisions Safely

Why does 'cannot hallucinate' need a careful reading?

'Cannot hallucinate' is a guarantee about format, not about correctness, and repeating it without that distinction is misleading. Because the valid answers are fixed by your schema before the call, returning an off-schema value or a type error is structurally impossible. Jev will never invent a fourth category when you defined three, and it will never write an essay instead of picking one. TypeSafe reports a 0% structured-output error rate and a 0% tool-call error rate on that basis.

What the guarantee does not cover is whether the valid value it returned is the right one. A billing ticket filed under 'technical' is a wrong answer that is perfectly on-schema. The model cannot break the shape of the answer; it can absolutely pick the wrong shape-valid answer. TypeSafe is reasonably candid that the 0% figure is asserted from schema design rather than measured empirically, which is the honest way to state it.

For an operations team this changes what you monitor. You stop worrying about parse failures and malformed output, because those genuinely cannot happen, and you spend that attention on label correctness instead. The discipline that keeps this honest is one borrowed from agent eval work: every failure becomes a rule. When a decision comes back wrong, the fix is rarely to the model. It is to the state you handed it, or to the threshold, or to the eval bench, so the same mistake cannot pass unseen twice. A model that cannot violate your schema is still a model that needs reviewed labels to prove it is choosing correctly within that schema.

How do you set a threshold you can defend?

You set it against reviewed labels, or you are guessing with a decimal point. A threshold picked from intuition looks precise and means nothing, because it was never checked against cases where being wrong is expensive. The number needs an evidence base before you let code act on it.

Build that base the way you would build an eval suite for any AI system: a frozen set of representative cases, scored against a rubric you wrote, run on every model change, prompt change, and question-definition change. The bar to ship anything is the bench staying green. Crucially, the set has to cover more than the easy path. Include easy cases, genuinely ambiguous ones, cases with missing context, edge cases, and above all the specific inputs that have already caused expensive mistakes. Every entry in the bench should trace to a real failure that happened once. A threshold validated only on clean inputs will pass in testing and fail on exactly the messy production traffic it was supposed to catch.

Structured content makes this bench tractable, because the scores can live next to the source content the model judged. When a decision is wrong, a reviewer's note can reference the exact document, field, and version the model saw, not a screenshot in a separate ticketing tool. In Sanity, the state you pass, the returned label, the confidence, and the reviewer's correction can all be documents in the same Content Lake, queryable with GROQ. 'Which auto-approved decisions later got overturned by a reviewer' becomes one query, and that query is how you learn whether your threshold is set too low.

Why is version discipline the failure that catches teams out?

Version discipline is the subtle one, because a moving alias can change the behaviour of your production system with no release on your side. If you call jev-latest or jev-preview, you are pointing at whatever version TypeSafe resolves that alias to today. Right now both resolve to jev-1.13.0, published 15 September 2026, the only published version. The day a new one ships, your thresholds were calibrated against a model you are no longer running, and nobody on your team merged a thing.

The fix is three habits. Pin the versioned model ID, jev-1.13.0, not the alias, so behaviour changes only when you decide it does. Log the model ID reported in every response, so you can prove after the fact which version made any given decision. And version your question definitions, the schemas and option sets themselves, alongside your application code, so a schema change and a threshold change land in the same reviewable pull request. A question with a new 'other' option, or a Score with a relabelled level, is a behaviour change every bit as real as swapping the model, and it deserves the same review.

This is where modelling process as data pays off. Sanity Workflows, in beta, lets you model the stages a document moves through in TypeScript, versioned and deployed like the rest of your code, so the process cannot drift from what is written down. When your question definitions and your review stages are both code and content under version control, a threshold change and a schema change are reviewable together, and 'what published without review' stays a single GROQ query rather than an archaeology project.

What happens to confidence when the request fails?

Confidence exists only when a request succeeds. A timeout, a rate limit, or an overload returns no answer and therefore no confidence, and the single most dangerous bug in this whole pattern is code that quietly turns that absence into a high-confidence default. That is how a refund gets auto-approved on a decision the model never made.

Handle failure deterministically instead. When a call does not come back, the case should hit a fixed fallback, deny by default, hold for review, retry with backoff, or queue, chosen for the cost of the action, never a defaulted 'yes'. Jev's published limits are worth designing against: total context 64,000 tokens with a 32,000-token state budget, a rate limit of 250,000 tokens per second and 1,200 requests per minute, on the endpoint POST https://api.typesafe.ai/v1/systemone. Sustained volume will hit those ceilings, and the queue is what absorbs the overflow without corrupting your outcomes.

This is exactly where the CMS earns its place, because a queued or failed case is a content state, not an exception log. Low-confidence results and no-answer cases both belong in a review queue, and a review queue is editorial workflow with a status field. Sanity Functions, which run on Content Lake and react to create and update events, can move a case into a 'needs review' state on the way in and out; Content Lake real-time subscriptions let a reviewer's dashboard update the moment a case lands. The failed decision and the low-confidence decision converge on the same queue, modelled as data next to the content they concern, which is where a human can actually see, correct, and clear them.

How content platforms support feeding, gating, and auditing a typed decision

FeatureSanityContentfulStrapi + LangChain.jsWordPress
Pass the exact field as stateGROQ selects the precise field, block, or section a question is about, so state stays inside the 32,000-token budget rather than dumping a whole document.Fields are retrievable, but assembling the exact slice a question needs is App Framework glue you build; presentation-first modelling tends toward whole entries.Possible via the REST or GraphQL layer, but selecting and shaping the exact field into model state is hand-built in your own orchestration code.Rich text is typically an HTML or markup blob, so the natural state is a wall of markup, not the specific field, which fights the token budget.
Rich text as clean statePortable Text stores rich text as structured blocks and marks, so you can hand the model text with its structure intact instead of parsing markup.Rich text fields serialize to a structured JSON tree, workable but presentation-oriented and less granular for isolating a single passage.Depends on the field type and plugins chosen; often a rich-text blob that needs cleaning before it is usable as model state.Post content is HTML, so state extraction means stripping and re-parsing markup before the model ever sees the words.
Re-judge on content changeFunctions react to create and update events on Content Lake, so a changed field can trigger a fresh decision automatically as part of the content pipeline.Webhooks plus App Framework code can trigger re-judging, but the trigger-to-decision pipeline is entirely yours to build and maintain.Lifecycle hooks and external orchestration can do it, assembled from community plugins and custom services rather than a native primitive.Save hooks exist via plugins, but wiring change events into a decision pipeline is bespoke PHP and external services.
Store the returned label and confidenceThe label, per-option probabilities, and confidence are just fields on a document in Content Lake, queryable with GROQ alongside the state that produced them.A returned label stores fine in a field; keeping probabilities and confidence together with the judged state is modelling work you design.Storable in custom fields or collections, but the schema tying state, label, and confidence together is entirely hand-modelled.Post meta can hold the values, but querying label plus confidence across many decisions is awkward without custom tables.
Route low-confidence results to reviewThe review queue is a content state: Workflows (beta) models the review stage as versioned data next to the content, so below-threshold cases become a status a reviewer clears.A review flow can be built in the App Framework, but it is UI-bound custom work, not a modelled content state you can query.Achievable with a status field and custom admin views, but the queue and its transitions are entirely self-assembled.Draft and pending statuses exist, but a confidence-driven review queue is a plugin-and-custom-code project.
Audit what the model sawState, label, confidence, and reviewer corrections live in one Content Lake; 'which auto-approved decisions were later overturned' is one GROQ query against a trail in the content itself.Audit data can be captured, but reconstructing the exact state, version, and outcome usually spans the CMS plus an external log.No native queryable audit of what the model saw; you build the logging and the queries or stitch an external store.Revisions track content edits, but linking a specific decision to the exact state and version it judged is not native.