AI Governance & Risk8 min read•

Auditing Decisions From a Model That Gives No Reasons

A compliance reviewer opens a decision log and finds a row that reads: choice "deny", confidence 0.71, timestamp. Nothing else.

A compliance reviewer opens a decision log and finds a row that reads: choice "deny", confidence 0.71, timestamp. Nothing else. The model that produced it, a System One model like TypeSafe AI's Jev, released 15 September 2026, returns probabilities and stops. It writes no prose, no rationale, no summary. So when someone asks "why did it deny this?", there is no answer to retrieve, because the model never wrote one. That is not a bug you can patch with a better prompt. It is what this class of model is: a parallel sampler that answers typed questions and emits no text of any kind.

The stakes are real for anyone routing money, access, or eligibility through a probability. A calibrated number is a filter, not evidence about a person, and "schema-valid" guarantees format, not correctness. Jev can return the wrong valid value at 0.9 confidence.

This article reframes the audit problem: you cannot recover a reason the model never produced, so the audit trail has to capture everything around the decision instead. That turns out to be a content-modelling problem before it is an AI problem, which is why it sits on a CMS site. Sanity, the AI Content Operating System, is the intelligent backend that lets you hand a typed model the exact field a question is about, and reproduce that exact state later.

Illustration for Auditing Decisions From a Model That Gives No Reasons
Illustration for Auditing Decisions From a Model That Gives No Reasons

Why can't you ask a System One model why it decided?

You cannot ask a System One model why it decided because the model produces no rationale to ask about. Jev, TypeSafe AI's first 'System One model', is not autoregressive. A parallel sampler computes all outputs in a single pass, trained with RLCD (Reinforcement Learning for Calibrated Decisions) rather than RLHF. You pass it state, meaning any text or JSON you want judged, plus typed questions you define in code, and it answers them and stops. It writes no prose, no code, no rationale, and no summary. There is nothing to parse for a because-clause, because none is emitted.

This is a deliberate trade, not an oversight. The category name borrows Daniel Kahneman's System 1 (fast, intuitive) against System 2 (slow, deliberate). A frontier model can narrate a chain of thought, which reads like an explanation, but that narration is generated text that may or may not reflect the computation that produced the answer, and it costs latency and money to produce. Jev drops the narration entirely. TypeSafe reports 70 to 500 milliseconds end to end and $0.042 per million input tokens with output free, which is only possible because there is no output to meter beyond the structured verdict.

The consequence for governance is blunt. If your definition of 'auditable' is 'the system can tell me its reasoning', a System One model fails that test by construction, and no amount of logging changes it. The workable definition is different: a decision is auditable when you can reproduce it and interrogate it from the outside. That reframing is the whole job, and it depends far more on how you store the state and the verdict than on the model itself.

What must an audit trail capture when the model gives no reasons?

When the model gives no reasons, the audit trail must capture everything the model cannot say, so that the decision can be reconstructed and challenged after the fact. Concretely, seven things per decision: the exact state the model saw (not a later version of the content), the question definition and its version, the model version reported in the response, every option's probability, the confidence value, the threshold in force at the time, the route taken (auto-approve, auto-deny, or escalate), and any human override with who made it.

Each field earns its place. The exact state matters because content changes; if you log a document ID and re-fetch it a month later, you are auditing a different input than the one that was judged. The question definition and its version matter because option lists and Noul criteria drift as the taxonomy evolves, and a Choice over five options is not comparable to the same question after a sixth was added. The full probability distribution matters more than the winner alone: a 'deny' at 0.34 versus 0.33 over 'allow' is a coin flip wearing a verdict, and only the distribution reveals it. The threshold in force records the policy, not just the outcome, so you can tell whether a decision was auto-actioned or sent to review.

Production hygiene follows from this. Pin the versioned model ID (jev-1.13.0, not just jev-latest), log the model the response actually reports, and version question definitions alongside application code. Evaluate thresholds against reviewed labels before you trust them, and never convert a timeout or a rate limit into a high-confidence default, because a silent failure that reads as 0.95 is the most dangerous log entry you can write.

How does content structure make a decision reproducible?

Content structure makes a decision reproducible by letting you reconstruct the exact state the model saw at the exact revision it saw it. This is the hinge of the whole audit problem, and it is a content-modelling question, not an AI question. If state is assembled from specific fields at a specific document revision, the decision can be replayed later; if state is a scraped HTML page, it cannot, because the page has changed and you cannot recover what it contained at judgement time.

The quality of a typed judgement depends entirely on the state it is handed. Jev works inside a 32,000-token state budget, so you want to pass exactly the field, block, or section a question concerns. An HTML blob can only hand it a wall of markup, most of which is navigation and styling noise that competes with the signal. Structured content lets you project one description field, one price, and one status flag as the state, which is both cheaper and cleaner. In Sanity, this is a GROQ query that selects the precise fields, so the state is a small, deterministic object rather than a rendered page.

Reproducibility then comes from revision addressing. Sanity's document version history and Content Releases let you point at content as it existed at a specific revision, so replaying a past decision means fetching the same field values, not today's. As one Sanity engineering principle puts it about tools feeding models, return schema-shaped responses the model can pass straight through, because a tool that returns prose forces paraphrasing and paraphrasing is where facts go to die. The same logic governs state: hand the model structured fields, log those exact fields, and the decision stops being a black box and becomes a repeatable experiment.

How do you answer 'why' without a rationale from the model?

You answer 'why' without a model-supplied rationale by interrogating the decision from the outside, using three techniques that do not require the model to explain itself. First, replay the exact state with the deciding field removed or changed and watch what moves. If deleting a single sentence from the state flips a 'deny' to an 'allow', you have found the load-bearing input, which is a far more honest explanation than any generated chain of thought. This only works if you logged the exact state, which is why the previous section matters.

Second, inspect how the probability mass was spread across options. Because a Choice returns a probability for every option, not just the winner, you can see whether the model was decisive (0.9 on one option, near zero on the rest) or genuinely torn (three options within a few points). A confident wrong answer and an uncertain right answer look identical if you only store the label, and completely different if you store the distribution. This is also why an explicit 'other' option is best practice: it gives the model somewhere to put mass when nothing fits, instead of forcing it to pick the closest wrong answer.

Third, compare the decision against reviewed labels. You cannot know if a probability is well calibrated in the abstract; you learn it by holding a sample of decisions against human-reviewed ground truth and seeing whether 0.8-confidence decisions are right about 80 percent of the time. Sanity's Workflows (in beta) model this reviewing process as data, so a question like 'what auto-actioned below threshold without human review' becomes one GROQ query, because the process left a trail in the content repository rather than in a separate system. Audits become queries, which is the property a governance-minded reader cares about most.

What are the limits of auditing a probability, and where must a human sign?

The hard limit of auditing a probability is that a probability is a filter, not evidence about a person, and no audit trail changes that. A 0.71 confidence that a resume matches a role is a routing signal, not a finding of fact about the candidate, and treating it as evidence is a category error that a clean log will not save you from. High-stakes and irreversible decisions need a human decision-maker recorded as such, with the probability as one input among several, never as the decision itself.

'Cannot hallucinate' does not rescue you here. That phrase is a guarantee about format, not correctness: valid answers are fixed by the schema before the call, so an off-schema value or a type error is structurally impossible, but Jev can still return the wrong valid value with high confidence. Critics also note that grammar-constrained decoding in ordinary generative models (structured-output modes, libraries such as Outlines) already drives schema violations near zero, so the format guarantee alone is not the novel part; the calibrated probability over every option, the single-pass speed, and the cost are. Attribute the numbers honestly: on TypeSafe's own four-workflow benchmark Jev scores 67.8 percent, behind GPT-5.6 Sol at 74.1 percent and Claude Opus 5 at 73.1 percent, and that 'accuracy' is agreement with two frontier models used as consensus labels, not ground truth.

Prompt injection is the other live risk. When attacker-controlled text sits inside the state, injected instructions can flip an allow into a deny or the reverse. A decision model shrinks what an attacker can make the system do, because the output space is your fixed option set rather than free text, but it does not make hostile state safe. Your audit trail should therefore record the provenance of state, so a reviewer can tell trusted fields from untrusted user input when a decision looks wrong.

How do confidence-gated routing and escalation actually work in production?

Confidence-gated routing works by separating the answer from the decision to act on it: the verdict says what to do, and the confidence decides whether it is safe to do automatically. Above a threshold, the system acts; below it, the decision escalates to a frontier model or a human. Thresholds are not universal numbers you can copy from a blog post. They rise with the cost and irreversibility of a mistake, so auto-deleting a page might auto-action at 0.85 while auto-denying a loan should escalate almost everything, and the right cutoff is found by evaluating against reviewed labels, not guessed.

The production shape is an escalation cascade. A System One model handles the confident majority at volume and low cost, and a slower, more expensive frontier model or a human handles the uncertain minority. This is where the two model classes stop competing and start composing: Jev's speed and price make it viable to judge everything, and a frontier model's ability to emit a rationale makes it the right tool for the small fraction of cases that a person will actually read. You get breadth cheaply and depth only where it is needed.

Making this auditable means the routing itself must be content you can govern, not logic buried in code. Sanity's approach of modelling process as data applies directly: workflows are defined in TypeScript, versioned and deployed like the rest of your code, so the process cannot drift from what is written down, and the threshold in force at any moment is recorded rather than reconstructed. Functions can re-judge content on change, so when a document is edited the decision is refreshed and the new verdict logged alongside the old, and Content Releases let you stage a threshold change and review it before it ships, the same way you stage a website change. Governance stops being a spreadsheet and becomes a query over your own repository.

Can it be self-hosted, and what does the lock-in picture look like?

A System One model like Jev cannot be self-hosted today. It is a closed managed API, not open weights, and there is no published paper describing the architecture beyond TypeSafe's own claims. Access has moved fast and unpredictably: it launched 15 September 2026 behind a waitlist, appeared on the Vercel AI Gateway and OpenRouter within a day or two, was called 'available to everyone' on 20 September, then paused new signups on 22 September citing capacity. Anyone planning a production dependency should treat availability as volatile and design for it.

That opacity is a governance fact, not a footnote. You cannot inspect the weights, you cannot reproduce the benchmark, and every performance and pricing figure is self-reported, so the honest posture is to verify calibration yourself against your own reviewed labels rather than trusting the vendor's numbers. The dependency is also a single closed endpoint, which means your continuity plan is the escalation cascade doing double duty: if the API is unavailable, decisions route to the human or frontier path rather than defaulting to a fabricated high-confidence answer.

The part you can own is the state and the record. This is where the content platform, not the model vendor, is the durable asset. Sanity, the AI Content Operating System, sits at the layer legacy CMSes never reached: they stop at publishing, while Sanity operates content end to end, which here means the same structured content that gets published is the content you project as state and the record you store the verdict against. Your content model already defines answer spaces (enumerated fields, references to taxonomy documents, content types, and workflow states are Choice sets that exist before anyone writes a prompt), and Content Lake, GROQ, document version history, and Content Releases are generally available surfaces you control regardless of what happens to any one model API. Swap the model, keep the audit trail.

How content platforms support auditing a reasonless decision

FeatureSanityContentfulWordPressStrapi
Passing the exact field as stateGROQ projects the precise field, block, or section as a small deterministic object, so state fits the 32,000-token budget without markup noise.Structured fields via the Content Delivery API let you select specific fields, so field-level state assembly is workable through the API.Core content lives in the post_content HTML blob, so passing one clean field means parsing markup or adding custom fields and meta.Code-defined content types expose fields over REST or GraphQL, so selecting a single field as state is straightforward.
Deriving option sets from the content modelEnumerated fields, references to taxonomy documents, content types, and workflow states are Choice sets that already exist in the schema.Validation lists and reference fields define enumerations that can seed option sets, mapped by hand into the question definition.Taxonomies exist (categories, tags), but free-form fields and plugin sprawl make a bounded, authoritative option set harder to pin down.Enumeration fields and relations in code-defined types give clear option sets, defined in the schema you author.
Reproducing state at the exact revisionDocument version history plus Content Releases address content at a specific revision, so a past decision replays against the values it actually saw.Versioning and scheduled releases exist; reconstructing the exact field snapshot a decision saw is possible with API work and careful timestamping.The revisions table stores prior post states, so history exists, though reassembling field-addressable state from a revision takes custom work.Draft and publish plus a history plugin can capture revisions, but field-level revision replay for state is largely self-built.
Storing the verdict, probabilities, and thresholdThe returned label, full distribution, confidence, and threshold store as structured fields on a decision document, queryable later with GROQ.Decision records fit as entries or metadata; querying distributions across records is doable via the API with your own indexing.Verdicts store as custom fields or a custom table; querying probability distributions at scale usually means custom SQL or an external store.A decision content type can hold the verdict and probabilities; cross-record analysis is built with your own queries or reporting layer.
Routing low confidence to reviewWorkflows (beta) model review as data and Functions re-judge on change, so 'what auto-actioned below threshold without review' is one GROQ query.Scheduled workflows and App Framework extensions in fixed slots can route items for review, wired up through the App Framework.Pending/review states via plugins can hold low-confidence items, with the routing logic living in custom plugin code.Draft/publish states plus custom controllers or lifecycle hooks can gate low-confidence decisions to a review queue you build.
Recording what the model sawBecause state is assembled from logged fields at a known revision, the audit record is field-addressable and reproducible by default, not bolted on.What the model saw can be recorded by snapshotting the fields sent; reproducibility depends on capturing the revision alongside the payload.Recording the exact input means capturing the rendered or parsed blob at decision time, since re-fetching later returns a changed page.Snapshotting the field values sent as state is possible; tying that snapshot to a specific revision is added application logic.