AI Operations8 min read•

Evaluating a Decision Model on Your Own Content, Not the Vendor's Benchmark

A vendor benchmark tells you a model agreed with two frontier models 67.8% of the time.

A vendor benchmark tells you a model agreed with two frontier models 67.8% of the time. It tells you nothing about whether it will tag your product pages correctly, catch the drafts your editors reject, or route the ambiguous cases to a human instead of guessing. When TypeSafe shipped Jev on 15 September 2026 as the first "System One model," that gap became urgent: the launch number, by TypeSafe's own account, measures agreement with GPT-5.6 Sol and Claude Opus 5 used as consensus labels, not correctness on your content.

So the benchmark that matters is the one you build. This article shows how to evaluate a typed decision model against your own labels, your own questions, and your own edge cases, measuring accuracy, calibration, cost, and the share you can safely automate, then re-running it whenever anything changes.

That makes it a content problem before it is an AI problem. A typed judgement is only as good as the state you hand it, and Sanity, the Content Operating System for the AI era, is the intelligent backend whose structured, queryable content lets you pass a model the exact field a question is about instead of a wall of markup.

Why does a vendor benchmark tell you so little about your own workflow?

A vendor benchmark tells you little about your own workflow because it measures a different thing than the one you care about. On TypeSafe's own four-workflow benchmark, Jev scores 67.8%, behind GPT-5.6 Sol at 74.1% and Claude Opus 5 at 73.1%, and TypeSafe is explicit that this 'accuracy' is agreement with two frontier models used as consensus labels, not ground truth. Every performance, pricing, and latency figure in the launch materials is self-reported and has not been independently reproduced, so attribute each one to TypeSafe rather than treating it as fact.

Jev is a class of model that makes fast, structured decisions instead of generating text. You hand it 'state,' any text or JSON up to a 32,000-token budget, plus typed questions you define in code, and it answers and stops, writing no prose, no rationale, and no summary. TypeSafe reports 70 to 500 milliseconds end-to-end at $0.042 per million input tokens with output free. Those economics are the entire reason to care, because they make it plausible to judge every page, draft, or comment at volume.

But agreement with a bigger model is a cheap proxy, not correctness. Two frontier models can be confidently wrong together on exactly the cases that hurt you: a locale you publish in but they underweight, a product category with in-house jargon, a moderation rule specific to your community. A benchmark built on their consensus systematically hides your hardest cases. The only way to know whether a decision model works for you is to test it on decisions your editors have already made, and to measure it against labels you trust rather than labels a vendor generated. That evaluation is the subject of the rest of this guide.

Illustration for Evaluating a Decision Model on Your Own Content, Not the Vendor's Benchmark
Illustration for Evaluating a Decision Model on Your Own Content, Not the Vendor's Benchmark

What should you actually measure when evaluating a decision model?

You should measure five things, and headline accuracy is only the first. Start with accuracy against reviewed labels: on a frozen set of cases where a human already made the call, how often does the model return the same value? Because Jev's Choice type returns a probability for every option plus a confidence value, and its Score and Noul types return calibrated probabilities, you can measure far more than a raw hit rate.

The second metric is calibration: when the model says 0.8, does that answer come out right about 80% of the time? Calibration is what makes confidence-gated routing possible, and it is the property TypeSafe claims it trained for with RLCD, Reinforcement Learning for Calibrated Decisions. If the probabilities are honest, you can set a threshold and trust it. If they are not, a high confidence score is just a number. You verify this against your labels, in a reliability plot, not against the vendor's word.

Third, measure the share routed to review at each threshold. Raising the auto-accept bar sends more cases to a human, and that tradeoff has a real operational cost. Fourth, measure cost and latency per decision on your actual state sizes. As one Sanity eval writeup on agent evaluation puts it, 'the second metric after success rate is cost per conversation,' and a customer quote in that same corpus is blunt: 'It absolutely eats through credits.' Fifth, break every result down by content type and locale. An aggregate of 90% can hide a Spanish-language product category sitting at 60%. The number that matters is the worst slice you would ship, not the mean across all of them.

How do you build a label set from decisions your editors already make?

You build the label set from decisions your editors have already made, because those decisions are your ground truth and they cost nothing new to collect. Every existing tag on a document, every draft an editor rejected, every moderation outcome, every merge-or-redirect call on duplicate pages is a labeled example of exactly the judgement you want a model to make. You do not need a labeling project. You need to find the labels already sitting in your content history and turn them into an evaluation set.

This is where a structured content backend earns its place. In Sanity, those decisions leave a trail in the Content Lake: document history records who changed a field and when, workflow states record what was approved or sent back, and reference fields to taxonomy documents record which tag won. A single GROQ query can pull the published pages that carry a given tag alongside the drafts that were rejected for it, giving you positive and negative examples in one read. Because the content is structured, you project the exact field the question concerns rather than exporting a wall of markup.

Sampling matters as much as volume. Deliberately draw four kinds of case: easy ones the model should never miss, ambiguous ones where editors themselves disagreed, missing-context ones where the state genuinely lacks the information to decide, and edge cases at the boundary of a category. A benchmark stacked with easy cases flatters every model. The ambiguous and missing-context cases are where you learn whether the confidence values are honest, because a well-calibrated model should return low confidence exactly there. Include an explicit 'other' option in your Choice sets so the model can say nothing fits rather than picking the closest wrong answer, and count 'other' as a correct escalation, not a failure.

Why is passing the right state a content-modelling problem, not an AI problem?

Passing the right state is a content-modelling problem because the quality of a typed judgement depends entirely on what you hand the model, and Jev has a 32,000-token state budget to work inside. A decision about a single product's category should see that product's title, attributes, and description, not the entire page including navigation, footer, and unrelated modules. Feed it a wall of markup and you spend your token budget on noise while starving the question of signal. Structured content lets you pass exactly the field, block, or section the question concerns.

There is a deeper point here that the launch conversation mostly missed. A content schema already defines answer spaces. Enumerated fields, references to taxonomy documents, content types, and workflow states are Choice sets that exist before anyone writes a prompt. If your schema says a page's status is one of draft, in review, approved, or archived, that is a four-option Choice, already validated, already the set of legal answers. You are not inventing an answer space for the model; you are reading one your content model already enforces. That is why this class of model fits a structured backend so naturally: the schema is doing half the work.

Sanity, the AI Content Operating System, is built so that the state you pass and the option set you derive come from the same queryable content model. A GROQ projection pulls the specific block a question is about; the same schema that renders your site defines the enumerations a Choice question chooses among. As a Sanity writeup on agent tools observes, 'a tool that returns prose forces the model to paraphrase, and paraphrasing is where facts go to die.' The equivalent failure for a decision model is handing it prose when you could have handed it structure. When the state is a clean, typed object, the model reads it once and every question you attach reads the same shared state.

How do you turn a passing evaluation into confidence-gated routing in production?

You turn an evaluation into production by using the confidence values your evaluation validated to decide what runs automatically and what goes to a human. The pattern is confidence-gated routing: the answer says what to do, and confidence decides whether it is safe to do it without review. Thresholds rise with the cost and irreversibility of a mistake, so auto-publishing a low-stakes internal tag can accept a lower bar than auto-rejecting a contributor's submission. Above the threshold, the decision model acts at volume; below it, an escalation cascade sends the uncertain minority to a frontier model or a human.

The non-negotiable rule is that you set those thresholds from your own reliability data, never from the vendor's confidence claim. If your evaluation showed that 0.85 answers on Spanish product pages are right only 70% of the time, then 0.85 is not your auto-accept bar for that slice. And you never convert a timeout or a rate limit into a high-confidence default: an infrastructure failure is an escalation, not an approval. Jev returns no rationale, so for regulated or high-stakes decisions there is no model-authored answer to 'why this Choice.' That means the telemetry around the decision has to carry the audit trail the model does not provide.

In Sanity, Functions give you the trigger and Workflows give you the gate. Functions are small, single-purpose pieces of code that run on Sanity's cloud infrastructure and react to content changes, reading and writing the dataset and calling external services, so a document Function can call the decision endpoint on create or update, store the returned label and full probability distribution back on the document, and move a low-confidence result into a review stage. Workflows, currently in Beta, model the editorial process as versioned TypeScript so an agent 'has authority to advance or reject, and the stages are where a human keeps control,' and because the process left a trail in the content repository, 'audits become queries': 'what published without legal review' is one GROQ query, not a ticket to another system.

What does the format guarantee not protect you from?

The format guarantee does not protect you from being wrong. When TypeSafe says Jev 'cannot hallucinate,' it means the valid answers are fixed by the schema before the call, so an off-schema value or a type error is structurally impossible. That is a guarantee about format, not correctness. Jev can and will return the wrong valid value, confidently, and no schema constraint catches that. Worth noting too: grammar-constrained decoding, the technique behind structured-output modes and libraries such as Outlines, already drives schema violations near zero for ordinary generative models, so the format guarantee alone is not new. What is distinctive is the calibrated probability over every option at single-pass speed and cost, which is exactly why your own calibration measurement is the thing worth trusting.

There are hard limitations to design around rather than dismiss. Jev is text-only, so images and video are out of scope. It is unreliable at counting and treats dates as text rather than ordered quantities, so 'is this the most recent version' or 'which of these five is cheapest' is the wrong shape of question for it. It needs a bounded, known answer space and does not write the schema for you. And a probability is a filter, not evidence about a person, so a résumé-matching or moderation score gates a workflow; it is not a finding about an individual.

Security is the limitation builders underweight most. Prompt injection still works when attacker-controlled text sits in the state: injected instructions can flip an allow into a deny or the reverse. A decision model shrinks what an attacker can make your system do to your defined option set, which is a real narrowing, but it does not make hostile state safe. When your state includes user-submitted comments, contributor drafts, or scraped content, treat that text as untrusted, and put the same review gate in front of decisions made on it that you would put in front of any action driven by input you did not write. Your evaluation should include adversarial cases for exactly this reason.

How do you keep the evaluation honest as your content and the model change?

You keep the evaluation honest by treating it as a standing gate, not a one-time report, and re-running it whenever any input to the decision changes. Four things change underneath you: the question definition, the option set, the model version, and the content model itself. Rename a tag, add a category, and the Choice set your evaluation validated no longer matches the one running in production. Jev publishes jev-1.13.0 today behind the jev-latest route, but a route can move, so pin the versioned model ID, log the exact model reported in each response, and version your question definitions alongside application code so a change to either is a diff you can review.

This is the same discipline Sanity's own eval guidance describes for agents: 'a frozen set of representative conversations, twenty to start, each scored against a rubric you wrote. Run the suite on every model change, every prompt change, every tool change. The bar to ship anything to production is the eval bench staying green.' The decision-model version substitutes typed questions for conversations, but the gate is identical. Every failure you find becomes a new case in the frozen set, so the bench grows in exactly the places the model has hurt you before.

Storing the evaluation as structured content compounds the benefit. When the labeled cases, the returned probabilities, and the reviewer's notes live in the same backend that serves your content, 'the scores live next to the source content the agent queries,' and a reviewer's note can reference the exact document and field the model misjudged. Content Lake real-time subscriptions and Functions let a change to a content type re-trigger the relevant slice of the evaluation automatically, so drift surfaces as a red bench rather than a silent regression discovered in production. The vendor benchmark froze on 15 September 2026. Your content did not, which is precisely why the benchmark that governs your rollout has to be one you own and keep running.

How content platforms support evaluating and running a typed decision model

FeatureSanityContentfulStrapi + payload-aiWordPress
Pass the exact field or block as stateGROQ projects the single field, block, or section a question concerns, keeping the 32,000-token state budget on signal, not markup.API-first delivery returns structured entries, but the editorial UI and schema are coupled to storage, so isolating one arbitrary block as state is less direct.Structured fields via API, but assembling the precise state slice means custom queries and glue code you maintain per model.Content is stored as rendered pages and a wall of markup, so you can hand the model a blob, not the exact field a question concerns.
Derive the option set from the content modelEnumerated fields, references to taxonomy documents, and workflow states are validated Choice sets that already exist in the schema before any prompt.Enumerations exist in content types, but tapping schema and business logic to build an option set is a custom-code exercise.Schema defines enum fields, but deriving and syncing a Choice set from them is app-side work, not a platform primitive.Taxonomies exist as categories and tags, but they are loosely structured, so option sets need cleanup before a model can trust them.
Re-judge automatically on content changeFunctions react to create and update events on the Content Lake, calling the decision endpoint the moment a document changes and writing the result back.The App Framework and custom code can re-judge on change, but freshness and the webhook plumbing are a roadmap item you own and maintain.Re-judging on change means wiring your own webhooks and index freshness; AI is plugin-bolted-on rather than native.Post-save hooks exist, but keeping a decision fresh across edits is bespoke plugin work with no managed event pipeline.
Store the returned label and full probabilitiesA document Function writes the winning value, the per-option probability distribution, and the confidence back onto the document in one write.Custom code can store a label and probabilities back on the entry; the schema and freshness handling around it are yours to build.You can add fields for label and probabilities, but persistence, versioning, and retrieval are all app-side responsibilities.Post meta can hold a label, but structured storage of a full probability distribution per option is awkward and custom.
Route low-confidence results to human reviewWorkflows (Beta) model review stages as versioned TypeScript so an agent can advance or reject and humans keep control at the gate.Workflow features exist in enterprise tiers, but confidence-gated routing from a model score is custom integration work.No native review-gate primitive; routing a low-confidence decision to a human is an application you build and run.Editorial roles exist, but confidence-based routing to review is entirely custom plugin and code territory.
Record what the model saw for auditDocument history records field-level change over time, and because the process leaves a trail, audits like 'what published without review' are one GROQ query.Entry versioning exists, but reconstructing the exact state passed to a model at decision time is not a built-in query.Draft and publish states exist; a field-level audit of what a model saw at decision time is custom logging you maintain.Post revisions track content edits, but tying a stored decision to the precise state the model read is manual reconstruction.