Comparison8 min read•

Jev vs Frontier LLMs: Choosing the Right Model for Each Step

A support workflow tags an incoming ticket as "technical," routes it to the wrong queue, and a billing dispute sits untouched for two days.

A support workflow tags an incoming ticket as "technical," routes it to the wrong queue, and a billing dispute sits untouched for two days. The team reaches for a chat model to do the classification, and now every ticket costs a few cents, takes a second or two, and occasionally returns an essay when it was asked for a category. Scale that to thousands of decisions an hour and you are paying frontier-model prices for a job that is really just picking one label from a known set.

TypeSafe released Jev on 15 September 2026 as the first "System One model," a class built to make fast, structured decisions instead of generating text. You hand it state and typed questions, it returns a label with a probability and stops. This guide is not a winner declaration. Frontier LLMs and decision models are complementary, and the interesting question is where the boundary between them falls in a real workflow.

That boundary is a content-modelling problem before it is an AI problem, because the quality of a typed judgement depends entirely on the quality of the state you hand it. Sanity, the AI-native content platform, matters here because its structured, queryable content is what lets you pass a typed model the exact field a question is about, not a wall of markup.

What is a System One model, and how does it differ from a frontier LLM?

A System One model is a class of model, introduced by TypeSafe with Jev on 15 September 2026, that consumes state and returns typed judgements rather than generating text. The name comes from Daniel Kahneman's fast, intuitive System 1, as against the slow, deliberative System 2 that chat models imitate. You pass Jev any text or JSON you want judged, plus typed questions you define in code, and it answers them and stops. It writes no prose, no code, no rationale, and no summary.

A frontier LLM like GPT-5.6 Sol or Claude Opus 5 does the opposite by design. It generates tokens one at a time, which is what makes it good at drafting, summarizing, writing code, and one-off complex reasoning, and also what makes it slower and more expensive per call. Jev is not autoregressive. A parallel sampler computes all outputs in a single pass, and TypeSafe reports end-to-end response times of 70 to 500 milliseconds.

The practical divide is what comes back. A frontier model returns words that your code has to parse and trust. Jev returns a decision drawn from a set of valid answers you fixed before the call. That difference shows up in cost, because TypeSafe bills input tokens only, at $0.042 per million, and leaves output unmetered on the grounds that a decision is not prose. It also shows up in reliability, which the next section covers. The two are not rivals so much as tools for different steps of the same job, and knowing which step you are on is the whole skill.

Illustration for Jev vs Frontier LLMs: Choosing the Right Model for Each Step
Illustration for Jev vs Frontier LLMs: Choosing the Right Model for Each Step

Which model should handle which step of a workflow?

Use a frontier LLM when you need generated text, generated code, a written rationale, or one-off complex reasoning. Use a decision model when you need a bounded decision made thousands of times with a confidence number attached. That is the entire decision rule, and most workflows contain both kinds of step rather than one or the other.

Jev offers three question types for the bounded case. Choice picks one option from a set you define, up to 255 options, and returns the winner, a probability for every option, and a confidence value; best practice is to include an explicit "other" option so it can say nothing fits rather than picking the closest wrong answer. Score places the state on an ordered scale of 2 to 10 word-described levels and returns a probability-weighted position that can land between levels, for severity, quality, urgency, or risk bands. Noul is a yes/no question returning a single number from 0 to 1, the probability that the answer is yes, with no separate confidence field because the probability is the certainty measure.

A content workflow is full of both kinds of step. Drafting a product description, translating a page into eight locales, or writing an editorial brief want a generative model, and in Sanity those are the jobs AI Assist and Agent Actions are built for. Tagging an article against a taxonomy, routing a submission, moderating a comment, or filtering results for relevance want a decision model, because the space of valid answers is bounded and known in advance. The mistake is using one model for the whole pipeline. The escalation cascade, which early adopters converged on within days, splits the work: a decision model classifies cheaply and at volume, code handles what it can, and anything below threshold goes to a frontier LLM or a human.

Does a decision model really eliminate hallucination?

A decision model eliminates format hallucination, not correctness errors, and the distinction is the whole story. Because the valid answers are fixed by your schema before the call, returning an off-schema value or a type error is structurally impossible. TypeSafe reports a 0% structured-output error rate and a 0% tool-call error rate. Jev will never invent a fourth category when you defined three, and it will never write an essay instead of returning a label.

It can still return the wrong valid value. A billing ticket filed under "technical" is a real mistake; it is simply a mistake inside the schema rather than outside it. TypeSafe is reasonably candid that the 0% figure is asserted from schema design rather than measured empirically, and any article repeating "cannot hallucinate" without that caveat is misleading. What you actually get is a guarantee about shape, plus a probability you can act on.

Contrast this with JSON mode and structured outputs on a chat model. Those constrain what the model writes, and they are genuinely useful, but the model is still generating tokens that your code must parse, and a malformed or truncated response remains possible. Returning a probability distribution over allowed answers is a different thing from emitting parseable text. One is a decision; the other is text that happens to look like a decision. For a content platform, that changes what you store. Instead of persisting a generated string and hoping it validates, you store a known label plus its confidence, which is a clean, queryable field you can filter, audit, and re-judge later.

How much cheaper and faster is a decision model, and what is the catch?

On TypeSafe's own numbers a decision model is dramatically cheaper and faster, and the catch is that it is a few points behind on accuracy and the accuracy figure is not measured against ground truth. TypeSafe reports Jev at roughly 40 to 200 times faster and 40 to 400 times cheaper than comparable frontier LLMs, with per-decision cost around $0.0004 against $0.0304 for GPT-5.6 Terra and $0.0836 for GPT-5.6 Sol. Those are self-reported figures that have not been independently reproduced.

The speed compounds when you batch. Questions in one request are evaluated independently and in parallel against one shared read of the same state, so a tenth question costs tokens but almost no extra time. TypeSafe's cookbook reports batching 13 questions into one call runs 12.2 times cheaper and 10 times faster than asking them separately, with identical answers. For a content pipeline that wants to tag, score, route, and moderate the same document, that is one call instead of four.

The honest catch is accuracy. On TypeSafe's four-workflow benchmark Jev scores 67.8%, level with GPT-5.6 Terra but behind GPT-5.6 Sol at 74.1% and Claude Opus 5 at 73.1%. That benchmark's "accuracy" measures agreement with GPT-6 Astra and Claude Fable 5.1 as consensus labels, not agreement with ground truth, so read it as agreement between models rather than correctness. The case for a decision model rests on cost and latency, not on being smarter than the frontier. There is also a floor: it is not a calculator, dates are text to it rather than ordered quantities, it is text-only so it cannot moderate images, and at low volume the integration effort can cost more than the inference it saves.

What does a content platform have to provide for a decision model to be useful?

A content platform has to let you hand a typed model the exact field a question is about, inside a bounded state budget, and this is where structured content earns its keep. Jev's total context is 64,000 tokens with a state budget of 32,000 tokens. A CMS that stores rich text as an opaque HTML blob can only pass a wall of markup as state, which wastes the budget and confuses the judgement. A platform that stores content as structured data can pass exactly the field, block, or section the question is about.

Sanity, the AI-native content platform, is well suited to this because content lives as structured data and rich text lives as Portable Text, a format whose blocks, marks, and annotations preserve structure across chunking and retrieval rather than collapsing into markup. GROQ lets you query for precisely the field or section you want to judge, so the state you assemble is the relevant slice and nothing else. This is an architectural fit, not a shipped integration: Sanity has announced no built-in Jev support, no partnership, and no System One feature, and the honest framing is what a content model must provide, not that a button exists.

The rest of the loop matters too. Content Lake real-time subscriptions give you a change event to re-judge against the moment content changes, so a label reflects the current state rather than a stale one. Functions give you serverless hooks to run tag-on-publish or moderate-on-publish pipelines that connect editors to a decision model. And because the returned label and confidence are just data, you store them as fields, filter on them in GROQ, and stage anything low-confidence through Content Releases for human review before it goes live.

What is the cascade architecture, and why did early adopters converge on it?

The cascade architecture is the pattern where a decision model classifies and routes at volume, code handles what it can, and the hard minority goes to a frontier LLM or a human. Early adopters converged on it within days of Jev's release because it puts each model on the step it is best at, and it makes the cost of the whole pipeline dominated by the cheap step.

The cleanest illustration is 1kpapers.com, where the author classified 1,018 research papers across 24 candidate topics for $0.08 total, at a median 256 milliseconds per paper. The summaries themselves were written by a generative model for $3.99. That is the boundary drawn in one project: a generative model produced the text, a decision model did the routing, and the routing cost eight cents. A second project, typesafe-computer-use, drives a Mac toward a plain-English goal at about $0.0002 per step, with OCR reading the screen, Jev picking the next action, and a writing model called only for free text, running roughly $0.003 per 12-step task against $0.40 to $0.90 for Opus 5 on raw screenshots.

Getting the cascade right in production is mostly discipline. Confidence-gated routing means the answer says what to do and the confidence decides whether it is safe to do automatically, with thresholds that rise as the cost and irreversibility of a mistake rise. Pin the versioned model ID, jev-1.13.0 rather than a moving alias, log the model reported in every response, and version question definitions alongside application code, or an alias can move and change behaviour with no release. Evaluate thresholds against reviewed labels covering easy, ambiguous, missing-context, and edge cases before trusting them. And never turn a timeout or a rate limit into a high-confidence default, because confidence only exists when a request actually succeeds; use a deterministic fallback or queue the case instead.

The eight-cent routing bill

On 1kpapers.com the author classified 1,018 research papers across 24 topics for a total of $0.08, at a median 256 milliseconds per paper, having already paid $3.99 to a generative model to write the summaries. That split is the cascade in one number: the generative model did the expensive creative work, the decision model did the high-volume routing, and the routing was effectively free. These are author-reported figures, not independently reproduced, but the shape of the win is the point. A workflow is full of both kinds of step, and paying frontier prices for the routing step is the waste worth removing.

How content platforms support feeding a typed decision model

FeatureSanityContentfulWordPressStrapi
Passing the exact field as stateGROQ queries return precisely the field, block, or section a question is about, so state stays within the 32,000-token budget without padding.GraphQL and REST return typed fields, so you can select the slice you need, though rich areas often arrive as combined entries.Post fields and meta are retrievable via REST, but core content ships as one rendered HTML body that is hard to slice cleanly.REST and GraphQL expose defined fields, so field-level selection works well for the schemas you build in the content-type editor.
Rich text as clean statePortable Text keeps blocks, marks, and annotations as structured data, so you hand a model structure rather than markup across chunking.Rich Text is a structured JSON document you can traverse, which passes as cleaner state than raw HTML in most cases.The classic editor stores HTML; blocks add structure, but the stored payload still leans on serialized markup for delivery.Rich text is typically stored as HTML or Markdown by default, so you often parse markup before it is usable as state.
Re-judging on changeContent Lake real-time subscriptions emit a change event on edit, giving a hook to re-run a decision so the label reflects current state.Webhooks fire on publish and change events, so you can trigger a re-classification externally when an entry updates.Hooks and REST polling can detect updates, but real-time change streams require extra plugins or custom infrastructure.Lifecycle hooks and webhooks fire on create and update, so you can trigger re-judging from your own service layer.
Running the pipeline serverlesslyFunctions run tag-on-publish or moderate-on-publish hooks in the platform, connecting the editor to a decision call without separate infra.App Framework and functions let you build automation, typically hosted and wired up as separate apps in the ecosystem.Automation runs through plugins or external cron and WP-Cron, so pipeline logic usually lives outside the core platform.Custom controllers, services, and plugins host pipeline logic, which you build and deploy as part of your own application.
Storing the returned label and confidenceThe label and confidence value are stored as typed fields you filter in GROQ, so the decision becomes queryable, first-class content.Custom fields hold the label and a number, and you query them via the content API like any other typed field.Post meta or custom fields can store the values, though querying by a numeric confidence often needs custom meta queries.Add fields to the content type for label and confidence, then filter through the API using your defined schema.
Routing low-confidence results to reviewContent Releases stage low-confidence items so a human reviews and schedules them before they go live, inside the editorial loop.Workflows in higher tiers gate content for review, so flagged items can wait on approval before publishing.Draft and pending-review statuses hold flagged items, with editorial workflow plugins adding finer review gates.Draft and publish states plus review plugins let you hold flagged entries, with the gating logic built in your app.
Auditing what the model sawContent Source Maps and Audit logs trace which content and version fed a decision, so you can reconstruct the state behind a label.Entry versioning and activity logs record changes, so you can reconstruct which revision existed at decision time.Post revisions track edits, but tying a specific stored state to an external model call usually needs custom logging.Draft history and your own audit tables record versions, so reconstruction depends on logging you implement.