How to Feed CMS Content to Jev: Structuring State for System One Models
You wire Jev into your support triage, point it at your knowledge base, and the accuracy is embarrassing. The model files a billing question under "technical" and grades a critical outage as "low severity." You blame the model.
You wire Jev into your support triage, point it at your knowledge base, and the accuracy is embarrassing. The model files a billing question under "technical" and grades a critical outage as "low severity." You blame the model. The model is fine. You handed it a 40,000-character HTML blob that overflowed the 32,000-token state budget and got truncated mid-article, so it judged half a document it could not fully read. Jev, released by TypeSafe AI on 15 September 2026 as the first "System One model," consumes state and returns typed judgements, and the quality of the judgement depends entirely on the quality of the state you feed it. That makes feeding a content repository to a decision model a content-modelling problem before it is an AI problem. This guide works through the practical core: how to query for exactly the fields a question concerns, flatten rich text into clean state, fit the token budget, and batch questions that share one read. Sanity, the AI-native content platform, stores content as structured, queryable data rather than opaque markup, which is precisely what lets you hand a typed model the exact field, block, or section a question is about. No Sanity-Jev integration exists; this is a pattern you build, and the shape of your content decides whether it works.
Why is an opaque HTML blob the worst possible state to feed a decision model?
An HTML blob is the worst-case state for a System One model for three concrete reasons, and they compound. First, you cannot select part of it. If the question is about a product's return policy and that policy lives in one paragraph of a 6,000-word article stored as a single markup string, you have no addressable handle on that paragraph. You send the whole string or you write a parser. Second, markup burns budget you do not have. Jev's state budget is 32,000 tokens, and every div, span, class attribute, and inline style is tokens spent on presentation rather than meaning. A wall of nested markup can double the token cost of the same information expressed as clean text, which pushes real content past the truncation point. Third, and most damaging, the model sees presentation instead of meaning. Jev is not autoregressive; a parallel sampler reads the state once and computes every answer against that single read. When that read is polluted with layout, the signal it is judging is diluted before the question is even asked. WordPress is the canonical example: post content is stored as an HTML string, and even Gutenberg blocks are serialized into HTML comments inside that string, so the default unit you can hand a decision model is a wall of markup rather than an addressable field. Selecting one section as clean state means parsing HTML first, every time. Structured content solves exactly this. When content is stored as data, the smallest meaningful unit, a field, a block, a section, is directly addressable, and you send meaning without the markup tax. That is the whole game for state quality, and it is decided by your content model long before the API call.

How do you query for exactly the fields a question concerns instead of the whole document?
The first and biggest budget win is to stop sending documents and start sending fields. A typed question is narrow by construction: a Choice over 24 topics, a Score on an urgency scale, a Noul asking whether a claim needs legal review. Each one concerns a specific slice of a document, not the whole thing, so the state should be that slice. In Sanity you express this as a GROQ query that projects only the fields the question touches. GROQ mode queries the dataset at request time, which is exactly right when content is structured and consistent and the schema tells you where to look. A projection like *[_type == "supportArticle" && _id == $id]{ title, summary, "policy": body[style == "policy"] } returns clean, addressable data rather than a rendered page, so the state you assemble for Jev is measured in hundreds of tokens instead of thousands. Getting the field names right matters here more than teams expect. Retrieval fails in production with a recognizable shape: a field called body that is actually a slug, a hero that is a reference to a mediaAsset rather than an image. The model, or your query, needed to know the shape of the data, not just its types. A content model where the schema tells you where the return policy lives is a content model where you can hand Jev exactly the return policy. This is the same discipline behind pure structured query generally: you write the predicate, you get exactly what you asked for. GROQ, SQL, and GraphQL all share that property, and it is the property you want when the space of valid answers is bounded and you already know which field carries the evidence. Query narrow, send narrow, and the 32,000-token budget stops being a constraint you fight and becomes headroom you rarely touch.
How do you flatten Portable Text into clean state without losing structure?
Rich text is where the field-selection win can quietly unravel. You have carefully queried one section, but if that section arrives as a tangle of markup or a paragraph of prose you paraphrased on the way out, you have reintroduced the blob problem at a smaller scale. The goal is clean text for the state that still keeps block boundaries meaningful, because a paragraph break, a list, or a heading often carries the signal a question is judging. Sanity stores rich text as Portable Text, a structured JSON representation where marks, annotations, and blocks are addressable data rather than embedded HTML. That structure is what makes clean flattening possible: you walk the block array, emit the text of each block, preserve the boundaries you care about (headings, list items, a callout block), and drop the presentational noise. You are not paraphrasing, you are serializing. This matters because paraphrasing is where facts go to die. When Sanity's own team watched agents get built against structured content, the ones that worked passed schema-shaped data straight through, and the ones that struggled got a wall of text back and re-narrated it, badly. The same principle applies to state you feed a decision model: do not summarize the section into prose and then ask Jev to judge your summary, because now the judgement is only as good as your lossy rewrite. Serialize Portable Text to clean, boundary-preserving text and let Jev read the actual content. The reason this is achievable at all is architectural. Portable Text preserves structure across chunking and retrieval by design, so the block boundaries survive the trip into the state budget. An HTML blob gives you no boundaries to preserve; you get a string and a regular expression and a prayer.
When should you send a section instead of the whole page?
Send the section when the question is section-shaped, which is more often than instinct suggests. A page is a container for many independent claims: a product page holds a description, specs, a return policy, shipping terms, and reviews. If your question is "does this return policy meet the 30-day standard," the specs and reviews are not context, they are noise that spends budget and dilutes the single read Jev performs against the state. Because a parallel sampler judges the whole state at once rather than reasoning step by step toward the relevant part, irrelevant content is not something the model learns to ignore; it is signal you paid to send. The rule is to match the granularity of the state to the granularity of the question. Section-shaped questions get sections. Field-shaped questions get fields. Only genuinely document-shaped questions, "is this whole article ready to publish," get the whole document, and even then you send the serialized content, not the rendered page. Structured content makes this granularity a query parameter rather than a parsing project. When your body is an array of Portable Text blocks tagged by role, selecting the policy section is a filter, and selecting the whole article is dropping the filter. The token math follows directly. TypeSafe reports batching many questions into one call is dramatically cheaper because they share one read, but that read is only cheap if the state is small. A 2,000-token section costs a fraction of a 20,000-token page across thousands of documents, and at $0.042 per million input tokens the difference between sending sections and sending pages is the difference between a project that pencils out and one that does not. Right-size the state to the question and you are paying for evidence, not for layout you queried by accident.
How do you batch many questions about one document into a single call?
Batching is the pattern that turns per-document judgement from expensive into nearly free, and it works because of how Jev reads state. Questions in one request are evaluated independently and in parallel against one shared read of the same state. A question does not condition on another question's answer, so asking ten questions about a document is not ten times the work; it is one read of the state and ten cheap judgements against it. TypeSafe's cookbook reports that batching 13 questions into one call runs 12.2x cheaper and 10x faster than asking them separately, with identical answers. The practical consequence is that once you have paid to assemble and send a document's state, you should ask everything you want to know about that document in the same call. For a support article: a Choice for the topic across 24 candidates with an explicit "other" option, a Score for urgency on a five-level scale, a Noul for "needs legal review," a Noul for "contains a pricing claim," all against one read of the same clean state. The batching also disciplines your content model, because it rewards assembling the fullest sensible state once rather than re-querying per question. There is a caveat worth stating plainly. Batching shares a read within one call, but sending the same state across many separate calls still meters input tokens each time. If you re-judge every document nightly, you re-send every state nightly, and the bill scales with volume even though output is free. That is an argument for re-judging on change rather than on a schedule, which the next section covers. It is also an argument against the counting anti-pattern: Jev is not a calculator, so if you need a judgement per item in a list, ask one Noul per item, but do it inside a single batched call against one shared read rather than one call per item.
How do you re-judge content when it changes and store the result honestly?
A judgement is only as current as the last time you ran it, so the feeding loop has to react to change rather than run on a timer. Re-judging nightly re-sends every document's state every night and pays for content that did not move; re-judging on change sends only what changed, when it changed. Sanity Functions are the native primitive for this: small, single-purpose pieces of code running on Sanity's cloud infrastructure that react to content changes. Document functions fire on the create, update, and delete events for documents in a dataset, and because publishing a document fires an update on the published document, a publish is a change event you can hang a re-judgement on. The pattern is: a Function triggers on publish, runs the GROQ query that assembles the document's state, calls Jev with the batched questions, and writes the typed answers back onto fields. Storing the result honestly is its own discipline. Write back the full return, not just the winner: the chosen label, the probability for every option, and the confidence value, because confidence is what gates whether you act automatically. Store the model ID reported in the response too, because jev-latest is an alias that can move, and logging the versioned ID that actually answered is how you keep an alias from silently changing behaviour with no release on your side. When a result lands below your threshold, stage it rather than publishing it: Content Releases lets you hold a low-confidence label for human review before it goes live. And because the process left a trail in the content repository, the audit becomes a query. "What got auto-classified below 0.8 confidence last week" is one GROQ query when the label, confidence, and model ID live next to the content, rather than a join across a separate log store. Workflows modelled as data mean the process cannot drift from what is written down.
What does 'cannot hallucinate' actually protect you from, and what does it not?
"Cannot hallucinate" is the claim that sells System One models, and repeating it without a distinction is how you build a system that fails silently. Here is the precise version: valid answers are fixed by your schema before the call, so returning an off-schema value or a type error is structurally impossible. Jev will never invent a fourth category when you defined three, and it will never write an essay instead of picking one. TypeSafe reports a 0% structured-output error rate and a 0% tool-call error rate, though it is candid that this figure is asserted from schema design rather than measured empirically. That is a guarantee about format, not about correctness. Jev can still return the wrong valid value: a billing ticket filed under technical is a perfectly well-formed answer and a wrong one. The format guarantee removes an entire class of integration bugs, the ones where you parse a response and it is not the shape you expected, and that is genuinely valuable. It does nothing about whether the judgement is right, which is exactly why the state you feed it matters so much and why confidence gating is not optional. This is also the honest boundary of what to claim in production. A returned probability is a filter, not evidence; it tells you whether to act automatically, not whether the decision would hold up as an audit trail about a person. Thresholds should rise with the cost and irreversibility of the mistake, and you should evaluate them against reviewed labels that cover easy, ambiguous, missing-context, and edge cases before you trust them. Two failure modes deserve naming. Dates are text to Jev, not ordered quantities, so extract them with a Choice over enumerated options and compare them in code. And confidence only exists when a request succeeds, so never convert a timeout or a rate limit into a high-confidence default; use a deterministic fallback or queue the case. The model that cannot hallucinate a format can absolutely be handed a truncated, malformed, or half-read state, and then it confidently judges the wrong thing. Good state is the correctness story. The schema is only the format story.