Select, Don't Generate: Answering From CMS Content Without Writing Prose at Request Time
A support bot returns a confident, fluent paragraph about your return policy. It says 45 days. Your actual policy, approved by legal last quarter, says 30.
A support bot returns a confident, fluent paragraph about your return policy. It says 45 days. Your actual policy, approved by legal last quarter, says 30. No one wrote the sentence the customer saw, no one reviewed it, and you find out when the chargebacks start. This is the failure mode of putting a generative model in the query path: every answer is a fresh composition, unread until a user reads it, and wrong answers are new sentences rather than logged, fixable choices.
There is another way to answer from content you already have. Instead of generating prose at request time, you select it. The copy is written and reviewed at ingest or by editors; at request time a decision model picks the right existing answer, section, or document from a bounded set of candidates. No LLM sits between the question and the approved text. Latency and cost collapse to a single decision call, and every answer shown is copy someone signed off on.
This belongs on a CMS site because selection is a content-modeling problem before it is an AI problem. It only works when answers exist as addressable units, a field, a block, or a document, rather than paragraphs buried in page HTML. Sanity, the AI Content Operating System, stores content as structured data in Content Lake, which is what lets you hand a typed model the exact field a question concerns and choose among candidates the schema already defined.
What does 'select, don't generate' mean for answering from content?
Select, don't generate is a pattern where the prose is authored and reviewed ahead of time, and at request time a model chooses which existing answer to show rather than writing a new one. The reader always sees copy a human approved. The model's only job is a decision: given this question and these candidates, which one fits?
The contrast with retrieval-augmented generation is the point. In a typical RAG setup you retrieve passages, then a large language model synthesizes them into a fresh answer, and that answer can drift, hedge, combine two policies into one that never existed, or invent a number. The synthesis step is where hallucination lives. If you remove the synthesis step and let the model only pick from a fixed set of reviewed answers, the worst it can do is pick the wrong approved answer. That is a wrong choice you can log, reproduce, and fix, not a novel sentence no one has read.
This is where a new class of model becomes interesting. TypeSafe AI released Jev on 15 September 2026, billing it as the first 'System One model', a model that makes fast, structured decisions instead of generating text. It writes no prose, no code, no rationale, and no summary. You pass it state plus typed questions, it answers, and it stops. Its Choice question type picks one option from a set of up to 255 and returns a probability for every option plus a confidence value. That is exactly the shape of a selection problem: here are the candidate answers, which one applies?
Treat the pattern as the takeaway and any single model as an implementation detail. The architectural claim holds regardless of vendor: if the answer already exists and has been reviewed, do not regenerate it, select it. The rest of this guide is about what your content platform has to provide for that selection to be fast, correct, and auditable.
Why does selecting approved content beat generating it at request time?
Selecting approved content beats generating it because it moves the writing and the review to a moment when a human is in the loop, and leaves only a decision for request time. Three consequences follow, and each maps to a cost you are otherwise paying quietly.
First, correctness becomes a property of your content, not your model. Every answer a user sees is copy an editor or subject-matter owner approved. A generative model composes a new sentence on every request, so your policy correctness depends on the model getting the synthesis right every time, forever. Select the answer instead and the model cannot phrase your refund window as 45 days when the approved field says 30, because 45 days is not one of the candidates.
Second, latency and cost drop to a single decision call. A generative answer streams tokens, and long answers cost more and take longer. A decision model returns a label. TypeSafe reports Jev runs end to end in 70 to 500ms, priced at $0.042 per million input tokens with output unmetered and free, and evaluates multiple questions against one shared read of the state so a tenth question costs tokens but almost no extra time. Attribute those figures to TypeSafe; they are self-reported and have not been independently reproduced. The structural point survives the caveat: a bounded decision is cheaper and faster than open-ended generation.
Third, wrong answers are debuggable. When selection misfires you have a request, a candidate set, a chosen label, and a probability distribution over every option. You can replay it, see the second-place candidate, adjust the option list or the state you passed, and confirm the fix. A hallucinated paragraph gives you none of that. It is a one-off artifact with no schema behind it. That is the difference between a bug you can close and a liability you can only apologize for.

Where does selection fit, and where does it fail?
Selection fits anywhere the right answer already exists as a discrete, reviewed unit and the question maps to a bounded set of them. It fails the moment the honest answer is something no one has written yet and would require synthesis across several sources.
The strong fits are the boring, high-volume surfaces where wrong answers are expensive: FAQs, help centers, product specifications, and policy answers. 'What is your return window for opened electronics?' has a correct answer that lives in a field or a document, and the job is to route the question to it. Support and inbox triage, deciding which macro or which article applies, is the same shape. So is deciding which of three approved disclaimers a given product page needs. In all of these the answer space is knowable in advance, which is precisely what a Choice question requires. Best practice with Jev is to include an explicit 'other' option so the model can say nothing fits rather than picking the closest wrong answer, which for a help center means falling through to search or a human rather than showing a confidently irrelevant article.
The failures are just as clear. Open-ended questions that need genuine synthesis, 'compare these two plans for a team of five who travel a lot and summarize the tradeoffs', do not have a pre-written answer to select. Forcing selection there means either a maintenance nightmare of hand-written permutations or a wrong pick. Counting and date arithmetic are also poor fits; Jev is unreliable at counting and treats dates as text rather than ordered quantities, so 'which of these three plans is cheapest' is a job for a query, not a judgment. The rule of thumb: if a competent editor could have written the answer in advance and filed it under a heading, selection fits. If the answer only exists once the question is asked, you still need generation, ideally labeled as such so the reader knows the difference.
How do you build the candidate set the model chooses from?
You build the candidate set with a query, and the quality of that query decides everything downstream, because a decision model can only be as good as the options you hand it. Give it the wrong ten documents and even a perfect pick is wrong. The candidate set is the real engineering surface of this pattern.
There are three ways to narrow content to candidates, and the best answer uses all three. A pure structured query is exact: you write the predicate and get precisely what you asked for. In Sanity that is a GROQ filter over your documents, and it falls over the moment the user says 'something like X' or 'the cozy one', because that intent lives in vibes, not fields. Pure semantic search ranks by meaning but does not respect the filters that have to hold, so it will happily surface an out-of-stock product or a draft. Hybrid retrieval combines them, and the field has measured why that matters: Anthropic's contextual retrieval research found contextual embeddings cut top-20 retrieval failures by 35%, adding contextual BM25 took that to 49%, and adding reranking brought it to 67%. None of the three layers alone was enough.
In Sanity this runs as one GROQ query. Predicates do the filtering that must hold, category, price, stock location, then a score() pipeline blends a keyword match, boost([title] match text::query($queryText), 2), with text::semanticSimilarity($queryText) across the document, ordered by _score. One note that matters: text::semanticSimilarity() is only valid as an argument to score(). Semantic search ranks, it does not filter, so you narrow the candidate set with a filter first and rank what is left. For a catalog where the useful detail sits in prose fields, enable dataset embeddings and stay in GROQ mode rather than standing up a separate vector store. The candidate set that comes back, the top handful of documents or sections, is what becomes your Choice options. The option list is derived from content, not hand-maintained, which means it stays current as the content does.
Why does selection only work when your content is addressable?
Selection only works when answers exist as addressable units, a field, a block, or a document, because both halves of the pattern depend on it. You cannot build a candidate set out of paragraphs buried in page HTML, and you cannot pass a typed model the exact thing a question is about if the smallest thing you can reach is a rendered page.
Start with the state budget. Jev accepts a 64,000-token context with a 32,000-token state budget. When your content is structured data, you pass exactly the field, block, or section the question concerns and spend that budget on signal. When your content is a wall of markup, you pass navigation, styling, and boilerplate, and the useful sentence competes for room with a cookie banner. Structured content is what lets the state be small and relevant, which is what lets the judgment be good. This is a content-modeling problem before it is an AI problem.
The second half is the option set itself. A schema already defines answer spaces before anyone writes a prompt. Enumerated fields, references to taxonomy documents, content types, and workflow states are Choice sets that exist in the content model. Sanity's Content Lake stores content as structured data addressed by field, block, and document, so 'which policy applies' can be a Choice over the policy documents your schema references, and 'which severity tier is this ticket' can be a Choice over an enumerated field's values. You are not inventing option lists, you are reading them off the model you already maintain. That is the difference between a page-oriented CMS, where the addressable unit is usually a rendered page you have to parse, and a structured content platform, where the addressable unit is the field the question is actually about. Structure is not a nice-to-have here. It is the precondition.
How do you route low-confidence selections and audit what was shown?
You route low-confidence selections with a threshold and a fallback, and you audit them by storing the decision next to the content it chose from. Both are made possible by the fact that a decision model returns a calibrated probability, not just an answer.
Confidence-gated routing is the core production pattern. The answer says what to do; the confidence decides whether it is safe to do automatically. Thresholds rise with the cost and irreversibility of a mistake: a confident pick above your bar shows the approved answer directly, and anything below it falls through to a fallback, a human, a search page, or a generative model clearly labeled as such. This is an escalation cascade, the decision model handling volume and a frontier model or a person handling the uncertain minority. Evaluate thresholds against reviewed labels before trusting them, and never convert a timeout or a rate limit into a high-confidence default, because a fail-open decision system is worse than none. In Sanity, the fallthrough and the follow-up work can run on Functions, single-purpose TypeScript that reacts to create, update, and delete document events, so a low-confidence selection can open a review task rather than ship silently.
Auditing is where this pattern quietly wins over generation. A decision model gives you no rationale, which critics rightly raise for regulated or high-stakes decisions. Selection over CMS content answers that differently: both the candidate set and the chosen answer are addressable, versioned content you can log, alongside the returned label and the probability for every option. You store what the model saw, what it picked, and how sure it was, next to the content itself. With Workflows, still in beta, the audit becomes one GROQ query because the process left a trail in the content repository rather than a separate system. That does not substitute for model-level explainability, but it does mean 'what did we show, from which candidates, at what confidence' is a question your content platform can answer.
Schema-valid is not the same as correct
What about prompt injection and untrusted text in the state?
Prompt injection still works when attacker-controlled text sits in the state, so a decision model is not a security boundary on its own. Injected text can flip a selection the way it can flip any judgment, turning a should-show into a should-not or the reverse. What a decision model changes is the blast radius: it shrinks what an attacker can make the system do to your fixed option set. It does not make hostile state safe.
The selection-over-CMS-content pattern has a genuine structural advantage here, and it is worth stating precisely without overclaiming. When the candidate answers are editor-reviewed content and the state you judge is the content itself, the text the model reads is trusted by construction. There is no user-supplied prose in the state to carry an injection, because the state is your approved copy. That is a real mitigation for the common case of answering from a help center or a policy library. It is not a cure the moment user text enters the state, for example when the question itself is free-form user input or when you fold a customer message into what the model reads. At that point the untrusted text is back, and you handle it the way you handle any untrusted input: separate it from instructions, constrain the option set, and gate on confidence.
So the honest posture is layered. Keep the reviewed content as the state wherever you can, which is most of the time for this pattern. Where user text must enter, treat the decision as advisory and route anything consequential through confidence gating and human review. And remember the closed nature of a managed model like Jev, no open weights and no published paper, is a lock-in and opacity consideration of its own, separate from injection. Structured, reviewed content narrows the attack surface a great deal. It is a smaller target, not an invulnerable one, and building as though selection alone sanitizes hostile input is how the mitigation becomes the vulnerability.
How content platforms support select-don't-generate over your content
| Feature | Sanity | Contentful | WordPress | Strapi |
|---|---|---|---|---|
| Pass the exact field or block as state | Content Lake stores content as data addressed by field, block, and document, so you pass exactly the unit a question concerns inside a bounded state budget. | API-first structured content with defined content types and fields, so answers can be addressed at the field level rather than the page. | Content is largely page and post HTML with custom fields bolted on, so the addressable unit is usually a rendered page you parse to isolate a section. | Structured content model with components exposed over REST and GraphQL, so fields and components are addressable units you can pass directly. |
| Derive the option set from the content model | Enumerated fields, taxonomy references, content types, and workflow states are Choice sets that already exist in the schema, read off the model rather than hand-maintained. | Content types and enumerated fields define answer spaces; deriving option lists is possible, with modeling more UI-bound than schema-as-code. | Taxonomies and custom fields can define option sets, but the default page shape means many answers live inside post body HTML. | Components, enumerations, and relations in the content model give you option sets you can query and derive candidates from. |
| Build the candidate set with hybrid retrieval | One GROQ query filters with predicates, then blends text::query() and text::semanticSimilarity() in a score() pipeline ordered by _score. | Structured filtering is native; semantic ranking over content typically relies on an external vector service you wire in. | Keyword and taxonomy queries are native; semantic ranking generally means a plugin or an external embedding service. | Filtering via REST and GraphQL is native; semantic ranking becomes a separate vector database plus indexing pipeline you build. |
| Keep the candidate index fresh on change | Content Lake keeps the index fresh, so a description edit, a price change, a publish, or a delete is reflected without a separate re-embedding project. | Content updates fire webhooks; keeping an external semantic index in sync is your integration to build and maintain. | Content changes fire hooks; re-indexing an external semantic store on every edit and delete is a pipeline you own. | Lifecycle hooks exist; incremental re-embedding, deletion handling, and backfill on schema changes are a permanent roadmap line item. |
| Route low-confidence selections to review | Functions react to create, update, and delete document events to open a review task; Workflows (beta) model the review process as versioned data. | App Framework and workflow features can route items for review; the low-confidence gate itself is application logic you add. | Editorial roles and plugins support review queues; wiring confidence-gated fallthrough is custom application code. | Roles, draft states, and lifecycle hooks support review routing you implement in your application layer. |
| Audit what was selected and shown | Store the chosen label and per-option probabilities alongside versioned content; with Workflows (beta) 'what published without review' is one GROQ query. | Version history and entries are queryable; storing decision metadata next to content and auditing it is a schema and query you design. | Revision history exists per post; assembling a decision trail across selections is largely custom logging you maintain. | Draft and publish states plus custom fields let you store decision metadata; the audit query is one you build against your model. |