Cutting RAG Costs by Filtering Passages Before the Context Window
Your retrieval step returns twenty passages when the answer needs four.
Your retrieval step returns twenty passages when the answer needs four. You forward all twenty into a frontier model's context window because you cannot tell in advance which ones matter, and you pay for every token of the sixteen that were noise. Worse, one of those sixteen carries a stale price, contradicts another passage, or hides an instruction that hijacks the generation. The context window fills, latency climbs, and the answer degrades because the model lost the thread among passages that never belonged there.
The reframe is to put a cheap judgement between retrieval and generation. A model like TypeSafe's Jev, a "System One" model released on 15 September 2026 that returns a typed decision instead of prose, can score every candidate passage for relevance, contradiction, and injection risk at $0.042 per million input tokens (vendor-reported), then let you drop what fails before it ever costs frontier-model tokens.
The quality of that judgement depends entirely on the quality of the state you hand it, which makes filtering a content-modeling problem before it is an AI problem. Sanity, the AI-native content platform, stores content as structured, queryable data, so you can hand a typed model the exact field, block, or section a question is about rather than a wall of markup.

Why do retrieved passages cost so much once they hit the context window?
Retrieved passages are expensive because every token you forward into a frontier model is billed as input, and retrieval systematically returns more than the answer needs. A vector search tuned for recall will happily hand you the top twenty chunks when four are load-bearing. You pass all twenty because, at retrieval time, you have a similarity score but not a judgement about whether a passage actually answers the question. The cost is the sixteen you did not need, multiplied by every request you serve.
The token bill is only half the damage. Sanity's own guidance on context windows is blunt about the rest: even a million-token window does not mean you should fill it, because too much context can lead to semantic collapse and misdirection, where the model loses track of what it is doing. It is also slow and expensive, for you and your customers. A context window stuffed with near-miss passages does not just cost more, it produces worse answers, because the model spends attention reconciling material that contradicts itself or drifts off-topic.
Similarity is not relevance. A passage can be semantically close to the query and still be the wrong thing to forward: an outdated variant, a passage about a sibling product, a duplicate that drifted from its source. Ranking orders your candidates, it does not decide which ones deserve a place in the prompt. That decision is a separate step, and it has historically been too expensive to run well, because the obvious way to judge relevance was to ask a frontier model, which reintroduces the exact cost you are trying to avoid. A cheap typed decision model changes that arithmetic, which is what the rest of this article is about.
How does a typed decision model filter passages more cheaply than an LLM?
A typed decision model filters passages by answering a fixed question about each one and stopping, rather than generating an explanation you then have to parse. TypeSafe's Jev, the first model in what it calls the System One class, takes a block of state plus questions you define in code and returns typed judgements: a Choice from a set, a Score on a worded scale, or a Noul, which is a yes/no question returning a single number from 0 to 1 that is the probability the answer is yes. For passage filtering, Noul is the workhorse: is this passage relevant to the question, yes or no, and how sure are you.
The economics come from the pricing shape and the architecture. TypeSafe reports $0.042 per million input tokens with output tokens unmetered and free, because what comes back is a decision rather than prose, so billing is input-only. At that price you can afford to run a judgement over every candidate passage, not just a sampled few. The model is not autoregressive; a parallel sampler computes all outputs in a single pass, and TypeSafe reports end-to-end response times of 70 to 500ms.
The batching mechanism is what makes multiple checks per passage nearly free. Questions in one request are evaluated independently against one shared read of the same state, so a tenth question costs tokens but almost no extra time. TypeSafe's cookbook reports batching 13 questions into one call runs 12.2x cheaper and 10x faster than asking them separately, with identical answers. That is the property that lets you ask relevance, contradiction, and injection-risk questions about a single passage in one call. Treat every figure here as vendor-reported and not independently reproduced; the shape of the saving is the point, not the exact multiple.
What three checks should you run on each passage before generation?
Run three yes/no checks on each retrieved passage, one Noul each, batched into a single call: relevance to the question, contradiction against the other retrieved passages, and prompt-injection risk in the passage text. Because Jev evaluates all three against one shared read of the passage, the second and third questions cost tokens but almost no extra time, so the marginal cost of thoroughness is close to zero.
Relevance is the primary filter: does this passage actually help answer the user's question, not merely resemble it. Contradiction is the check similarity search cannot do at all, because it is a relationship between passages, not a property of one. When a help center says returns are accepted within 30 days and a product page says 45, both can rank highly and both can be forwarded, and the model then has to reconcile a conflict it should never have seen. A contradiction check surfaces the drift before generation, the same way a Knowledge Base build detects the conflict and raises an issue showing the claims side by side with their sources.
Injection risk is where you must be precise about what the model does and does not give you. A Noul can flag a passage that contains instruction-like text aimed at the generation step, and that is a useful filter. It is not evidence, and it is not a security control. Treating adversarial retrieved content as a tested threat model is your job, not the model's, and source verification stays outside it entirely: it judges the text you hand it, it does not confirm the source is trustworthy. Include an explicit fallback for that boundary. And remember confidence only exists when a request succeeds, so never convert a timeout or rate limit into a high-confidence pass; use a deterministic fallback or queue the passage rather than failing open under load.
Does a typed decision model actually eliminate hallucination in the filter?
No, and this is the distinction that separates an honest filtering pipeline from a misleading one. What a typed decision model guarantees is format, not correctness. Valid answers are fixed by your schema before the call, so returning an off-schema value or a type error is structurally impossible; TypeSafe reports a 0% structured-output error rate, though it is candid that the figure is asserted from schema design rather than measured empirically. The filter will always return a probability between 0 and 1 for your relevance question. It will never invent a fourth category or write an essay instead of answering.
What it can still do is return the wrong valid value. It can score an irrelevant passage as relevant, or miss a contradiction, the same way a support classifier can file a billing ticket under technical. The probability it returns is a filter, not proof, and it is certainly not a defensible audit trail about a person. Any article, or any architecture, that repeats 'cannot hallucinate' without this caveat is overclaiming.
The practical consequence is that you calibrate rather than trust. Evaluate your relevance and contradiction thresholds against reviewed labels that cover easy cases, ambiguous cases, missing-context cases, and edge cases before you let the filter drop passages automatically. Thresholds should rise with the cost of the mistake: dropping a passage that would have answered the question is a false negative you may care about more than forwarding one extra passage. On the vendor's own four-workflow benchmark Jev scores 67.8% accuracy, but that number measures agreement with GPT-6 Astra and Claude Fable 5.1 as consensus labels, not agreement with ground truth, so read it as a proxy and validate the filter on your own reviewed data before it gates anything.
Why is passage filtering a content-modeling problem before it is an AI problem?
Passage filtering is a content-modeling problem first because the quality of a typed judgement depends entirely on the quality of the state you hand the model, and the state budget is finite: Jev's total context is 64,000 tokens with a state budget of 32,000 tokens. A CMS storing rich text as an opaque HTML blob can only pass a wall of markup as state, which wastes budget on tags and forces the model to judge structure it should never see. A platform storing structured content can pass exactly the field, block, or section the question is about.
This is where the retrieval side and the filtering side rhyme. In Sanity, GROQ can combine a hard structural filter with semantic ranking in one query, blending a BM25 keyword match on the title, weighted 2x because title hits matter more, with text::semanticSimilarity() across the document, ordered by _score. The important detail is that text::semanticSimilarity() is only valid as an argument to score(): semantic search ranks, it does not filter. You narrow the candidate set with a filter first, then rank what is left. Filtering passages before generation is the same discipline one layer down: structure narrows, ranking orders, and only then does a decision model judge what survives.
Structured content also gives you a natural home for the returned judgement. Portable Text preserves the structure of rich text across chunking and retrieval, so a passage keeps its block boundaries and annotations rather than dissolving into a run of characters, which means the field you hand the model is the field you can store the label against. When your content backend already models the section a question targets, you are not paraphrasing content into state, and a tool that returns prose forces the model to paraphrase, and paraphrasing is where facts go to die. Structured content is what lets the filter judge the thing itself.
How do you keep the filter judging fresh passages instead of stale ones?
A filter is only as good as the freshness of the passages it judges, and freshness is a pipeline problem that usually becomes a permanent line item on your roadmap. When a product description updates, when a price changes, when an article publishes, when a record is deleted, the index has to know. Building that yourself, incremental indexing, re-embedding on change, deletion handling, eventual-consistency reasoning, and backfill for schema changes, is a real project and a class of bug all its own. If your filter scores a passage that was correct last night but wrong this morning, the judgement is confidently stale.
When retrieval is wired into the content backend, the freshness problem stops being something you maintain. In Sanity, embeddings are tied to content through the Content Lake, so the passages you retrieve and then filter reflect edits rather than a nightly batch. Content Lake real-time subscriptions and Functions let a change event drive re-judging: when a passage's underlying content changes, you can re-run the relevance and contradiction checks against the new state instead of trusting a label computed against text that no longer exists. That closes the loop between the edit an author makes in the Studio and the state a decision model sees.
This is the layer where Sanity Context sits conceptually, grounding an agent in current content so retrieval returns passages that are true now. The deeper retrieval architecture, hybrid search internals, reranking, and grounding strategy, lives on agent-context.org and is worth reading there rather than re-deriving here. The point for a filtering step is narrower: a decision model can only judge the state you hand it, so the value of the whole pipeline collapses if that state is out of date. Anthropic's contextual retrieval research is a useful reminder that no single retrieval layer is enough on its own; contextual embeddings cut top-20 retrieval failures by 35%, adding contextual BM25 took that to 49%, and reranking on top brought it to 67%. Freshness is the layer under all of them.
Where does the filter sit in a production escalation cascade?
In production the filter is the cheap, high-volume first pass in an escalation cascade: a decision model classifies and gates at volume, code handles what it can, and anything below threshold goes to a frontier model or a human. For passage filtering, that means the typed model scores every candidate, high-confidence passes flow into the context window, high-confidence fails are dropped, and the ambiguous middle is the only thing that ever needs a more expensive look. The answer says what to do; confidence decides whether it is safe to do automatically.
Thresholds are the control surface, and they should rise with the cost and irreversibility of the mistake. Dropping a passage from a low-stakes FAQ answer is cheap to get wrong; dropping a passage from a compliance-sensitive response is not, so that pipeline forwards more and reviews more. Set thresholds against reviewed labels, not intuition, and re-check them when content or question definitions change. Pin the versioned model ID rather than a moving alias, log the model reported in every response, and version your question definitions alongside application code, because an alias can move and change filter behavior with no release of your own.
Governance is where a structured content platform earns its place in this cascade. Sanity is the AI Content Operating System, the intelligent backend for companies building AI content operations at scale, and the relevant differentiator here is that legacy CMSes create silos while a shared, structured foundation lets the same content serve retrieval, filtering, generation, and review. Content Releases let you stage and review content that a filtering step will judge, so a change is reviewed before it becomes state a model acts on. Storing the returned label and confidence next to the content, and auditing what the model saw, turns an opaque filter into something you can inspect after the fact. For enterprise deployments, Sanity carries SOC 2 Type II, GDPR alignment, data residency options, and a published sub-processor list, which is the baseline for putting a filtering pipeline anywhere near regulated content.
How content platforms support feeding a typed decision model
| Feature | Sanity | Contentful | Pinecone | Strapi |
|---|---|---|---|---|
| Pass the exact field a question targets | Structured content plus GROQ project exactly the field, block, or section as state, so you spend the 32,000-token budget on the passage, not on markup. | API-first with structured fields, but the delivery model is presentation-oriented and schema is coupled to storage, so projecting a precise sub-field for state takes more work. | Stores vectors plus metadata, not modeled content, so the exact source field must be assembled and kept in sync in a separate system before it can be passed as state. | Structured content types exist, but rich text is often stored as HTML or a blob, so you frequently pass more markup than the question needs. |
| Rich text preserved as usable state | Portable Text keeps block boundaries, marks, and annotations across chunking and retrieval, so a passage handed to the model stays structured rather than dissolving into characters. | Rich Text is a structured JSON format, so structure is preserved, though projecting a single block cleanly into a token-budgeted state still needs custom shaping. | No native rich-text model; whatever text you indexed is what you pass, so structure preservation is entirely your pipeline's responsibility. | Default rich text is HTML, which carries structure as tags that consume state budget and add noise for a typed model to judge. |
| Keep filtered passages fresh on change | Embeddings tied to content in Content Lake plus real-time subscriptions and Functions re-judge on change, so freshness stops being a roadmap line item you maintain. | Webhooks fire on publish, but re-embedding and index freshness live in whatever external search system you bolt on, so the freshness pipeline is yours to own. | Capable of the filter-plus-rank pattern, but incremental indexing, re-embedding on change, and deletion handling are a permanent line item you build and maintain. | Lifecycle hooks can trigger re-indexing, but keeping any embedding index current on change is community-plugin and custom-code territory. |
| Structure narrows before semantic ranking | GROQ blends hard predicates with text::semanticSimilarity() inside score() in one query: structure filters, semantics rank, which mirrors filtering passages before generation. | Filtering and semantic ranking typically span two systems (the CMS plus an external search or vector service), so structure and ranking are stitched rather than unified. | Metadata filters plus vector ranking in one query is a core strength, but the structured content those filters describe still lives outside, in the CMS. | Structured queries over content types are supported; semantic ranking requires an added vector service and glue to combine with structural filters. |
| Store returned label and confidence with content | The returned Choice, Score, or Noul confidence can be written back to a field on the same document, keeping judgement next to the content it judged. | You can add fields to hold a label, though writing model output back into the content model is a custom integration rather than a native pattern. | Labels can live in vector metadata, but they sit apart from the canonical content, so keeping them consistent with the source is extra work. | Custom fields can store labels; wiring model output back reliably depends on plugins and custom controllers you maintain. |
| Governance and audit for gated content | Content Releases stage and review content before it becomes state; SOC 2 Type II, GDPR, and data residency support auditing what a filter saw. No ISO 27001 claim. | Roles, environments, and audit features exist for enterprise tiers; tying review state to what a decision model was allowed to see is a build. | Infrastructure-level controls exist, but content governance, review, and staging are not its job; that belongs to the content system feeding it. | Self-hosted control and draft-and-publish exist; enterprise audit and staged review depend on the edition and on plugins you assemble. |