Content Modelling8 min read•

Why Structured Content Beats HTML Blobs When Models Make Decisions

A support team wires a decision model into their triage queue.

A support team wires a decision model into their triage queue. It classifies tickets into six categories, and it works, until they realize the "state" they are handing it is the entire rendered ticket page: the HTML wrapper, the nav chrome, three help-center promo blocks, and somewhere in the middle, the customer's actual complaint. The model has 32,000 tokens of state budget, and they are spending most of it on markup. Accuracy drops on long threads, cost climbs, and nobody can say which part of the page drove any given decision.

Here is the thing that machine exposes: the problem was never the model. It was the content model. A generative chatbot handed the same messy dump usually copes and hides the mess. A model that only makes typed judgements, TypeSafe's Jev, released 15 September 2026, cannot cope its way around bad structure, because there is no room to. This article argues that a decision model puts a price on content modelling that generative AI let teams avoid, and that the fix belongs to Sanity, the AI-native content platform whose structured, queryable content lets you hand a typed model the exact field a question is about.

Illustration for Why Structured Content Beats HTML Blobs When Models Make Decisions
Illustration for Why Structured Content Beats HTML Blobs When Models Make Decisions

What is a decision model, and why does it care about your content model?

A decision model is a model that reads state and returns typed judgements instead of writing text. TypeSafe's Jev, released on 15 September 2026 as the first "System One model", is the working example: you pass it state (any text or JSON you want judged) plus questions you define in code, it answers them, and it stops. It produces no prose, no rationale, and no summary. The name nods to Kahneman's fast, intuitive System 1, against the deliberative System 2 that chat models imitate.

You ask it three kinds of question. A Choice picks one option from a set of up to 255, returning the winner, a probability for every option, and a confidence value. A Score places the state on a worded scale of 2 to 10 levels and can land between them, so it might return 1.035. A Noul is a yes/no that comes back as a single number from 0 to 1, where the probability is itself the certainty measure. Built by Diogo Almeida, a co-author of the InstructGPT paper, Jev runs on a parallel sampler rather than autoregressive token generation, so all outputs compute in one pass.

The reason this belongs on a content site is simple: the quality of a typed judgement is bounded by the quality of the state you hand it. A generative model given a wall of markup will improvise around the noise. A decision model has nowhere to hide, because it does not narrate and it does not fill space. So the question "what exactly do we send?" moves from optional to mandatory, and that question is a content-modelling question wearing an AI costume.

Why does a 32,000-token state budget expose bad content modelling?

The 32,000-token state budget is where the cost of bad modelling becomes visible. Jev's total context is 64,000 tokens, with 32,000 reserved for the state you are judging. That is a hard ceiling, and every token spent on HTML tags, navigation chrome, cookie banners, or footer boilerplate is a token not spent on the content the question is about. A generative model handed the same bloated dump usually copes, which is exactly the problem: coping hides the waste.

Consider a product page passed as a raw rendered HTML blob. A large fraction of those tokens describe layout, scripts, and repeated site furniture. If your question is "is the return policy clearly stated?", you are paying to send the mega-menu, the related-products carousel, and the newsletter modal so the model can find two paragraphs buried inside them. On a long page you can blow the budget before you reach the content that matters, and truncation silently drops the part you needed.

The Sanity Context team's production data makes the general point sharply: even a million-token window does not mean you should fill it, because too much context leads to semantic collapse and misdirection where the model loses track of what it is doing, and it is slow and expensive for you and your customers. A tight budget is not a constraint to route around. It is a forcing function that rewards a content model where you can address and extract the exact field, block, or section a question concerns, and send only that.

HTML blob vs markdown vs structured rich text: three ways to store the same article

Take one article stored three ways, and watch what each representation lets you do when a decision model asks about it. As an HTML blob, the content is a single opaque string. You cannot address "the pricing section" without parsing the markup, and parsing HTML reliably at scale is its own tax: counter-intuitive nesting, inline styles, and wrapper divs that mean nothing. You either send the whole blob and burn the budget, or you write brittle scrapers that break the next time the template changes.

Stored as a markdown field, you are better off. It is human-readable and lighter than HTML, but it is still a string. To pull out one section you split on heading markers heuristically, and heuristics fail on the edge cases: a code block containing a hash, an inconsistent heading level, a table that spans what you assumed was a boundary. You can make it work, but every consumer of that content re-implements the same fragile splitting logic, and none of them agree.

Stored as structured rich text, the article is a sequence of typed blocks with marks and annotations, and structure survives selection. A heading is a heading because it is typed as one, not because it starts with a hash. A section is addressable because it is data, not because you guessed where it began. This is the difference between hoping you can extract the pricing section and querying for it. When the consumer is a decision model on a fixed state budget, that difference is the whole game: one representation lets you pass exactly the section the question is about, and the other two make you pay for the rest of the page to reach it.

Why does section-level addressability make a typed decision cheaper and more accurate?

Section-level addressability is what turns a page-level question into a section-level one, and section-level questions are both cheaper and more accurate. Cheaper because you send fewer tokens: TypeSafe bills input only, at $0.042 per million input tokens, so trimming a 20,000-token page to the 800-token section the question concerns is a direct cost cut on every call. More accurate because the model is not asked to first locate the relevant part and then judge it. You have already located it. The judgement is all that is left.

This is where a query language earns its keep. In Sanity, content is stored as Portable Text, a structured rich-text format where blocks, marks, and annotations are addressable data, and GROQ can project exactly the part you want. A projection that selects the pricing section of a page, rather than the page, is the artefact that makes the whole pattern work. As the Sanity Context team puts it, pure structured query means you write the predicate and you get exactly what you asked for. For a decision model that is not a limitation, it is the entire point: you know what you are judging, so you can hand over precisely that.

The economics compound when you batch. Because Jev evaluates questions independently against one shared read of the state, a tenth question costs tokens but almost no extra time. TypeSafe's cookbook reports batching 13 questions into one call ran 12.2x cheaper and 10x faster than asking them separately, with identical answers. Feed one clean, well-scoped section as state, ask every question you have about it at once, and the shared read does the rest. The prerequisite for all of it is content you can address.

What does "cannot hallucinate" actually guarantee, and what does it not?

"Cannot hallucinate" is true about format and false about correctness, and conflating the two is where these systems go wrong. Because the valid answers are fixed by your schema before the call, returning an off-schema value or a type error is structurally impossible. TypeSafe reports a 0% structured-output error rate and a 0% tool-call error rate. Jev will never invent a fourth category when you defined three, and it will never write an essay when you asked for a Noul. That is a real guarantee, and it is worth having.

What it does not guarantee is that the valid answer is the right one. A billing ticket can still come back filed under technical. The category is valid, the schema is satisfied, and the decision is wrong. TypeSafe is reasonably candid that the 0% figure is asserted from schema design rather than measured empirically, and any write-up that repeats "cannot hallucinate" without drawing that line is misleading. Format safety removes one class of failure. It does not remove judgement error, and confidence-gated routing exists precisely because judgement error remains.

The content-modelling consequence is direct. If format is guaranteed but correctness is not, then correctness is the thing you have to defend, and correctness depends on the state. Send a section stripped of the context that disambiguates it, and you have manufactured a wrong-but-valid answer that no schema will catch. The state you hand over is the part still under your control, which is another way of saying the content model is the part still under your control.

How do you keep the state fresh when the content changes?

Fresh state means re-judging content the moment it changes, not on a nightly batch, and that is a pipeline problem your content platform either solves or hands to you. A decision that was correct against last week's pricing section is wrong the instant the price changes, and if your state is a cached HTML scrape, nothing tells you it went stale. The judgement quietly rots while the number that produced it sits in a database looking authoritative.

The fix is a content pipeline that fires on change. In Sanity, Content Lake keeps the search index fresh as content moves: when a product description updates, when a price changes, when an article publishes, when a record is deleted, the index has to know, and that is handled for you rather than being your cron job to maintain. Real-time subscriptions give you a change event to react to, and Functions let you run serverless automation on publish, so re-judging a changed section becomes a hook rather than a rebuild. The state a decision model reads can track the content it came from.

There is a governance corollary the brief on Jev makes plainly: confidence only exists when a request succeeds, so never turn a timeout or a rate limit into a high-confidence default. Queue the case or fall back deterministically. The same discipline applies to freshness. A stale-but-confident label is as dangerous as a timeout coerced into certainty. Staging and reviewing content changes through Content Releases before they go live gives you a place to re-run judgements against the new state, so a label and the content it describes move together instead of drifting apart.

Where does a decision model belong in a real content workflow?

A decision model belongs in the classify-cheaply-at-volume slot, with generative models and humans handling everything else, and the clearest early evidence is a workflow that used both. On 1kpapers.com, Hassan El Mghari classified 1,018 research papers across 24 candidate topics for $0.08 total, at a median 256ms per paper, having first paid $3.99 to a generative model to write the summaries. Different models for different parts of one job: prose from the generator, the typed routing decision from Jev. Neither replaces the other.

That split maps cleanly onto content operations. Where you need words, generation writes candidates; AI Assist can draft, rewrite a block in a different voice, or translate a page's headings into several locales inside the Studio. Where you need a judgement, a decision model routes, scores, or gates. Where content needs to become clean state for either, GROQ projects the exact block and Portable Text keeps it structured across chunking and retrieval. Each tool does the part it is actually good at, and the content model is the shared surface they all read from.

Here is the honest boundary to hold: Sanity, the AI Content Operating System and intelligent backend for companies building AI content operations at scale, has no announced integration, partnership, or shipped support for System One models, and this argument does not need one. The claim is architectural, not a product feature. A platform that stores content as addressable, queryable structure is what lets you hand any typed model the exact field a question is about. The decision model is one more consumer. The content model is still the protagonist, which is the whole point of caring about how you store an article before you care about who reads it.

How content platforms support feeding clean state to a typed decision model

FeatureSanityContentfulWordPressStrapi
Passing the exact field a question is aboutGROQ projections select a single field, block, or section as state, so a page-level question becomes a section-level one you can query for directly.Structured content model with typed fields; the Content Delivery API returns named fields, so field-level state is straightforward once modeled.Post body lives in one post_content field by default, so addressing a section means parsing the stored markup rather than querying for it.Structured fields exposed over REST and GraphQL, so selecting a specific field as state is well supported by the API.
Rich text as clean, structured statePortable Text stores rich text as typed blocks with marks and annotations, so structure survives chunking and selection instead of collapsing to a string.Rich Text field stores as structured JSON nodes, so rich text can be passed as structure rather than as raw HTML.Block editor stores serialized block markup inside an HTML blob; it is parseable, but sections are recovered from markup, not read as data.Rich text is a Markdown string or a blocks field; the blocks field is structured, while Markdown remains a string you split heuristically.
Fitting a 32,000-token state budgetProject only the section a question concerns, dropping nav, chrome, and boilerplate before the call, so budget goes to content not markup.Query individual fields to avoid sending whole entries, though rich-text JSON nodes carry some structural overhead into the budget.Sending rendered pages spends budget on layout and site furniture unless you build extraction to isolate the content first.Field-level API responses keep payloads lean; Markdown rich text still arrives as one string to trim yourself.
Re-judging state when content changesContent Lake keeps the index fresh on every change, and real-time subscriptions plus Functions give a publish event to re-judge against automatically.Webhooks fire on publish, so you can trigger re-judgement, though keeping any external index fresh is your pipeline to build.save_post hooks and REST expose change events; freshness of any downstream state is left to plugins or custom code.Lifecycle hooks and webhooks fire on entry changes, so re-judgement can be triggered with pipeline code you maintain.
Storing the returned label and confidenceModel a decision object (label, per-option probabilities, confidence) as typed fields on the document, queryable alongside the content it judged.Add fields to the content type to store the label and confidence, kept next to the entry they describe.Store results in post meta or a custom table; queryable, though not modeled as first-class typed fields by default.Extend the content type with fields for label and confidence, exposed over the same REST and GraphQL APIs.
Section-level projection of rich text as stateNative: a single GROQ query projects one Portable Text section as clean, structured state, no external parsing or extraction step.Read the Rich Text JSON and walk its nodes to isolate a section; supported, but section selection is application code, not a query.Requires parsing serialized block markup to isolate a section, so section-level state depends on custom extraction logic.Blocks field can be walked to a section in code; Markdown fields need heuristic splitting to reach one section.
Auditing what the model sawThe projected state is a deterministic GROQ result you can log and reproduce, and Content Releases stage changes so you know which version was judged.Versioning and webhooks let you reconstruct an entry's state at a point in time with supporting logging you add.Revisions track post changes, but reproducing the exact extracted state a model saw depends on your extraction being logged.Draft and publish states plus your own logging let you reconstruct what was sent, built with pipeline code.