Content Operations7 min read•

Auto-Tagging and Classifying Content at Scale with a Decision Model

A content lead inherits a library of four thousand articles tagged by twelve different people over six years.

A content lead inherits a library of four thousand articles tagged by twelve different people over six years. "Product" means three different things depending on who filed it, half the back catalogue predates the current taxonomy, and the search facets are quietly useless because the underlying tags are inconsistent. Retagging by hand is a quarter's worth of someone's time, so it never happens, and the rot compounds with every publish.

Auto-tagging with a decision model reframes that project as a batch job. Instead of a person reading every item, you enumerate your taxonomy as fixed options, hand a typed model the exact content each question is about, and get back a label plus a confidence value for every document. TypeSafe reports that one project classified 1,018 research papers across 24 candidate topics for $0.08 total, at a median 256ms per paper. A back catalogue of a few thousand articles can be retagged for the price of a coffee, and the real cost moves to designing the taxonomy.

The judgement is only as good as the state you feed it, which makes this a content-modeling problem before it is an AI problem. This is where Sanity, the AI-native content platform, earns its place in the workflow: because its content is structured and queryable, you can hand a typed model the exact field, block, or section a question is about, inside a tight state budget, rather than a wall of markup.

What is a decision model, and why use one for tagging?

A decision model is a class of model that makes fast, structured decisions instead of generating text. TypeSafe released the first of these, Jev, on 15 September 2026, calling it a 'System One model' after Daniel Kahneman's fast, intuitive mode of thinking. You pass it state, meaning any text or JSON you want judged, plus typed questions you define in code. It answers them and stops. It writes no prose, no rationale, no summary, and no code.

That narrowness is exactly what makes it right for tagging. A chat model asked to classify an article can drift into explaining itself, invent a category that is not in your taxonomy, or return prose you then have to parse. A decision model cannot do any of those things, because the valid answers are fixed by your schema before the call. For auto-tagging you mostly use one question type: Choice, which picks one option from a set you define, up to 255 options, and returns the winning option, a probability for every option, and a confidence value.

The economics are the argument. TypeSafe reports pricing of $0.042 per million input tokens, with output tokens unmetered and free because what comes back is a decision rather than prose. On the vendor's own benchmark it reports a per-decision cost of about $0.0004, against $0.0304 and $0.0836 for two comparable frontier chat models. For a content operation, that turns 'retagging the archive' from a headcount question into a line item you can approve without a meeting. Every one of these figures is self-reported by TypeSafe and has not been independently reproduced, so treat them as vendor claims when you plan against them.

Illustration for Auto-Tagging and Classifying Content at Scale with a Decision Model
Illustration for Auto-Tagging and Classifying Content at Scale with a Decision Model

How do you turn a taxonomy into a classification workflow?

Turning a taxonomy into a classification workflow means enumerating your controlled vocabulary as the options for a Choice question, then running one classification pass per document. The steps are concrete. First, list every valid tag in the category you are classifying, for example your twenty-four content topics, as the Choice options. Second, add an explicit 'other' option, which TypeSafe recommends as best practice so the model can say nothing fits rather than picking the closest wrong answer. Third, select the state to hand the model with a query that pulls the specific fields the question is about, not the whole document. Fourth, write the returned label and confidence back onto the document as fields.

Most content items need more than one tag, and here the architecture pays off. Questions in one request are evaluated independently and in parallel against a single shared read of the same state, so you can batch several questions per document, for example section, audience, funnel stage, and content type, in one call. Because they share that read, a tenth question costs tokens but almost no extra time. TypeSafe's cookbook reports that batching 13 questions into one call runs 12.2x cheaper and 10x faster than asking them separately, with identical answers.

In Sanity, this maps cleanly onto surfaces you already have. Your controlled vocabulary can live as a reference list or an enumerated field, which is the same list you enumerate as Choice options. A GROQ query selects exactly the fields a question needs as state, so you stay inside the 32,000-token state budget instead of shipping a wall of markup. The classification results write back into structured fields, ready to power facets, personalization, and downstream retrieval.

Why does structured content make or break the result?

Structured content makes or breaks the result because the quality of a typed judgement depends entirely on the quality of the state it is handed. A decision model does not fetch context or browse. It reads exactly what you pass, once, and answers. If the state is noisy, ambiguous, or padded with irrelevant markup, the label degrades, and no confidence threshold fully rescues a bad input.

This is where a content model that stores rich text as an opaque HTML blob actively hurts you. All you can pass is the whole blob, wrapped in tags, which burns the 32,000-token state budget on markup and dilutes the signal the question actually depends on. A platform that stores content as structured data lets you pass the exact field, block, or section the question is about. Classifying funnel stage? Pass the intro and the call to action, not the entire body. Classifying audience? Pass the summary field and the product references. The tighter and more relevant the state, the more reliable the Choice.

Sanity stores rich text as Portable Text, a structured format where annotations, marks, and blocks are addressable rather than flattened into a string. That means a query can pull a single named block or a specific section and hand it over as clean state, and it means the structure survives when you slice content for the model. Combined with GROQ, which lets you project precisely the shape you want, you get the practical thing a decision model needs most: the ability to feed it a small, high-signal window onto each document rather than everything at once. The content-modeling work you do here is not overhead. It is the accuracy lever.

What does 'cannot hallucinate' actually guarantee?

'Cannot hallucinate' guarantees format, not correctness, and getting that distinction right matters more than any single accuracy number. Because the valid answers are fixed by your schema before the call, returning an off-schema value or a type error is structurally impossible. The model will never invent a twenty-fifth category or write an essay instead of picking a tag. TypeSafe reports a 0% structured-output error rate on this basis, though the company is reasonably candid that the figure is asserted from schema design rather than measured empirically.

What it does not guarantee is the right valid answer. A decision model can still file a billing document under 'technical', confidently, because 'technical' is a legitimate option in your set. It will pick a valid tag every time; it will not always pick the correct one. Any workflow that treats a returned label as automatically true is misreading the guarantee. The format is safe. The judgement is still a judgement.

That is precisely why the confidence value is load-bearing and a low-confidence queue is not optional. On TypeSafe's own four-workflow benchmark Jev scores 67.8% accuracy, level with one frontier model and behind two others, at roughly one two-hundredth of the cost. Crucially, that benchmark measures agreement with two other large models used as consensus labels, not agreement with ground truth, so it tells you how often the decision model matches a committee of chat models, not how often it is right. For tagging your library, the number that matters is one you generate yourself: how the model performs against a set of reviewed labels from your own content.

How do you catch the wrong-but-valid tags before they ship?

You catch wrong-but-valid tags with confidence-gated routing: the answer says what to tag, and the confidence decides whether it is safe to apply automatically. Anything at or above your threshold writes straight to the field. Anything below goes to an editor review queue. This is the escalation cascade in miniature. The decision model classifies cheaply and at volume, code applies what clears the bar, and ambiguous cases route to a person or, if you prefer, to a frontier model for a second opinion.

Thresholds should rise with the cost and irreversibility of the mistake. A wrong internal facet tag is cheap to fix; a wrong compliance or regional tag that changes who sees a document is not, so it deserves a higher bar and human sign-off. Set these against reviewed labels covering easy, ambiguous, missing-context, and edge cases before you trust them, rather than picking a round number and hoping. And treat failure honestly: confidence only exists when a request succeeds, so never convert a timeout or a rate limit into a high-confidence default. Use a deterministic fallback or queue the case.

Two weaknesses are documented and worth designing around. A decision model is not a calculator, so any tag that depends on counting is unreliable and the error grows with the size of the thing counted. And dates are text to it, not ordered quantities, so 'which came first' or 'does this fall in a window' are unreliable; extract dates with a Choice over enumerated options and compare them in code. In Sanity, Content Releases give you a natural staging surface for the queue: classify into a draft or a release, let editors approve the low-confidence items, and publish when the batch is clean, so nothing wrong-but-valid reaches production unreviewed.

How do you keep tags current as content changes and taxonomies shift?

You keep tags current by re-judging on change rather than on a schedule, and by running the whole library through a batch pass whenever the taxonomy itself moves. Both cases are cheap enough to do properly. The first is an event problem: when a document is edited or published, that is your trigger to re-run its classification against the current question definitions, so tags never drift out of sync with the content they describe. The second is the migration case, and it is the strongest argument for this whole approach.

Retagging a library after a taxonomy change is normally a project. Someone renames three categories, splits a fourth, and now every historical document is mislabeled against the new scheme, so it either gets a quarter of manual cleanup or it silently rots. A decision model turns that into a batch job. You enumerate the new taxonomy as Choice options, run every document through, write back the new labels, and send everything below threshold to review. The leap of faith becomes a review queue. Given the reported economics, a few thousand documents reclassify for pocket change, and the meaningful cost is the taxonomy design, not the inference.

In Sanity, the change trigger is Functions, serverless content automation hooks that run on publish. A classify-on-publish Function can call the decision model with the document's relevant fields as state, batch the several questions you care about into one call, and write the labels and confidence back before the document goes live. Content Lake real-time subscriptions give you the same freshness signal for bulk backfills. Pin the versioned model ID rather than a moving alias, log the model reported in every response, and version your question definitions alongside your application code, so a tagging behaviour change is always a release you chose, never an alias that quietly moved under you.