When the Taxonomy Changes: Handling Option-Set Drift in Production Decision Models
A recommendation engine tags every incoming article with one of eighteen content categories.
A recommendation engine tags every incoming article with one of eighteen content categories. Six weeks later, product marketing splits "Analytics" into "Product Analytics" and "Marketing Analytics," retires "General," and renames "How-to" to "Guides." Nobody tells the decision model. It keeps returning "General" for hundreds of documents, confidently, against an option that no longer means anything. The labels are schema-valid and completely wrong, and because the model emits no rationale, nobody notices until a downstream dashboard shows a category that the taxonomy team deleted a month ago.
This is option-set drift, and it is the failure mode nobody planned for when they wired a fast typed-decision model like TypeSafe's Jev into a production content pipeline. A Choice question picks one option from a fixed set, but taxonomies are not fixed. They split, merge, and get renamed while the model keeps answering against a snapshot frozen in code.
This article reframes drift as a content-modelling problem, not an AI problem. The fix is to stop hard-coding option sets and start deriving them from the taxonomy's source of truth. That is where Sanity, the Content Operating System for the AI era, earns its place on this page: when your categories are queryable documents rather than a string array, the option list, the labelled history, and the migration all become things you can query, re-run, and audit.
What is option-set drift, and why does a decision model make it worse?
Option-set drift is what happens when the fixed set of choices a typed-decision model picks from stops matching the taxonomy the business actually uses. There are two distinct kinds, and conflating them is the first mistake.
The first kind is the option set changing. Someone adds a category, merges two tags, renames a product line, or retires a label. TypeSafe's Jev, released on 15 September 2026, answers a Choice question by picking one option from a set you define in code, up to 255 options, and returning a probability for every option plus a confidence value. If that set is a hard-coded array and the taxonomy moves, the model is now answering yesterday's question. It cannot pick the new category because the new category is not in its option list, and it will not tell you it is confused. It will pick the closest surviving option and hand you a schema-valid, business-wrong label.
The second kind is subtler: the content changes under a fixed option set. The categories stay the same, but new topics arrive that none of the old labels fit. The model is still forced to choose, so it rounds every genuinely novel document to the nearest existing bucket. The taxonomy looks stable while quietly losing resolution.
What makes a decision model worse than a human here is that it fails silently and at volume. TypeSafe reports 70 to 500ms end-to-end, so a drifted option set produces thousands of confidently mislabelled documents before anyone reviews a single one. And because Jev writes no prose, no rationale, and no summary, there is no sentence in the output that says "none of these fit well." The guarantee that it cannot emit an off-schema value is a guarantee about format, not correctness. Drift lives entirely in the gap between the two.
Why should you derive option sets from the taxonomy instead of hard-coding them?
Deriving option sets from the taxonomy's source of truth is the single change that turns drift from a silent outage into a controlled migration. When the option list is generated from the same place the business edits its categories, the model and the taxonomy cannot fall out of sync by accident, because there is only one list.
Hard-coding is seductive because it is fast. You write the Choice options as a literal array next to the API call, ship it, and move on. The cost arrives later, at the exact moment the taxonomy changes, and it arrives as a distributed bug: the array in the decision service, a dropdown in the editor, a filter in the frontend, and a report in the warehouse all drift apart because each holds its own copy of the truth. Reconciling them is archaeology.
The alternative is to treat taxonomy terms as data with an identity, not as strings. When each category, tag, or workflow state is a document with a stable ID, a human-readable title, and a status, you generate the Choice option set by querying that collection at call time or at deploy time. Adding a category is a content edit, not a code change. Retiring one is a status flip, not a search-and-replace across four repositories.
This is where structured content stops being hygiene and becomes leverage. A schema already defines answer spaces: enumerated fields, references to taxonomy documents, content types, and workflow states are Choice sets that exist before anyone writes a prompt. In Sanity, taxonomy terms modelled as documents are queryable with GROQ, so the list you feed Jev is the same list your editors see in the Studio and the same list your frontend filters on. Best practice with Jev is to include an explicit "other" option so the model can decline rather than force a wrong pick; derive the real options from the taxonomy and append "other" as the safety valve, and you have a set that stays honest as the business moves.

How do you detect drift before it corrupts the dataset?
You detect drift by watching two numbers the model already gives you for free: the rate of "other" answers and the rate of low-confidence answers. Together they are your drift alarm, and they fire before a human ever notices a bad label in a report.
Every Jev Choice response returns a probability for every option plus a confidence value, and every Noul response returns a probability from 0 to 1 that doubles as its own certainty. If you include an explicit "other" option and log every response, a rising share of "other" picks means the content is arriving that your current labels do not fit. That is the second kind of drift, content moving under a fixed option set, showing up as a measurable trend rather than a vague sense that categorisation has gotten worse.
Low confidence is the companion signal. When the model spreads its probability mass thinly across several options instead of concentrating it, the state it was handed does not map cleanly onto the choices. A cluster of low-confidence answers on documents that used to score high is often the first visible symptom that a category was split or renamed upstream, because the split leaves the old option straddling two new meanings.
The operational move is to make both rates first-class metrics, not log lines nobody reads. Emit them per question and per taxonomy version, chart them over time, and alarm on a step change. Because Jev evaluates questions independently against one shared read of the state, and a tenth question costs tokens but almost no extra time, you can afford to run a small panel of diagnostic questions alongside the real one without meaningfully slowing the pipeline.
One discipline matters more than any threshold: never convert a timeout or a rate limit into a high-confidence default. A dropped call is not a decision. If you silently backfill failures with the last known label, you will mask exactly the drift these metrics exist to catch.
How do you migrate a taxonomy without silently corrupting old labels?
You migrate by adding before you remove, running both versions side by side on a sample, reviewing where they disagree, and only then switching the pipeline over. Ripping out an old option and dropping in a new one in a single deploy is how you get a dataset full of labels that were correct under one taxonomy and meaningless under the next.
Here is a concrete playbook for an option-set change. First, add the new option to the derived set while keeping the old ones and the explicit "other" fallback in place, so nothing the model can pick disappears mid-flight. Second, run both the old question definition and the new one over a representative sample of documents, capturing the full probability distribution from each, not just the winning label. Third, review the disagreements: the documents where the old version and the new version pick different options are exactly the population your migration affects, and their volume tells you how big a relabelling job you have signed up for. Fourth, once the disagreements look right to a human reviewer, switch the production pipeline to the new definition and retire the old option.
The reason this stays tractable is that everything in the loop is queryable when your taxonomy lives in the content model. When categories are documents referenced by content, the affected population is not a guess; it is a query for every document pointing at the retired or split term. A schema migration can then re-run the decision on exactly those documents rather than reprocessing the whole corpus. In Sanity, document history and Content Releases let you stage the relabelling as a reviewable batch, so a taxonomy change lands as a controlled release with an audit trail rather than an in-place overwrite that erases what the labels used to mean.
How do you audit a decision that comes with no reason?
You audit a reasonless decision by recording everything around it, because the model itself will never explain a pick. Jev returns a winner, a probability for every option, and a confidence value, and that is all. It writes no rationale. For a regulated or high-stakes taxonomy decision, "why this category?" has no answer from the model, so the answerable version of that question has to be assembled from telemetry you capture yourself.
The minimum record for every decision is four things: the exact state you handed the model, the full question definition including the option set at that moment, the complete response including the probability distribution and confidence, and the model version that produced it. Pin the versioned model ID rather than trusting a floating alias, log the model reported in each response, and version your question definitions alongside application code so you can reconstruct not just what was decided but what could have been decided. Jev publishes jev-1.13.0 as its only version today, but pinning is a discipline for the day there are more.
The state is the part builders forget, and it is the part that matters most for content. If you handed the model a wall of HTML, your audit trail records a wall of HTML, and reconstructing which field actually drove the pick is guesswork. If you handed it a specific block, field, or section from structured content, the record shows precisely what the model saw. Storing the taxonomy version alongside the decision closes the loop: when someone asks why a document carries a category that no longer exists, you can answer with the option set that was live when the call was made. This is content modelling doing the work an opaque model cannot, and it is why the platform holding the state, not just the model making the call, is where auditability is won or lost.
Where does confidence gating fit when the taxonomy is in flux?
Confidence gating is what keeps a drifting taxonomy from auto-labelling its way into a mess: the answer says what to do, and the confidence decides whether it is safe to do automatically. During a migration, when disagreement and "other" rates are elevated, that gate is the difference between a controlled rollout and a silent corruption event.
The pattern is an escalation cascade. Run the fast decision model at volume, and route only the uncertain minority to a frontier model or a human reviewer. Thresholds are not a single global constant; they rise with the cost and irreversibility of a mistake. Auto-applying a blog category that is trivial to fix can sit at a low bar. Auto-applying a compliance classification, a content-visibility state, or anything that gates publication should demand much higher confidence and send everything below it to review.
Crucially, thresholds have to be earned against reviewed labels before you trust them. Pick a threshold, apply it to a sample where a human has already assigned the correct label, and measure how often the auto-applied decisions were right. Only then do you know what confidence level actually corresponds to acceptable error for your taxonomy and your risk tolerance. A threshold guessed at is a threshold that will fail exactly when the content shifts.
This is also the honest boundary of what a decision model buys you on hostile input. Prompt injection still works when attacker-controlled text sits in the state, and injected text can flip a picked option from one value to another. A typed decision model shrinks what an attacker can make the system do to your fixed option set; it does not make untrusted state safe. Confidence gating and human review on the low-confidence tail are part of how you contain that, not a substitute for treating attacker-controlled text as attacker-controlled. In Sanity, Functions can run the decision on publish and route low-confidence results into a review workflow rather than straight to the live label, so the gate is enforced in the pipeline, not just documented in a wiki.