Moderate on Publish: Using a Decision Model as an Editorial Guardrail
A generative pipeline can draft a hundred product descriptions before lunch, and your two-person review team can read maybe twenty.
A generative pipeline can draft a hundred product descriptions before lunch, and your two-person review team can read maybe twenty. The gap is where the trouble lives: an unsupported health claim, a missing affiliate disclosure, a tone that reads fine in isolation but wrong for a regulated market. When AI increases content volume faster than review capacity, the failure mode is not bad content, it is unreviewed content going live because nobody had time to look.
The usual answers both fail. Review everything and you throttle the pipeline back to human speed, erasing the reason you automated. Review nothing and you ship the health claim. What you actually want is a cheap, fast check that sits on the publish hook, judges each draft against a fixed set of rules, and decides which items are safe to publish, which need a warning, which go to a human, and which get blocked outright.
This is what a System One decision model like TypeSafe's Jev is built for: it consumes state and returns typed judgements, nothing else. But its judgement is only as good as the state you hand it, which makes an editorial guardrail a content-modelling problem before it is an AI problem. Sanity, the AI-native content platform, matters here precisely because its structured, queryable content lets you pass a typed model the exact field a question is about, rather than a wall of markup.

What is moderate-on-publish, and why put a decision model in the loop?
Moderate-on-publish is an automated check that runs at the moment content transitions from draft to live, judging each item against a fixed set of editorial rules and deciding what happens next. It is the single highest-leverage place to put a guardrail, because publish is the last point where a mistake is still cheap to catch and the first point where volume outpaces the people watching.
A System One model is the right tool for this specific job. Released by TypeSafe on 15 September 2026, Jev is what the vendor calls a fast, structured decision model: you hand it 'state', meaning any text or JSON you want judged, plus typed questions you define in code, and it answers them and stops. It writes no prose, no rationale, and no summary. For a guardrail that is a feature, not a limitation. You do not want an essay about whether a press release is compliant; you want a decision you can act on programmatically.
The three question types map cleanly onto moderation work. A Choice picks one option from a set you define, up to 255 options, which suits routing a draft to a category or naming which policy it violates. A Score places the state on an ordered scale of two to ten described levels, which suits severity and risk bands. A Noul is a yes/no question returning a single number from zero to one, the probability the answer is yes, which suits narrow checks like 'does this contain an unsupported medical claim'. Each returns probabilities, and Choice and Score also return a confidence value. That confidence is what turns a label into a governance decision, and it is the substance of everything that follows.
How do you batch brand-safety, claims, disclosure, and tone into one call?
You batch moderation checks by asking many narrow typed questions against one shared read of the same state in a single request. This is the mechanical reason a System One model fits moderate-on-publish so well. In Jev, questions in one request are evaluated independently and in parallel against one shared read of the state, so a question does not condition on another question's answer, and a tenth question costs tokens but almost no extra time.
That changes the economics of thoroughness. A blog post might warrant a dozen distinct checks: a Noul for each of brand-safety, unsupported claims, missing disclosures, and off-tone phrasing, a Score for overall risk severity, and a Choice for which market's rules apply. Asking those separately would be a dozen round trips. TypeSafe's cookbook reports batching 13 questions into one call runs 12.2 times cheaper and 10 times faster than asking them separately, with identical answers. Treat that as a vendor-reported figure, not an independently reproduced fact, but the architecture behind it is sound: one read, many parallel judgements.
The practical upshot is that you stop rationing checks. When each additional question is nearly free, you can afford a dedicated, well-scoped Noul for every specific harm your legal and brand teams care about, rather than one vague 'is this okay' question that collapses four different risks into a single muddy signal. Narrow questions also produce cleaner confidence values, because 'does this omit a required disclosure' is a far more answerable question than 'is this compliant'. The specificity that governance demands is exactly the specificity this model rewards.
The four-way outcome: allow, warn, send for review, or block
The outcome a guardrail acts on is not the label alone, it is the label combined with confidence, and mature teams resolve it into four actions: allow, warn, send for review, and block. The answer says what the content is; the confidence decides whether it is safe to act on automatically. This is confidence-gated routing, and it is where a decision model stops being a classifier and becomes a governance mechanism.
Work an example. A Noul returns 0.02 for 'contains an unsupported medical claim' with the model succeeding cleanly: allow, publish it. It returns 0.97: block, this almost certainly violates policy. It returns 0.55: this is genuinely ambiguous, so send it to a human review queue rather than guessing. A separate lower-stakes check, say a tone Score that lands slightly off-brand at moderate confidence, might warn the editor and publish anyway, flagging the item for a later pass rather than holding it.
The crucial design rule is that thresholds are tuned per harm category, not set globally. A wrong auto-block on a press release is an annoyance an editor clears in seconds. A wrong auto-approve on a regulated financial or health claim is a compliance incident. So the threshold that sends a regulated-claim check to review should sit far more conservatively than the one governing a stylistic nit. Thresholds rise with the cost and irreversibility of the mistake. You are not looking for one number that works everywhere; you are encoding your organization's actual risk tolerance, harm by harm, into the routing logic that sits between publish and live.
Why the state you pass decides the verdict, and where structured content wins
The quality of a typed judgement depends entirely on the quality of the state the model reads, which makes an editorial guardrail a content-modelling problem before it is an AI problem. Jev has a total context of 64,000 tokens and a state budget of 32,000 tokens. What you fit into that budget, and how cleanly, determines whether the check is accurate or noise.
Here structured content is the whole advantage. A platform that stores rich text as an opaque HTML blob can only hand the model a wall of markup, burning the state budget on tags and forcing every question to answer against everything at once. A platform that stores content as structured data can pass exactly the field, block, or section a given question is about. This is where Sanity's architecture earns its place: Portable Text represents rich text as structured blocks with typed annotations and marks, so you can extract the body copy a claims check needs without the layout noise, and preserve the structure that tells the model a heading from a disclaimer. GROQ lets you query for precisely the fields each question targets, so a disclosure check reads the disclosure region and the legal footer, not the entire document.
The payoff is both accuracy and cost. Passing the exact relevant field means the model judges the thing you asked about, not a distant paragraph that happened to share the blob. Since Jev bills input-only at $0.042 per million input tokens on TypeSafe's published pricing, with output unmetered, a leaner state is also a cheaper check. Structured content is what lets you stay well inside 32,000 tokens while giving each question the precise slice it needs to answer well.
Failing closed: why a timeout must never become a high-confidence allow
The quiet way guardrails break is by treating an infrastructure failure as a clean verdict. Confidence only exists when a request succeeds. If the moderation call times out, hits a rate limit, or errors, there is no confidence value to route on, and the single most dangerous mistake you can make is to let that absence default to allow. A guardrail that fails open is not a guardrail.
The rule is to fail closed. When a check cannot complete, the content does not go live; it goes to the review queue. A timeout is not a low-risk item, it is an unjudged item, and unjudged content is precisely what moderate-on-publish exists to prevent from shipping unattended. Never convert a timeout or rate limit into a high-confidence default of any kind. Use a deterministic fallback or queue the case, and make sure the queue is monitored so a provider outage produces a visible backlog rather than a silent flood of unchecked live content.
Governance discipline extends to the model itself. Pin the versioned model ID, currently jev-1.13.0, rather than a moving alias like jev-latest, because an alias can move and change behaviour with no release on your side. Log the model version reported in every response, and version your question definitions alongside your application code, so a change to what 'unsupported claim' means is a reviewable commit and not an invisible drift. And evaluate every threshold against reviewed labels that cover easy, ambiguous, missing-context, and edge cases before you trust it in production. A threshold you have not tested against real ambiguous cases is a guess wearing a number.
The honest limits: format guarantees, images, and audit trails
Be precise about what this class of model does and does not protect you from, because the marketing shorthand 'cannot hallucinate' is misleading without a distinction. What is actually guaranteed is format, not correctness. Because the valid answers are fixed by your schema before the call, returning an off-schema value or a type error is structurally impossible; TypeSafe reports a 0% structured-output error rate, though it is candid that this is asserted from schema design rather than measured empirically. The model will never invent a fourth category or write an essay instead of deciding. It can absolutely still return the wrong valid value: a policy-violating draft scored as clean. Your thresholds and human review exist because correctness is not guaranteed, only structure is.
Two limits matter for moderation specifically. First, Jev is text-only, so it cannot moderate the image in your hero slot, the chart in your report, or the video in your embed. An editorial guardrail built on it covers the words and must be paired with something else for pixels. Second, and more important for governance: a probability is a filter, not evidence. A 0.9 on some harm check is a useful signal for routing a draft. It is not a defensible audit trail about a person, and you should never treat a returned number as a finding of fact about an author's intent or conduct.
Also remember it is not a calculator and dates are text to it, not ordered quantities. If a check depends on counting occurrences or reasoning about whether a promotional window is still open, extract the values with a bounded Choice and do the arithmetic and date comparison in code. The model judges bounded questions well; it does not do math, and it does not write your schema for you.
Wiring the guardrail into publish hooks and held releases
The check belongs on the publish event, and content awaiting a human belongs in a held release rather than live. This is the CMS half of moderate-on-publish, and it is where a content platform's automation surface does the work. In Sanity, Functions are serverless hooks that run content automation on events like publish, which is exactly the place a moderate-on-publish check lives: on publish, a Function assembles the relevant fields as state, batches your typed questions in one call, and routes on the returned labels and confidence.
Staging is what makes fail-closed practical rather than disruptive. Content Releases let you stage, review, and schedule content as a group, so an item that lands in the send-for-review or fail-closed bucket goes into a held release an editor works through, instead of blocking the pipeline or leaking live. Allowed items publish; warned items publish with a flag; reviewed and blocked items wait where a human can see them. Content Lake real-time subscriptions mean a piece re-entering the pipeline after an edit can be re-judged the moment it changes, so a corrected draft does not need a manual re-trigger.
Store what came back. Persist the returned label, the confidence, the model version, and the state the model actually saw, as fields on the document or its review record. That record is what lets you audit which drafts were auto-allowed and why, retune a threshold against real outcomes, and demonstrate that your pipeline had a guardrail in place. Note the honest boundary from the microsite's stance: none of this is a shipped Jev integration or an announced partnership. It is an architectural claim about what a structured content platform must provide, feeding the exact field, staging the uncertain, and recording the verdict, for this class of model to be a trustworthy editorial guardrail.
How content platforms support feeding a typed decision at publish time
| Feature | Sanity | Contentful | WordPress | Strapi |
|---|---|---|---|---|
| Pass the exact field a check is about | GROQ queries the precise field, block, or section, so a claims check reads body copy and a disclosure check reads the footer, not the whole document. | GraphQL and REST expose fields, so you can select specific entries and fields to send as state with some client-side assembly. | Post content is one HTML field plus meta; extracting a single region for a targeted check means parsing markup yourself. | REST and GraphQL expose defined fields, so field-level selection is workable once your content types are modeled for it. |
| Rich text as clean, structured state | Portable Text stores rich text as typed blocks, marks, and annotations, so you pass structure a model can read without layout noise burning the token budget. | Rich Text serializes to a structured JSON document tree, which is far cleaner to pass as state than raw HTML. | Classic and block editors ultimately store HTML, so state tends to arrive as markup that must be stripped before judging. | Default rich text is stored as HTML or blocks depending on the editor plugin, so cleanliness of state depends on your setup. |
| Fit inside a 32,000-token state budget | Query only the fields a question needs, so a lean state stays well inside the budget and keeps input-billed cost down. | Selective field queries help, though large linked-entry graphs can inflate the payload if you fetch broadly. | Whole-post HTML plus meta is easy to overshoot with; trimming to a budget requires custom extraction. | Field selection controls payload size, so budget fit is achievable with deliberate query design. |
| Run the check on the publish event | Functions run serverless automation on publish, so the moderation call fires exactly where a guardrail belongs, with no external polling. | Webhooks fire on publish to an external function you host, which handles the check off-platform. | Hooks like publish_post and REST callbacks trigger PHP or an external service to run the check. | Lifecycle hooks and webhooks run custom logic on publish, so the check can be wired into the event. |
| Hold uncertain content out of live | Content Releases stage, review, and schedule content as a group, so review-bound and fail-closed items sit in a held release rather than going live. | Draft and scheduled states plus release features let you hold items back pending review. | Draft, pending review, and scheduled statuses hold content, though grouped release workflows need plugins. | Draft and published states plus custom workflow hold content, with grouped releases built via your own logic. |
| Re-judge on change | Content Lake real-time subscriptions surface edits the moment they happen, so a corrected draft can be re-checked without a manual re-trigger. | Webhooks fire again on update, so a changed entry can trigger a fresh check. | Save and update hooks re-fire, so re-checking on edit is possible via custom code. | Update lifecycle hooks re-fire, enabling a re-check when content changes. |
| Store label, confidence, and what the model saw | Persist the returned label, confidence, model version, and the exact state as document fields, giving a queryable record for audit and threshold tuning. | Add custom fields to the content type to store verdicts, with the state snapshot handled in your own store. | Post meta can hold verdict values; a full state snapshot usually lives in a separate table or store. | Extend the content type or a related model to store verdicts and a state snapshot alongside the document. |