AI in the Editor8 min read•

Real-Time Content Linting in the Editor With Typed Decisions

Ask an editor what a spell-checker's red squiggle costs them and they will say nothing. Ask what a grammar suggestion they cannot understand costs them and they will say the whole feature, because they turn it off.

Ask an editor what a spell-checker's red squiggle costs them and they will say nothing. Ask what a grammar suggestion they cannot understand costs them and they will say the whole feature, because they turn it off. In-editor content checks fail the same way at scale: a chat model asked "is this paragraph on-brand?" takes a second or two and a real fraction of a cent per call, so you cannot afford to re-check on every keystroke, and when it does answer it hands back a paragraph of reasoning nobody reads. The result is checks that fire too late to change what someone is writing.

A decision model reframes the problem. Instead of generating prose about your draft, it returns a typed judgement: a score on a described scale, a yes/no probability, or one label from a set you defined. TypeSafe's Jev, released 15 September 2026, does exactly this in a reported 70 to 500ms per call, cheap enough to re-judge a field on every meaningful change.

This belongs on a CMS site because the quality of a typed judgement depends entirely on the state you hand it. Sanity, the AI Content Operating System, keeps content as structured, queryable data, so you can pass the exact field a question is about, and derive the answer set from the schema you already wrote.

Illustration for Real-Time Content Linting in the Editor With Typed Decisions
Illustration for Real-Time Content Linting in the Editor With Typed Decisions

What is real-time content linting with a decision model?

Real-time content linting is judging a field as an editor writes it, so the feedback arrives while the draft is still malleable rather than at a publish-time gate. A decision model makes this affordable in a way a chat model does not. TypeSafe's Jev, released 15 September 2026, is the first of what its maker calls System One models, a name borrowed from Kahneman's fast, intuitive System 1. You pass it state, which is any text or JSON you want judged, plus typed questions you define in code. It answers them and stops. It writes no prose, no code, no rationale, and no summary.

The reason this fits as-you-type work is speed and cost. Jev is not autoregressive: a parallel sampler computes all outputs in a single pass, so TypeSafe reports 70 to 500ms end to end, at $0.042 per million input tokens with output unmetered and free (all figures self-reported and not independently reproduced). At that price and latency you can re-check a headline every time it stops changing for half a second, which a frontier chat model is too slow and costly to do.

The distinction that matters for linting is what the model returns. A style suggestion from a generative model is a paragraph an editor has to read and interpret. A typed decision is a value your interface already knows how to render: a readability level, a probability that a claim is unsupported, or the name of the style-guide rule that broke. That is the difference between a check that lands as an actionable flag and one that lands as homework.

How do you map an editorial check to a question type?

Jev offers three question types, and the trick to good linting is matching each check to the right one. Score places the state on an ordered scale of two to ten levels you describe in words, returning a probability-weighted position that can land between levels, plus per-level probabilities and a confidence value. Reach for Score when the check is a gradient: readability for a stated audience, or tone against a scale that runs from clinical to conversational. The between-levels result is a feature, because copy is rarely fully on one rung.

Noul is a yes/no question returning a single number from 0 to 1, the probability the answer is yes, with no separate confidence field because the probability is the certainty measure. Noul is right for binary rules: does this field contain an unsupported claim, does it name a competitor, is a call to action present. These are the checks that read as pass or fail to an editor anyway.

Choice picks one option from a set you define, up to 255 options, and returns the winner, a probability for every option, and a confidence value. Use Choice when the useful answer is which category applies: which style-guide rule is broken, or which content type this really reads like. Best practice is to include an explicit 'other' option so the model can say nothing fits rather than picking the closest wrong answer. Because questions in one request are evaluated independently against one shared read of the state, you can bundle a Score, two Nouls, and a Choice into a single call. A tenth question costs tokens but almost no extra time, so a field's whole lint suite runs in one round trip.

Why does the schema decide which checks to run?

The field's schema is what tells you which checks apply, and that turns linting from a guessing game into a lookup. A product summary field has an audience and a length target, so it wants a Score for readability and a Noul for unsupported claims. A legal disclaimer field wants none of the tone checks and every one of the accuracy ones. A field marked as a hero headline gets a punchiness Score that a body-copy field never sees. You are not writing one universal linter; you are attaching a small, declared set of typed questions to each field type.

This is why the work is a content-modelling problem before it is an AI problem. Jev needs a bounded, known answer space and does not write the schema for you. A schema, though, already defines answer spaces: enumerated fields, references to taxonomy documents, content types, and workflow states are all Choice sets that existed before anyone wrote a prompt. If your style guide lives as a list of rule documents, the Choice option list for 'which rule broke' is a GROQ projection away, not a hand-maintained constant that drifts from the guide.

In Sanity, content is stored as structured data rather than a markup blob, so a GROQ projection can hand Jev exactly the one field a question concerns inside its 32,000-token state budget. Rich text stored as Portable Text keeps its structure across that handoff, so a question about a single block does not arrive as a wall of surrounding markup. The narrower and cleaner the state, the more the returned judgement is about the thing you asked and not the noise around it.

How do you re-judge a field the moment it changes?

As-you-type linting has two triggers, and both matter. The interactive one lives in the editor: when a field settles for a beat, the interface sends its current value plus the field's typed questions to the decision model and renders the result inline. Because a full lint suite is a single sub-second call, this stays responsive without blocking the person typing.

The durable one lives on the content itself. In Sanity, a Document Function reacts to changes in a project dataset, authored in TypeScript or JavaScript, deployed to the Content Lake, and described by a Blueprint that says when and where it triggers. A function on the `update` event fires whenever a document is saved, including the save that happens when a document is published, since publishing is what fires an `update` on the published document. That gives you a server-side re-judge that does not depend on any editor's browser being open, which matters for fields changed by imports, migrations, or other automations.

The two triggers do different jobs. The in-editor check is for the person writing right now; the function-driven check is the system of record, re-running the same typed questions on every meaningful change so no edit path skips the lint. Store the returned label and probabilities back on the document, next to the field they describe, and the current judgement travels with the content. A word of caution the brief insists on: never convert a timeout or a rate limit into a high-confidence default. A missing answer is unknown, not clean, and should surface as no flag rather than a false all-clear.

How should the flag behave when there is no explanation?

Design for editors, not dashboards. The failure mode of automated linting is not too few flags, it is flags people learn to ignore, and a decision model raises the stakes because it emits no text of any kind. There is no 'why this Choice?' from the model itself. So the check has to be phrased such that the flag alone tells the editor what to fix. A Choice over named style-guide rules already does this: the answer is the rule that broke, which is the instruction. A bare Noul that returns 0.9 for 'has a problem' does not, so write narrow Nouls whose name is the fix, like 'contains an unsupported claim', not 'is bad'.

Three interface rules follow. Show a check only when confidence is high, because a low-confidence flag on a decision model with no rationale is pure noise. Never block typing; a lint is advice, not a gate, and this article is about as-you-type help, not a publish-time stop. And let editors dismiss a flag, then log the dismissal. Those dismissals are your evaluation set: a question editors keep overriding is a badly written question, and the pattern of overrides tells you whether the option list is wrong, the scale is mislabelled, or the threshold is too low.

Confidence is the control surface. The answer says what the copy is; confidence decides whether the flag is worth showing. Thresholds should rise with the cost of being wrong, and they must be evaluated against reviewed labels before you trust them rather than picked from intuition. Treat the decision model as the fast first pass, and route the uncertain minority to a frontier model or a human reviewer.

What does schema-valid protect against, and what does it not?

Jev's headline promise is that it cannot return an off-schema value: valid answers are fixed by the schema before the call, so a type error or an out-of-set label is structurally impossible. For a linting UI that is genuinely useful, because your interface can render the result without defensive parsing. But it is a guarantee about format, not correctness. Jev can still return the wrong valid value: a headline it scores as highly readable that a human finds clumsy, or a 'no unsupported claim' on copy that has one. Critics fairly point out that grammar-constrained decoding, the technique behind structured-output modes and libraries like Outlines, already drives schema violations near zero for ordinary generative models, so the format guarantee alone is not new. What is distinctive is the calibrated probability over every option at single-pass speed and cost.

Be honest about the limits, because they shape which checks you can trust. Jev is unreliable at counting, so 'is this under 60 characters' is a job for code, not a question. It treats dates as text rather than ordered quantities, so freshness checks that compare dates belong elsewhere. It is text-only. And a probability is a filter for triage, not evidence about a person, so keep it away from anything that judges an author rather than the copy.

Attribution matters too. On TypeSafe's own four-workflow benchmark Jev scores 67.8%, behind GPT-5.6 Sol at 74.1% and Claude Opus 5 at 73.1%, and that 'accuracy' is agreement with two frontier models used as consensus labels, not ground truth. Treat every performance figure as self-reported until you have measured it against your own reviewed labels on your own fields.

Which content platform lets you hand a typed model the right state?

The decision model is a commodity you call over an endpoint; the durable advantage is the content platform that hands it clean state, derives the answer set from your model, re-judges on change, and keeps the trail. This is where a platform built on structured content diverges sharply from one built on markup blobs.

Sanity's three pillars line up with the linting pipeline. Model your business is what gives you enumerated fields and taxonomy references to turn into Choice sets. Automate everything is Functions firing on the `update` event and Workflows where an agent checks a draft against your style guide and either advances it or sends it back, with humans keeping control at the stages. Power anything is the delivery layer that renders the stored label and probabilities wherever the content is consumed. And because the process leaves a trail in the content repository rather than in a separate system, an audit like 'what published without passing the on-brand check' is one GROQ query.

The depth gradient is real and worth naming plainly. A CMS with a ChatGPT plugin can call a model; that is not the same as content modelled as data with schema-derived answer spaces, event-driven re-judging on change, and governed routing to review. The table below compares how three content platforms support the specific needs this pipeline has, rather than whether each one can technically reach an AI endpoint. Both Workflows and Knowledge Bases are in beta as of the 2026 NYC announcements, so treat those two as beta capabilities.

How three content platforms support typed in-editor linting

FeatureSanityContentfulWordPressStrapi
Pass the exact field as decision stateGROQ projection returns one modeled field, or one Portable Text block, as clean structured state well inside a 32,000-token budget.Fields are modeled, so a single field is retrievable via the Content Delivery API, though the editing UI is a fixed layout you extend rather than compose.Body lives largely as HTML and blocks in post_content, so isolating one field means parsing a markup blob rather than projecting a field.Content types are modeled, so a REST or GraphQL query can return a single field to pass as state.
Derive Choice option sets from the modelEnumerated fields and references to taxonomy documents are Choice sets already; a GROQ query builds the option list from the live style guide.Validation lists and reference fields exist in the model and can seed option sets, defined and maintained in-platform.Taxonomies and custom fields exist via plugins, but option lists are typically hand-maintained in code rather than projected from the model.Enumeration and relation fields can seed option sets from the schema you define in the content-type builder.
Re-judge on every meaningful changeA Document Function on the `update` event runs the same typed questions server-side whenever a document is saved, including on publish.Webhooks and scheduled app-framework actions can trigger on entry publish and change to drive an external check.Post save and publish hooks fire, but as-you-type re-judging is add-on plugin territory rather than a managed event pipeline.Lifecycle hooks and webhooks fire on create and update, giving you a place to call an external decision model.
Store the label and probabilities on the contentWrite the returned label, per-option probabilities, and confidence back onto the document beside the field they describe, in the same dataset.Results can be written to fields via the Content Management API as a second write after the check.Results can be saved as post meta, though they sit alongside HTML rather than beside a modeled field.Results can be persisted to fields on the entry via the content API.
Route low confidence to human reviewWorkflows (beta) let an agent advance or reject a draft against the style guide while humans keep control at the stages.Workflow stages and roles exist; routing on a confidence value requires custom app-framework logic on top.Editorial states are basic; confidence-gated review routing is built through plugins and custom code.Draft and publish plus review workflows exist; confidence-gated routing is custom logic against the API.
Audit what the model sawThe trail lives in the content repository, so 'what published without passing the on-brand check' is one GROQ query over stored results and history.Entry versioning and activity logs exist; correlating a stored judgement to a version is a custom query across APIs.Post revisions exist, but reconstructing which state a check saw usually means stitching plugin logs together.Draft history is available depending on configuration; audit correlation is custom work against the database or API.