Listicle6 min read•

Top 5 Content Operations Worth Automating With a System One Model

A back catalogue of 40,000 articles sits with half its tags missing or wrong, so search returns junk, related-content modules fire blanks, and the retrieval step feeding your chatbot pulls the wrong pages.

A back catalogue of 40,000 articles sits with half its tags missing or wrong, so search returns junk, related-content modules fire blanks, and the retrieval step feeding your chatbot pulls the wrong pages. You could pay a frontier LLM to re-tag the whole archive, but the per-document cost turns a cleanup into a budget line nobody approves. So the archive stays broken, and every downstream AI feature inherits the mess.

On 15 September 2026, TypeSafe released Jev, which it calls the first "System One model": a model that reads state and returns a typed judgement, a Choice, a Score, or a yes/no probability, and stops. No prose, no rationale. TypeSafe bills input only at $0.042 per million tokens, leaves output free, and reports that batching thirteen questions into one call runs 12.2x cheaper and 10x faster than asking them separately. That economics turns "possible" into "viable" for whole classes of content operations.

The catch is that a typed decision is only as good as the state you hand it, which makes this a content-modelling problem first. Sanity, the AI-native content platform, is where that groundwork pays off: because content is stored as structured, queryable data rather than opaque markup, you can hand a typed model the exact field, block, or section a question is about, inside a 32,000-token budget.

Illustration for Top 5 Content Operations Worth Automating With a System One Model
Illustration for Top 5 Content Operations Worth Automating With a System One Model

1. Taxonomy and auto-tagging across a back catalogue

Taxonomy and auto-tagging is the clearest win because the economics are documented and the work is embarrassingly parallel. The state is one document at a time: its title, its summary, and the body fields that carry meaning. The question is a Choice over your controlled vocabulary, up to 255 options, with an explicit "other" option so the model can say nothing fits rather than forcing a wrong tag. For a multi-facet taxonomy you batch several Choice questions into one call, one per facet, against a single read of the same state, and the second through tenth questions cost tokens but almost no extra time.

The concrete example is Hassan El Mghari's 1kpapers.com, which classified 1,018 research papers across 24 candidate topics for $0.08 total at a median 256ms per paper. He had already paid $3.99 to a generative model for the summaries, which is the point worth internalizing: one model writes the prose, a different, far cheaper model makes the decision. Re-tagging a 40,000-document archive at that rate is a coffee-sized expense, not a capital request.

Where it fits poorly: Jev is not a calculator and it treats dates as text, not ordered quantities, so a Choice like "which decade does this cover" over free-form date mentions is unreliable. Enumerate the buckets explicitly and let it pick a labeled option. The failure mode to design around is the closest-wrong-answer problem: without an "other" escape hatch, a document that belongs to none of your 24 topics gets filed under the least bad one, quietly polluting the taxonomy you built this to clean.

This is a content-modelling problem before it is an AI one. Content stored as Portable Text lets you pass exactly the heading, the intro block, or the abstract as state instead of a wall of HTML, and a GROQ query can select just those fields so the 32,000-token state budget goes to signal, not markup.

2. Moderation and brand-safety checks at publish time

Moderation and brand-safety checks pay back fast because the cost of one bad publish, a slur in user-generated copy, a competitor named in a place style forbids, a claim legal has not cleared, dwarfs the fraction of a cent it costs to check. The state is the draft, or the specific fields most likely to carry risk. The natural question type is a Score on an ordered severity scale of two to ten levels described in words, returning a probability-weighted position that can land between levels, plus a confidence value. A Noul, yes-or-no returning a single probability from 0 to 1, fits a narrow gate like "does this contain profanity."

The production pattern here is confidence-gated routing. The Score says how risky the content is; the confidence decides whether it is safe to act on automatically. Thresholds should rise with the cost and irreversibility of the mistake, so a low-severity, high-confidence result publishes, an anything-ambiguous result queues for a human, and a high-severity result blocks. Because Jev returns a typed value fixed by your schema, it structurally cannot return an off-schema label or write an essay instead of a verdict; TypeSafe reports a 0% structured-output error rate.

Read that guarantee precisely. It is about format, not correctness. Jev cannot invent a fourth severity band, but it can still place genuinely unsafe content one level too low. The 0% figure is asserted from schema design, not measured, and repeating "cannot hallucinate" without that distinction is misleading. It is also text-only, so image and video moderation are out of scope entirely.

The failure mode to design around is the timeout. Confidence only exists when a request succeeds, so never convert a rate limit into a high-confidence pass. Use a deterministic fallback that blocks or queues. In Sanity, a Function on publish can call the model, and Content Releases give you the staging and review surface for anything the check sends back for a human decision.

3. Relevance filtering before an expensive retrieval step

Relevance filtering pays back by moving a cheap decision in front of an expensive one. Before you embed a chunk, run a full-text retrieval, or feed a passage into a frontier model's context window, a System One model can ask whether the candidate is actually relevant to the query at hand, and drop the two-thirds that are not. The state is the query plus a candidate passage; the question is a Noul, "is this passage relevant to this query," returning a probability that is itself the certainty measure, with no separate confidence field because the probability IS the confidence.

The economics are the whole argument. At $0.042 per million input tokens with free output, filtering a thousand candidates costs cents, while the retrieval and generation steps you are gating cost dollars. TypeSafe reports end-to-end response times of 70 to 500ms, so the filter adds little latency, and because output is unmetered you pay only for the state you send in. The escalation cascade is the shape to build: Jev filters cheaply and at volume, code handles what it can, and only what clears the relevance bar reaches the expensive model.

Where it fits poorly: a probability is a filter, not evidence, and it is not a ranking function you should trust to fine-grained ordering. It tells you likely-relevant from likely-not, not that passage A outranks passage B by a hair. For true semantic ranking you still want vector similarity. The failure mode is threshold drift: a Noul cutoff tuned on easy cases silently discards good passages once queries get ambiguous or context goes missing, so evaluate thresholds against reviewed labels that cover easy, ambiguous, missing-context, and edge cases before trusting them.

This is where content structure and embeddings meet. Sanity's Embeddings Index API ties embeddings to content so freshness is automatic, and because content is queryable you can hand the filter exactly the passage in question rather than a whole document, keeping each check inside the state budget.

4. Routing inbound content to the right owner

Routing inbound content, submissions, reader feedback, support articles, bug reports, is a triage problem, and triage is exactly what a Choice question does. The state is the incoming item: the subject line, the body, whatever structured metadata came with it. The question picks one owner or queue from a defined set of up to 255 options, returns the winner, a probability for every option, and a confidence value, and you route on the winner while gating on the confidence. Add an explicit "needs a human to triage" option so ambiguous items land somewhere deliberate instead of being forced into the nearest wrong bucket.

Batching is what makes this cheap at scale. In one call against a single read of the item you can ask several Choice questions at once, owner, priority, and language, and the second and third questions add tokens but almost no time. TypeSafe's cookbook reports batching thirteen questions into one call at 12.2x cheaper and 10x faster than asking them separately, with identical answers. A support desk fielding thousands of items a day routes them for the price of a rounding error.

Where it fits poorly: routing that depends on knowing which of two dates came first, or on counting how many prior tickets a sender filed, is off the table, because Jev treats dates as text and its counting is unreliable and gets worse as the count grows. Do that arithmetic in code and hand Jev only the bounded classification. The failure mode to design around is the moving alias: pin the versioned model ID, jev-1.13.0 rather than jev-latest, and log the model reported in each response, or an alias can shift under you and reroute everything with no release.

Sanity fits this as the destination as much as the classifier. Content Lake real-time subscriptions surface an inbound item the moment it lands, a Function runs the Choice call, and the routed item drops into the right Studio workspace with its label and confidence stored on the document.

5. Scoring drafts against an editorial rubric before human review

Scoring drafts against an editorial rubric pays back last of the five because the volume is lower and the judgement is softer, but it still earns its place by front-loading the easy calls so editors spend attention where it matters. The state is the draft, or the section under review. The question is a Score over a rubric expressed as an ordered scale of levels described in words: clarity, structure, adherence to house style, each returning a probability-weighted position that can land between levels, so a draft can score 1.035 rather than being flattened to a whole number. Batch one Score per rubric dimension into a single call.

The point is not to replace the editor. It is to let a human open the drafts that scored poorly first and skim the ones that scored well, turning a flat review queue into a triaged one. Because output is free and billing is input-only, re-scoring a draft on every save costs almost nothing, so the rubric score can live on the document and update as the writer revises.

Where it fits poorly, and this is the honest limit of the whole class: Jev emits no text of any kind, no summary, no rationale, no suggested rewrite. It can tell you a draft scores low on clarity; it cannot tell you which sentence to fix. If you want words, a generative model produces the feedback and Jev scores it, two different models for two different jobs. The failure mode to design around is treating a low score as a verdict rather than a filter. A probability is not a defensible record about a writer's competence, so keep it as a routing signal, not a performance metric.

Sanity's AI Assist covers the generative half of this split honestly: it can rewrite a block in a different voice or fact-check a claim against a knowledge base inside the Studio, while a typed Score decides which drafts need that attention first. Structured content is what lets you score the section that matters instead of the whole page.

How content platforms support feeding a typed decision model

FeatureSanityContentfulWordPressStrapi
Pass the exact field a question is aboutGROQ selects a single block or section, so state is signal not markup; Portable Text keeps structure intact under selection.REST and GraphQL return fields, though rich text ships as a nested JSON tree you often flatten before sending as state.Rich text is stored as one HTML blob per post, so isolating a single block usually means parsing markup yourself.REST and GraphQL expose fields; rich-text blocks are addressable, though selecting sub-blocks takes custom querying.
Rich text as clean, structured statePortable Text preserves marks, annotations, and blocks as data, so structure survives chunking without an HTML round-trip.Rich Text is a structured document node; usable as state after you serialize the tree to text.HTML strings must be stripped or parsed before use, which risks dropping structure the model could use.Blocks plugin stores structured rich text; default field stores markdown or HTML depending on setup.
Fit the 32,000-token state budgetField-level GROQ projections send only the fields in scope, keeping large documents inside budget without truncation.Achievable by selecting fields in the query, though nested references can inflate payloads if not pruned.Whole-post payloads are common; trimming to budget typically needs custom code or a serialization filter.Field selection and populate control keep payloads lean when configured deliberately.
Re-judge content the moment it changesContent Lake real-time subscriptions and Functions fire on publish, giving a change event to re-run the decision against.Webhooks fire on publish and change, so a serverless handler can trigger a re-judge.Hooks and REST webhooks (often via a plugin) can trigger on save; reliability depends on the plugin.Lifecycle hooks and webhooks trigger on create and update, wired in the application layer.
Store the returned label and confidenceWrite the Choice, Score, or Noul value plus confidence back onto typed fields on the same document via the write API.Add fields to the content type and write results back through the Management API.Store as post meta via the REST API or a custom field plugin.Add columns to the content type and persist via the REST or entity API.
Route a low-confidence result to reviewContent Releases and Studio Workspaces stage and review flagged items; Roles & Permissions gate who resolves them.Workflows (on higher tiers) and tasks support review states for flagged content.Draft and pending states plus editorial-workflow plugins provide a review queue.Draft and publish states plus custom review logic in the admin panel.
Audit what the model sawAudit logs and Content Source Maps trace which content version was live; store the input snapshot alongside the result.Version history and API logs help reconstruct state; capturing the exact input requires storing it yourself.Post revisions record content changes; capturing the model input is a custom responsibility.Revisions where enabled record changes; input capture is left to the application.