// Journal · Sep 24, 2026 · 5 min read
What are System One models? TypeSafe's Jev, explained

What are System One models? They are a new class of AI model that returns typed, structured decisions with calibrated probabilities instead of generated text. The name comes from TypeSafe AI, a San Francisco AI lab founded in 2026, whose first public System One model — Jev, currently in early access — takes your program’s state plus a question with a predefined answer schema, and returns one of the allowed answers with a probability attached. Because every valid output is declared up front, the model can’t produce a type error or hallucinate a value that doesn’t exist. The trade: it decides, it doesn’t generate.
The name is a nod to Kahneman’s fast-versus-slow thinking. Frontier LLMs are the deliberate System Two: they reason, generate, and explain. TypeSafe’s pitch, laid out in their announcement of System One models and Jev, is that most decisions inside an automated workflow don’t need any of that — they need a fast, cheap, reliable answer from a known set. That matches what we see building automations: the expensive model spends most of its tokens on classification and routing calls that never needed generation in the first place.
What TypeSafe AI’s Jev actually does
Jev accepts unstructured input — text, with an emphasis on structured program state — and answers questions through three primitives, batched into a single request:
- Choice — pick one option from a declared set (up to 255 options): which team handles this ticket?
- Score — place the input on a declared spectrum of 2–10 levels, with between-level precision: how frustrated is this customer?
- Noul — a yes/no returned as a probability from 0 to 1: does the customer explicitly request a refund?
Every answer comes back as a typed value plus a per-option probability breakdown, and (for Choice and Score) a confidence measure. The calibration is the actual product. Jev is trained with a method TypeSafe calls Reinforcement Learning for Calibrated Decisions (RLCD), which optimises the probabilities against real outcomes — so, in their words, “higher confidence really does mean higher accuracy.” SDKs exist for Python 3.10+ (typesafe-sdk) and Node 20+ (@typesafe-ai/sdk).
That calibration claim is the one worth testing hardest in your own evaluation, because everything useful downstream depends on it.
Why typed decisions matter when software acts on AI output
Anyone who has wired an LLM into a real workflow knows the failure modes: the model returns JSON with a category that isn’t in your enum, wraps the answer in prose, or — worse — returns a perfectly parseable wrong answer with no signal that it was guessing. JSON mode and function calling fix the syntax; they don’t tell you when to trust the value.
A schema-constrained decision with a calibrated probability changes what your code can do:
- Act automatically when confidence is high. Route the ticket, apply the label, approve the match — no human involved.
- Escalate when confidence is low. Below the threshold, the case goes to a person instead of silently proceeding on a guess.
That second branch is exactly the review queue pattern we ship in every automation — except the threshold is now a number the model is trained to make meaningful, rather than one you eyeball from logprobs and hope generalises.
The numbers: 70–500ms, $0.042 per million input tokens
TypeSafe’s published figures for Jev, from their announcement:
- Latency: 70ms–500ms end-to-end — 40–200× faster than frontier LLMs on equivalent tasks, by their measurement.
- Price: $0.042 per million input tokens; output tokens are free (“too cheap to meter”).
- Accuracy: on their workflow evaluations they claim the Pareto frontier by “almost 2 orders of magnitude.”
For calibration of your own: frontier model APIs in 2026 price input at roughly $2–$15 per million tokens. If those vendor numbers hold up independently, a decision that costs a frontier model a visible line on your invoice costs Jev effectively nothing — which is what makes “run it on every record” viable rather than aspirational. The usual caveat applies: these are the vendor’s own benchmarks for an early-access product. Measure on your data before you believe any of it.
When to use a System One model instead of an LLM
Jev is deliberately narrow, and the practical guides are refreshingly blunt about it. Use it when:
- The answer space is known up front. Routing, triage, tagging, relevance filtering, yes/no verification — anything you could express as an enum.
- Volume is high and the decision repeats. Thousands of tickets, records, or events through the same question.
- Latency is user-facing. At 100ms you can put a model decision inside an interaction without the UX paying for it.
Don’t use it when you need generation, explanation, or open-ended reasoning. Jev cannot write text, can’t reliably count or treat dates as quantities, and reads instructions literally — negations and scoping words land at face value. It’s also not exempt from the context problems that degrade AI agents: accuracy drops as you stuff its input, so retrieve precisely before you ask.
The deployment pattern TypeSafe’s early users describe is the sensible one: Jev classifies and routes cheaply, code handles what it can, a frontier model takes the hard minority. Three tiers, each doing the job it’s actually good at.
Our take: the model is becoming a component, not the system
Havoric is an AI automation and web development agency — we automate repetitive manual processes and build the web and mobile apps around them, and most of the decisions inside those automations are exactly the shape Jev targets: pick a queue, score an urgency, answer yes/no. Today we make those calls with general-purpose LLMs constrained by schemas, and they are routinely the slowest, priciest step in the pipeline for the least interesting work.
We haven’t shipped Jev to production — it’s early access, and vendor benchmarks are vendor benchmarks. But the direction matters regardless of whether TypeSafe wins it: purpose-built decision models with honest confidence scores turn “should a human check this?” from a heuristic into an engineering parameter. That’s the question every AI automation project we scope ultimately turns on — and it’s a better foundation for autonomy than a bigger model that’s still, underneath, guessing in prose.