In September 2026 TypeSafe AI released Jev, a model that cannot chat. You give it state and typed questions, and it returns typed answers with probabilities in milliseconds, for a fraction of a cent. This course covers what Jev is and the argument behind it. It explains what “calibrated” and “can’t hallucinate” actually guarantee, what the published evidence shows, and why a model named after William Stanley Jevons is really a bet on the economics of cheap intelligence.
Built on Introducing System One Models & Jev (Diogo Almeida, 15 Sep 2026) and the TypeSafe documentation
Jev is two weeks old and every performance number about it so far is TypeSafe’s own. TypeSafe is unusually candid about that. Its launch post attaches a “Nuance” box to each claim, and its docs publish a list of the model’s known failure modes. The course uses both. Objections that go beyond what TypeSafe concedes are in Module 7 and are labelled as objections.
Short on time? Do Module 1 (what it is), Module 4 (what “can’t hallucinate” means) and Module 6 (the Jevons bet).
The launch post opens with a question: “Models have been superhuman at chat for years, so where is all the automation?” It comes from Diogo Almeida, TypeSafe’s founder and an author of the InstructGPT work at OpenAI that became the research behind ChatGPT. His answer is a new class of model:
“Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.” Diogo Almeida, TypeSafe, 15 September 2026
You send a state (text: a string, a JSON object or an array of text) and a set of questions. Jev evaluates every question against the state in parallel and returns typed answers. It never produces prose.
| Primitive | Asks | Returns | Example |
|---|---|---|---|
| Choice | Which of these options? | choice, probabilities, confidence | Route a ticket to billing, technical or account |
| Score | Which level on a scale you define? | score, probabilities, confidence | How frustrated is this customer, 0–2 |
| Noul | Is this true? | noul, a probability from 0 to 1 | Does this message request a refund? |
response = client.system_one(
state=ticket_and_policy,
questions={
"refund_requested": Noul(instructions="Does the customer ask for a refund?"),
"team": Choice(instructions="Which team should handle this?",
criteria={"billing": ..., "technical": ..., "account": ...}),
"frustration": Score(...),
})
if response.answers["team"].confidence < 0.5:
route_to_human(ticket)
The docs’ design rule matters as much as the API: ask atomic questions, “the kind of judgment a highly knowledgeable person could make in a few seconds,” and combine them in code. Instead of “rate this startup pitch,” ask about market, feasibility and differentiation separately, then weight them yourself. When priorities shift, “change a coefficient in your code rather than rewriting a prompt.”
System One is from Daniel Kahneman’s Thinking, Fast and Slow: fast, intuitive judgment rather than slow, deliberate reasoning. Jev is from William Stanley Jevons, the economist who observed that more efficient steam engines increased coal consumption. That second name is a thesis, and Module 6 takes it seriously.
P3 has a name in TypeSafe’s writing. Rich Sutton’s bitter lesson says general methods that use more compute beat clever algorithms. Almeida extends it: “The bitterest lesson in ML is that doing the right task > data > compute > algorithms.” His evidence is the precedent he worked on. At OpenAI, InstructGPT models more than 100× smaller than GPT-3, trained on the right task (following instructions), were preferred to GPT-3 itself. The claim for Jev is the same move again: chat was the right task for assistants, and calibrated decisions are the right task for automation.
P1 is widely shared, and P3 has a real precedent. The premise that decides whether Jev matters is P2’s forecast: that the bulk of valuable automation is many small, fast, bounded judgments embedded in code, rather than long open-ended agent tasks. If that is right, a System One model is the natural component. If most value turns out to be in open-ended work, Jev is a very good router for systems whose core is still an LLM.
Across many predictions from a well-calibrated model, outcomes given probability 0.8 happen about 80% of the time, and those given 0.2 about 20%. The docs are exact about the limit: “Calibration is measured across groups of predictions; it does not guarantee that an individual answer is correct.” Calibration is a property of the population of answers. That makes it the right tool for setting thresholds and estimating error rates, and useless as a promise about any one case.
The launch post puts the case in one line: “If a model can do a task 95% of the time but doesn’t say when it’s in the 5%, it can’t automate that task.” LLMs asked for a confidence estimate tend to be “overconfident and inconsistent.” A calibrated model lets code act on the 95% and route the 5%.
For Choice and Score, confidence summarises how peaked the distribution is: 1.0 when all probability sits on one option, 0 when
it is spread evenly. The docs’ illustration uses (n × largest probability − 1) / (n − 1) for n options. For three options at
90/6/4, that is (3 × 0.9 − 1) / 2 = 0.85. Noul answers have no separate confidence: the probability is the signal.
The docs’ banking example routes anything below 0.5 confidence to a human. It shows a balance on any confident read, because the wrong screen is recoverable. It lets a transfer proceed only above 0.9, and even then with confirmation. One model, one confidence number, different thresholds per action, because the cost of being wrong differs. The code, not the model, encodes the risk tolerance.
The launch post says Jev “can’t hallucinate” and, more precisely, that it “never makes type errors”: possible outputs are defined in advance, so an answer is always one of your options. TypeSafe calls this falsifiable by a single counter-example and “mathematically impossible” to violate, and in its hallucination chart it enters 0% for Jev with the note, “Our number is not empirical. Schema matching is guaranteed.” That is a genuine property with real value. In a system with latency guarantees, a malformed tool call “buried several layers deep in a dependency chain” is a deal-breaker, and Jev cannot produce one.
Type safety removes invented values. It does not remove wrong choices. A Choice between “billing” and “technical” can never return “refund-department”, but it can return “billing” when the answer is “technical”. What “can’t hallucinate” buys is that every error is a legal error, one your code can handle, measure and route. That is valuable, and it is not the same as being right.
TypeSafe’s own jaggedness page gives the clearest example. On the ticket “I was charged twice for the same order,” two Nouls asked “is the customer asking for a refund?” and “is the customer asking for something other than a refund?” The answers were 0.72 and 0.47, summing to 1.19. Each question is evaluated in isolation, so nothing forces logically related answers to agree. The docs’ advice: “don’t hold the model to arithmetic identities between separate questions,” and enforce the identities you need in code.
| Failure mode | Their fix |
|---|---|
| Literal reading: “answers the question you wrote, not the one you meant” | State the exact condition; put boundary cases in the criteria |
| Math, counting, numeric representations | “Jev is not a calculator.” Keep arithmetic in code |
| Date and time comparison | Extract components with Choice; compare in code |
| Indirection and double negatives | Fewer hops; name the relevant part of the state |
| Large state full of irrelevant detail | Filter first; send only what the question needs |
| Adversarial content in the state | Precise criteria; test edge cases before deploying |
| Generation | “There are other models for that” |
A refund workflow. Assign each step to Jev, plain code, a generative LLM, or a human.
TypeSafe is careful even here. On speed, its evals “are generally run from our laptops on the West Coast,” where the service is based. On price: “We can’t prove it isn’t subsidized,” though it expects prices to go down, not up.
The headline ratios on TypeSafe’s home page, 193.6× faster and 444.6× cheaper, come from its “workflow evals”. The design is specific and worth understanding. There is no ground-truth label set. Each workflow is fixed in code, and the reference answer is the average prediction of GPT-6 Astra and Fable 5.1, the largest and most expensive models. Every model is scored on how closely it matches that reference. TypeSafe’s own caveats:
TypeSafe also deliberately publishes no public-benchmark results. It argues users should build evals for their own use cases, since “System One tasks are much easier to evaluate.”
Matching the average of two frontier models at a fraction of the cost is strong evidence that Jev is a cheap substitute for frontier judgment on those workflows. It cannot show Jev is more accurate than the reference, because disagreeing with the reference counts as error even when the reference is wrong. It cannot show accuracy against real outcomes either. For your own use, the implication is concrete: the eval that matters is one you build with labelled outcomes from your own data, which is what TypeSafe tells you to do.
In 1865 William Stanley Jevons argued that more efficient steam engines would raise, not lower, Britain’s coal consumption, because cheaper useful work would find many more uses. TypeSafe’s FAQ applies it directly: “We expect machine intelligence to follow a similar path to coal… Every order of magnitude drop in the cost of intelligence unlocks orders of magnitude more use cases.” The paradox holds only when demand is elastic, meaning a price fall produces a proportionally bigger rise in use. That is the empirical question.
The launch demos show it. A bot plays Doom by asking Jev 10 questions a second, which the team puts at about $7 an hour. An agent wiki-races across Wikipedia, choosing among hundreds of links at each step. Neither is a sensible use of a model that takes seconds per call. The docs’ patterns go further: speculative fan-out (ask every question you might need in one call and let code ignore the rest), map-reduce over big data (turn a corpus into features), and verify everything (score and guardrail every LLM output). These are workloads that only exist when a judgment costs almost nothing.
The 13×/yr rate is Epoch’s average for reaching fixed benchmark scores (“The plunging price of thought”, Sep 2026), not a measurement on classification workloads. It is a yardstick, not a forecast.
The arithmetic cuts both ways. A 100× lead is under two years of frontier price decline, and even the 444.6× headline is under two and a half. So a pure price lead is not durable if LLMs keep falling and Jev stands still. But two things are not captured by the price decline. Latency is architectural: parallel evaluation of all questions in one pass does not arrive just because tokens got cheaper. And Jev’s own cost falls too. TypeSafe expects its prices to go down. Set both decline rates equal and the gap never closes. The durable claim is the interface and the architecture, not the price.
LLM providers have offered JSON schemas, constrained decoding and log-probabilities for years. A team can get typed answers from an LLM today. What is Jev adding beyond price and speed?
Decomposing judgments into atomic questions, filtering state, keeping dates and arithmetic in code, and enforcing cross-question identities are all engineering. A system built from dozens of small calls is more legible, but it is also more to design and maintain than one well-prompted LLM call.
Every performance number is TypeSafe’s, measured against a reference-model average on workflows its own team wrote, with no public benchmarks by choice. Calibration, the central property, has not yet been independently measured at scale.
If fast, calibrated, typed decisions are valuable, the large labs can train for them too, and they already own the distribution. The bitterest lesson cuts both ways: if the right task is what matters, anyone can pick the right task once it has been shown.
Reach for a System One model when all four hold: the answer is one of a known set of options or levels; the judgment is fast and semantic rather than arithmetic or multi-hop; the decision runs at volume or needs sub-second latency; and you can build a labelled eval to set thresholds. Where any fails, use code, a generative model or a person, or put the System One model in front of them as a router.