Jev: AI That Doesn’t Talk

In September 2026 TypeSafe AI released Jev, a model that cannot chat. You give it state and typed questions, and it returns typed answers with probabilities in milliseconds, for a fraction of a cent. This course covers what Jev is and the argument behind it. It explains what “calibrated” and “can’t hallucinate” actually guarantee, what the published evidence shows, and why a model named after William Stanley Jevons is really a bet on the economics of cheap intelligence.

7 modules
2 interactive tools
15 reasoning questions
~55 min

Built on Introducing System One Models & Jev (Diogo Almeida, 15 Sep 2026) and the TypeSafe documentation

How to read this course

Jev is two weeks old and every performance number about it so far is TypeSafe’s own. TypeSafe is unusually candid about that. Its launch post attaches a “Nuance” box to each claim, and its docs publish a list of the model’s known failure modes. The course uses both. Objections that go beyond what TypeSafe concedes are in Module 7 and are labelled as objections.

Short on time? Do Module 1 (what it is), Module 4 (what “can’t hallucinate” means) and Module 6 (the Jevons bet).

Course Modules

  1. A function call, not a conversationStart here
  2. The argument: the interface is the bottleneckPremises
  3. Calibration and confidenceCore concept
  4. What “can’t hallucinate” guaranteesInteractive
  5. The evidence, and how to read itEvaluation
  6. Why it is named after JevonsInteractive
  7. Objections, and when to use itInteractive
1

A function call, not a conversation

State in, typed probabilistic decisions out
By the end of this module you will
  • Describe what Jev takes in and returns, and why that differs in kind from an LLM
  • Know the three question types and when each fits

The launch post opens with a question: “Models have been superhuman at chat for years, so where is all the automation?” It comes from Diogo Almeida, TypeSafe’s founder and an author of the InstructGPT work at OpenAI that became the research behind ChatGPT. His answer is a new class of model:

“Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.” Diogo Almeida, TypeSafe, 15 September 2026

What goes in, what comes out

You send a state (text: a string, a JSON object or an array of text) and a set of questions. Jev evaluates every question against the state in parallel and returns typed answers. It never produces prose.

PrimitiveAsksReturnsExample
ChoiceWhich of these options?choice, probabilities, confidenceRoute a ticket to billing, technical or account
ScoreWhich level on a scale you define?score, probabilities, confidenceHow frustrated is this customer, 0–2
NoulIs this true?noul, a probability from 0 to 1Does this message request a refund?
response = client.system_one(
    state=ticket_and_policy,
    questions={
        "refund_requested": Noul(instructions="Does the customer ask for a refund?"),
        "team": Choice(instructions="Which team should handle this?",
                       criteria={"billing": ..., "technical": ..., "account": ...}),
        "frustration": Score(...),
    })
if response.answers["team"].confidence < 0.5:
    route_to_human(ticket)

The docs’ design rule matters as much as the API: ask atomic questions, “the kind of judgment a highly knowledgeable person could make in a few seconds,” and combine them in code. Instead of “rate this startup pitch,” ask about market, feasibility and differentiation separately, then weight them yourself. When priorities shift, “change a coefficient in your code rather than rewriting a prompt.”

The names

System One is from Daniel Kahneman’s Thinking, Fast and Slow: fast, intuitive judgment rather than slow, deliberate reasoning. Jev is from William Stanley Jevons, the economist who observed that more efficient steam engines increased coal consumption. That second name is a thesis, and Module 6 takes it seriously.

Takeaways
  • Jev returns typed answers with probabilities, never text
  • Choice, Score and Noul; many questions evaluated in parallel against one state
  • Design rule: atomic judgments in the model, logic in code
2

The argument: the interface is the bottleneck

TypeSafe’s thesis as a chain, and the premise it stands on
By the end of this module you will
  • State TypeSafe’s thesis and its load-bearing premise
  • Explain the “bitterest lesson” and the precedent it rests on
↓
↓
↓
↓

The bitterest lesson

P3 has a name in TypeSafe’s writing. Rich Sutton’s bitter lesson says general methods that use more compute beat clever algorithms. Almeida extends it: “The bitterest lesson in ML is that doing the right task > data > compute > algorithms.” His evidence is the precedent he worked on. At OpenAI, InstructGPT models more than 100× smaller than GPT-3, trained on the right task (following instructions), were preferred to GPT-3 itself. The claim for Jev is the same move again: chat was the right task for assistants, and calibrated decisions are the right task for automation.

Where the weight sits

P1 is widely shared, and P3 has a real precedent. The premise that decides whether Jev matters is P2’s forecast: that the bulk of valuable automation is many small, fast, bounded judgments embedded in code, rather than long open-ended agent tasks. If that is right, a System One model is the natural component. If most value turns out to be in open-ended work, Jev is a very good router for systems whose core is still an LLM.

Takeaways
  • Thesis: intelligence is abundant; a trustworthy machine interface is scarce
  • Method: optimise for the right task (RLCD for calibrated decisions)
  • Load-bearing forecast: automation is mostly many small judgments inside code
3

Calibration and confidence

What the probabilities promise, what they do not, and how to use them
By the end of this module you will
  • Define calibration precisely, including its limit
  • Compute Jev’s confidence from a probability distribution
  • Set thresholds that scale with stakes

Calibration

Across many predictions from a well-calibrated model, outcomes given probability 0.8 happen about 80% of the time, and those given 0.2 about 20%. The docs are exact about the limit: “Calibration is measured across groups of predictions; it does not guarantee that an individual answer is correct.” Calibration is a property of the population of answers. That makes it the right tool for setting thresholds and estimating error rates, and useless as a promise about any one case.

Why it matters for automation

The launch post puts the case in one line: “If a model can do a task 95% of the time but doesn’t say when it’s in the 5%, it can’t automate that task.” LLMs asked for a confidence estimate tend to be “overconfident and inconsistent.” A calibrated model lets code act on the 95% and route the 5%.

Confidence is computed from the probabilities

For Choice and Score, confidence summarises how peaked the distribution is: 1.0 when all probability sits on one option, 0 when it is spread evenly. The docs’ illustration uses (n × largest probability − 1) / (n − 1) for n options. For three options at 90/6/4, that is (3 × 0.9 − 1) / 2 = 0.85. Noul answers have no separate confidence: the probability is the signal.

Thresholds scale with risk

The docs’ banking example routes anything below 0.5 confidence to a human. It shows a balance on any confident read, because the wrong screen is recoverable. It lets a transfer proceed only above 0.9, and even then with confirmation. One model, one confidence number, different thresholds per action, because the cost of being wrong differs. The code, not the model, encodes the risk tolerance.

Takeaways
  • Calibration is a property of many answers, never a guarantee for one
  • Confidence = how peaked the distribution is; it is not the winner’s probability
  • Thresholds are set per action by the cost of being wrong
4

What “can’t hallucinate” guarantees

Type safety is real. Correctness and coherence are separate properties
By the end of this module you will
  • Say exactly which failure the type guarantee removes
  • Recognise the failure modes TypeSafe itself documents
  • Assign each step of a workflow to the right component

The guarantee

The launch post says Jev “can’t hallucinate” and, more precisely, that it “never makes type errors”: possible outputs are defined in advance, so an answer is always one of your options. TypeSafe calls this falsifiable by a single counter-example and “mathematically impossible” to violate, and in its hallucination chart it enters 0% for Jev with the note, “Our number is not empirical. Schema matching is guaranteed.” That is a genuine property with real value. In a system with latency guarantees, a malformed tool call “buried several layers deep in a dependency chain” is a deal-breaker, and Jev cannot produce one.

The distinction to hold

Type safety removes invented values. It does not remove wrong choices. A Choice between “billing” and “technical” can never return “refund-department”, but it can return “billing” when the answer is “technical”. What “can’t hallucinate” buys is that every error is a legal error, one your code can handle, measure and route. That is valuable, and it is not the same as being right.

Calibration is not coherence

TypeSafe’s own jaggedness page gives the clearest example. On the ticket “I was charged twice for the same order,” two Nouls asked “is the customer asking for a refund?” and “is the customer asking for something other than a refund?” The answers were 0.72 and 0.47, summing to 1.19. Each question is evaluated in isolation, so nothing forces logically related answers to agree. The docs’ advice: “don’t hold the model to arithmetic identities between separate questions,” and enforce the identities you need in code.

The documented jagged edges (jev-1.13)

Failure modeTheir fix
Literal reading: “answers the question you wrote, not the one you meant”State the exact condition; put boundary cases in the criteria
Math, counting, numeric representations“Jev is not a calculator.” Keep arithmetic in code
Date and time comparisonExtract components with Choice; compare in code
Indirection and double negativesFewer hops; name the relevant part of the state
Large state full of irrelevant detailFilter first; send only what the question needs
Adversarial content in the statePrecise criteria; test edge cases before deploying
Generation“There are other models for that”

Tool · Who should do this step?

Commit, then see why0 of 7 sorted

A refund workflow. Assign each step to Jev, plain code, a generative LLM, or a human.

Decide whether the customer’s message is asking for a refund
Jev. A bounded semantic judgment, the textbook Noul. This is the docs’ own example.
Check whether the purchase date falls inside the 30-day refund window
Code. Date comparison is a documented weakness and is exact arithmetic anyway. If the date is buried in prose, use Jev to extract it as Choices, then compare in code.
Confirm that two charges on the account have identical amounts and order IDs
Code. It is an exact match on structured fields. “Asking the model something code can compute exactly” is on the docs’ avoid list.
Write a friendly reply explaining the refund decision to the customer
LLM, or a template. Jev does not generate text. System One decides; something else writes.
Score how frustrated the customer is, to set ticket priority
Jev. A Score with defined levels, combined in code with other Scores (the docs’ composite-scoring pattern).
A $4,000 refund where Jev’s “duplicate charge” Choice came back at confidence 0.62
Human. High stakes and medium confidence: the docs’ own pattern routes exactly this case for review. Code decides the routing; a person decides the case.
Filter 400 retrieved policy passages down to the few relevant to this ticket before asking the main questions
Jev. One relevance Noul per passage, in parallel. It is cheap and fast, and it addresses the documented weakness that irrelevant state lowers accuracy.
Takeaways
  • Type safety turns every error into a legal error; it does not make answers correct
  • Separate questions are not forced to agree; enforce identities in code
  • Model for judgment, code for computation, LLM for prose, human for high-stakes doubt
5

The evidence, and how to read it

Speed and price are checkable. “Same intelligence” is measured against a reference, not against ground truth
By the end of this module you will
  • Separate the claims anyone can verify from those that rest on TypeSafe’s evaluation design
  • Explain what “agreement with a reference model” can and cannot show

Claims you can check yourself

70–500 msend-to-end response time (their figure for frontier LLMs: 3–329 s)
$0.042per million input tokens; output tokens free
64ktokens per request (32k for state plus the longest question)
255maximum options in one choice

TypeSafe is careful even here. On speed, its evals “are generally run from our laptops on the West Coast,” where the service is based. On price: “We can’t prove it isn’t subsidized,” though it expects prices to go down, not up.

The bolder claim: similar intelligence

The headline ratios on TypeSafe’s home page, 193.6× faster and 444.6× cheaper, come from its “workflow evals”. The design is specific and worth understanding. There is no ground-truth label set. Each workflow is fixed in code, and the reference answer is the average prediction of GPT-6 Astra and Fable 5.1, the largest and most expensive models. Every model is scored on how closely it matches that reference. TypeSafe’s own caveats:

  • The ratios are “on the higher end of real world gains.”
  • The workflows were not built to flatter Jev and are outside its training distribution, but they were made by TypeSafe’s capabilities team, “so some bias could exist.”
  • Using OpenAI and Anthropic models as the reference biases results toward those labs’ models.
  • The competing LLMs ran through TypeSafe’s structured-decision wrapper, which it finds most accurate but slower and costlier than asking for decisions without probabilities.

TypeSafe also deliberately publishes no public-benchmark results. It argues users should build evals for their own use cases, since “System One tasks are much easier to evaluate.”

What a reference-model eval can show

Matching the average of two frontier models at a fraction of the cost is strong evidence that Jev is a cheap substitute for frontier judgment on those workflows. It cannot show Jev is more accurate than the reference, because disagreeing with the reference counts as error even when the reference is wrong. It cannot show accuracy against real outcomes either. For your own use, the implication is concrete: the eval that matters is one you build with labelled outcomes from your own data, which is what TypeSafe tells you to do.

Takeaways
  • Speed, price and type safety are checkable
  • “Similar intelligence” means agreement with an expensive reference model, not ground truth
  • Build your own labelled eval before trusting thresholds
6

Why it is named after Jevons

Cheap judgment as an economic bet: what gets unlocked, and how long a price lead lasts
By the end of this module you will
  • State the Jevons argument and the condition it needs
  • See what per-decision prices of fractions of a cent make possible
  • Compute how long a price lead lasts against the measured decline in LLM costs

The paradox

In 1865 William Stanley Jevons argued that more efficient steam engines would raise, not lower, Britain’s coal consumption, because cheaper useful work would find many more uses. TypeSafe’s FAQ applies it directly: “We expect machine intelligence to follow a similar path to coal… Every order of magnitude drop in the cost of intelligence unlocks orders of magnitude more use cases.” The paradox holds only when demand is elastic, meaning a price fall produces a proportionally bigger rise in use. That is the empirical question.

What becomes possible

The launch demos show it. A bot plays Doom by asking Jev 10 questions a second, which the team puts at about $7 an hour. An agent wiki-races across Wikipedia, choosing among hundreds of links at each step. Neither is a sensible use of a model that takes seconds per call. The docs’ patterns go further: speculative fan-out (ask every question you might need in one call and let code ignore the rest), map-reduce over big data (turn a corpus into features), and verify everything (score and guardrail every LLM output). These are workloads that only exist when a judgment costs almost nothing.

Tool · How long does a price lead last?

Predict first
Jev’s price advantage today100×
Annual LLM cost decline13×
Annual Jev cost decline1.0×
What this shows

The 13×/yr rate is Epoch’s average for reaching fixed benchmark scores (“The plunging price of thought”, Sep 2026), not a measurement on classification workloads. It is a yardstick, not a forecast.

The arithmetic cuts both ways. A 100× lead is under two years of frontier price decline, and even the 444.6× headline is under two and a half. So a pure price lead is not durable if LLMs keep falling and Jev stands still. But two things are not captured by the price decline. Latency is architectural: parallel evaluation of all questions in one pass does not arrive just because tokens got cheaper. And Jev’s own cost falls too. TypeSafe expects its prices to go down. Set both decline rates equal and the gap never closes. The durable claim is the interface and the architecture, not the price.

Takeaways
  • The name is the thesis: cheaper judgment means far more judgment is used
  • It needs elastic demand, and the demos are evidence of new, cost-gated workloads
  • A price lead alone lasts about two years against frontier declines; latency and the interface are the durable part
7

Objections, and when to use it

Good-faith objections beyond TypeSafe’s own caveats, and a practical decision rule
By the end of this module you will
  • Argue against the System One thesis as well as for it
  • Decide, for a given task, whether a System One model belongs in it
Objection 1

Structured outputs already exist

LLM providers have offered JSON schemas, constrained decoding and log-probabilities for years. A team can get typed answers from an LLM today. What is Jev adding beyond price and speed?

The best reply: constrained decoding guarantees format, but not probabilities trained to be calibrated. TypeSafe’s argument is about the training objective, with RLCD rather than preference, plus parallel evaluation of many questions in one pass. Whether calibration beats a well-prompted LLM with log-probabilities on your task is testable, and worth testing.
Objection 2

The shift of work into the harness is a cost

Decomposing judgments into atomic questions, filtering state, keeping dates and arithmetic in code, and enforcing cross-question identities are all engineering. A system built from dozens of small calls is more legible, but it is also more to design and maintain than one well-prompted LLM call.

The best reply: that is the point of the composability thesis: the logic lives in code where it can be tested and changed. For high-volume, stable decisions the engineering pays back. For one-off or fast-changing tasks, an LLM may be the cheaper total solution.
Objection 3

The evidence is early and self-produced

Every performance number is TypeSafe’s, measured against a reference-model average on workflows its own team wrote, with no public benchmarks by choice. Calibration, the central property, has not yet been independently measured at scale.

The best reply: the evidence is unusually well caveated, the checkable claims check out, and the recommended practice (evaluate on your own labelled data) is the right one regardless of vendor. Early access is the time to run exactly that test.
Objection 4

Frontier labs can close the gap

If fast, calibrated, typed decisions are valuable, the large labs can train for them too, and they already own the distribution. The bitterest lesson cuts both ways: if the right task is what matters, anyone can pick the right task once it has been shown.

The best reply: the same was true of instruction following in 2022, and being first with the right objective still mattered. TypeSafe also says it makes all its own training data and describes itself as “primarily a data research lab.” Whether that is a moat is the open question.
A decision rule

Reach for a System One model when all four hold: the answer is one of a known set of options or levels; the judgment is fast and semantic rather than arithmetic or multi-hop; the decision runs at volume or needs sub-second latency; and you can build a labelled eval to set thresholds. Where any fails, use code, a generative model or a person, or put the System One model in front of them as a router.

Takeaways
  • Structured LLM outputs exist; the claimed difference is the training objective and the architecture
  • Composability moves work into the harness, which pays off at volume
  • Use it where the options are bounded, judgments are fast, volume is high and you can measure