Field guides · Jev
A frontier model that won't talk.
Jev can't write a sentence. It answers yes or no, picks from a list, or gives a score — with a probability attached — in the time a chat model takes to start typing. That turns out to be a very big deal, and not only for the cloud.
01 · The same question, twice
Is this email asking for a refund?
One customer message. A frontier chat model on the left, Jev on the right. Both get it right. Watch how they get there.
What came back
A paragraph vs a number
Your code has to read the paragraph and hope the word “yes” is in it. The number goes straight into an if.
How long
Seconds vs a blink
TypeSafe quotes 70–500 ms end to end. Fast enough to run on every keystroke, not just at the end of a job.
What it cost
Output tokens: free
Input is $0.042 per million tokens. There is no output to bill — it never generates any.
The wiring
The chat model is paying an autoregressive tax: every word of the explanation is a forward pass, and the word you actually needed arrives last. Jev skips generation entirely. You post a state (any text or JSON) and a set of typed questions; it returns a probability per question in one shot. Kev, the open re-implementation below, speaks the same /v1/systemone contract, so this is what the request looks like against either.
{
"state": "Shoes arrived two weeks late and in the wrong size. I want my money back.",
"questions": {
"refund": { "type": "noul", "instructions": "Does the message explicitly request a refund?" },
"department": { "type": "choice", "instructions": "Which team owns this?",
"criteria": { "returns": "Refunds and returns", "shipping": "Delivery problems", "other": "None of the above" } },
"frustration": { "type": "score", "instructions": "How upset is the customer?",
"criteria": ["Calm", "Frustrated", "Very angry"] }
}
}
// → answers.refund.noul 0.97
// → answers.department.choice "returns" probabilities { returns .81, shipping .16, other .03 }
// → answers.frustration.score 1.44 legend { 0 Calm, 1 Frustrated, 2 Very angry }
// → usage.output_tokens 0 billed latency_ms ~150
02 · What Jev is
A frontier brain with three buttons.
Jev is TypeSafe AI's first “System One” model, out in early access since 15 September 2026. It reads anything you give it — an email, a form, a block of JSON — and answers only in one of three shapes.
?
Noul · yes or no
“Is this spam?” “Does it mention a competitor?” Back comes one number from 0 to 1: how likely the statement is true.
⋮
Choice · pick one
“Which team should get this?” You give the list. Back comes a pick plus a probability for every option — up to 255 of them.
▁▃▆
Score · how much
“How urgent, from calm to furious?” You describe the rungs. Back comes a position on the ladder, with a spread showing how sure it is.
That's the whole vocabulary. No prose, no code, no “Certainly! Here's a summary.” The name comes from William Stanley Jevons, the 19th‑century economist who noticed that making something cheaper makes people use far more of it. Keep that in mind for beat 05.
- Maker
- TypeSafe AI, San Francisco
- Founder
- Diogo Almeida, ex‑OpenAI (RLHF, ChatGPT)
- Funding
- $40M seed, led by DCVC
- Weights
- Closed. API only, waitlist.
The wiring
Think of it as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out. Three properties matter for how you build on it. Type safety is structural — the output space is the criteria you passed, so a malformed answer is impossible (a wrong answer is not; see beat 04). Questions are independent — each one sees the state and its own instructions only, so forty questions cost about the same wall-clock as one. Confidence is separate from probability — a choice at 0.47 with the runner‑up at 0.28 reports low confidence, and that gap is the signal you threshold on.
llm -m jev 'Please refund my last payment.' \
-s 'Does this message explicitly request a refund?'
# → 0.97
03 · Why it's interesting
The number means what it says.
Ask a chat model how sure it is and you get vibes. Jev was trained so that when it says 80%, it is right about 80% of the time. That one property is what lets software act on the answer without a human reading it first.
100 emails · “is this a refund request?”
Each dot is one email. Colour is what Jev said; the bucket is how sure it was. A calibrated model fills each bucket to match its label.
So your code can say
if p > 0.90 → refund automatically if p < 0.10 → route to sales else → a person looks (≈ 1 in 6)
You choose how much risk to automate. The model tells you, honestly, which cases it isn't sure about.
Why chat models can't do this
They were trained to be liked (RLHF). A confident tone gets rated up whether or not it's warranted. Jev was trained with RLCD — reinforcement learning for calibrated decisions — where the only reward is a probability that ages well.
The wiring
RLCD swaps the RLHF reward (“did a rater prefer this?”) for a proper-scoring-rule style objective on synthetic tasks with known answers: the policy is rewarded for probabilities that are both accurate and honest about their own uncertainty. TypeSafe hasn't published the exact reward, so treat the mechanism as described, not verified. Measure it yourself: log (p, outcome) pairs for a week, then compute Brier score and a reliability curve. Kev's authors report Jev's “confident errors” — cases over 0.9 that were wrong — at about 3.7% on their held-out sources; that's the kind of number to hold your own thresholds against.
Two traps. Calibration doesn't compose: three well-calibrated decisions chained through thresholds and branches don't give you a well-calibrated workflow, so measure at the workflow's output too. And a choice with no escape hatch forces probability into the wrong bucket — always include an "other" / "unknown" option.
04 · The claims, and the honest version
200× is the brochure. 7–25× is the job.
The launch numbers are real but measured on TypeSafe's own workflow evals. Here's what they say next to what people outside the company have found in the first week.
TypeSafe says
- 40–200×faster than frontier LLMs on System One tasks; 193.6× on their best workflow eval.
- 40–400×cheaper; 444.6× on that same eval. Input $0.042/M tokens, output free.
- 70–500 msend to end, against 3–329 s for the models they compared (GPT‑5.6 Terra, GPT‑6 Astra, Fable 5.1).
- 0%type errors. True by construction — and only about the shape of the answer.
Outside the company
- 5–18×faster than OpenAI's Luna 5.6 on safety classification, with better accuracy — a Vercel engineer, via TechCrunch.
- 10–20×cheaper than Gemini for email classification — Bryo AI's CTO.
- 7–25×the realistic end-to-end gain once a decision sits inside a real pipeline with network and parsing around it.
- 61.8%on an invoice-processing eval where the frontier reference scored 79.1%. It is not smarter; it is faster at the things it can do.
It's a black box with one number
No reasoning comes back. Simon Willison asked it to rank Bay Area cities and watched Cupertino land on top and East Palo Alto at the bottom. Don't point it at people without checking for that.
“Can't hallucinate” is narrow
It can't produce an option you didn't offer. It can absolutely pick the wrong one, or be 0.9 sure about something false. Calibration is the promise; a probability is not a guarantee.
The yardstick is other models
TypeSafe's eval labels are the averaged answers of GPT‑6 Astra and Fable 5.1 at high reasoning — not human ground truth. Good enough to show speed; not proof of correctness.
The wiring
Where the speed actually comes from: no decoding loop, so latency is one prefill; a parallel sampler that answers every question in the same pass; and an attention mask that lets the state be encoded once and reused across questions. The cost story follows — you pay for prefill only, which is why output is free rather than discounted. The corollary is the limitation: anything that needs a generated answer (an extracted name, a rewritten sentence, a number outside a fixed ladder) is out of scope, and the public community stunts — jevchat, jev‑leftpad, jev‑2048 — work by turning generation into thousands of tiny choices, which is funny and not a plan.
05 · What it does to frontier models
The big model stops doing the small jobs.
Look at what an agent actually does across a task. Most of the calls aren't writing. They're deciding: which tool, is this done, does this pass, which bucket. Today every one of those is a frontier call with a paragraph of reasoning you throw away.
1→2
A second kind of model
Kahneman's fast System 1 and slow System 2, as an architecture. The frontier model drafts, plans and explains. A decision model routes, scores, gates and checks. Expect every lab to ship one.
↓$
Fewer, harder frontier calls
The classifier, router and LLM‑as‑judge jobs peel off. What's left for the big model is the work only it can do — and the bill for a workflow falls by more than any price cut has.
×∞
Jevons, on cue
When a judgment costs a fraction of a cent and 150 ms, you stop rationing them. A check after every step. A score on every row. Far more AI decisions in total, most of them never touching a frontier model.
The wiring
Concretely, in an agent loop: tool routing becomes a choice over the tool list; guardrails become a noul on the draft before it ships (“does this reveal a client name?”); done‑checks become a score against your acceptance rungs; LLM‑as‑judge in evals becomes a calibrated probability instead of a 1–10 that means nothing. The pattern is DSPy‑style typed signatures with a model that actually returns the type. Keep the frontier model for the draft and the plan; hand it back the low‑confidence cases only. That escalation — decide fast, escalate the uncertain slice — is the design, and the confidence field is what makes it work.
06 · Local and private AI
The first frontier trick that fits on a laptop.
Jev itself is closed and cloud-only. But the shape of the task — read, then point at an option — is tiny compared with writing. Within a week, open re-implementations appeared that run on a Mac, and the good ones land within a point of Jev on the same tests.
Kev
Apache 2.0 · Jared Palmer
A family of decision models on Qwen3.5 (0.8B, 4B, 9B, 27B). Speaks Jev's API, so TypeSafe SDK code points at it unchanged. Runs on CUDA, ROCm and Apple Silicon.
- Accuracy, new sources
- Kev‑27B 0.848 · Jev 0.857
- Kev‑4B on an M5, 32 GB
- 136 ms cached · 721 ms new text
- Kev‑4B on one H100
- ~101 req/s · 18 ms model time
- Confident errors, 9B
- 4.0% · Jev 3.7%
Laya
Apache 2.0 · Convai Innovations
A 421M‑parameter encoder (ModernBERT‑large) with a 322M multilingual sibling. Small enough for a CPU box or an air‑gapped rack; fine‑tune notebooks included.
- Latency, local GPU
- 7–33 ms
- Banking77 (77 options)
- 0.425 · Jev 0.870
- Calibration error
- 0.213 · Jev 0.144
- Sweet spot
- ≤ 20 options
⌂
Private data never leaves
Client inboxes, HR forms, contracts, health notes: classify, route and flag them on the machine they already live on. Only the hard, low‑confidence slice ever goes to a cloud model — and you can strip it first.
4B
Small models get a real job
A 4B model writing prose is a toy. A 4B model answering typed questions is within a few points of the frontier. Deciding is the first task where local hardware is genuinely good enough.
⚠
Small models lie small
Laya's English model reported 0.945 confidence on Bengali text while scoring 0%. Local calibration only holds inside the distribution you tested. Route languages and domains you haven't checked to “unknown.”
The wiring
Why this fits locally when generation doesn't: Kev is a stock Qwen with a rank‑16 LoRA and a pointer head that scores each option's </opt> token against the question's <decide> token — one prefill, no decode loop, KV cache shared across questions. Memory is the base model; compute is a single pass. A fitted temperature (2.41 for Kev‑4B) does the calibration at serve time. That's also the recipe if you want to train your own on a private label set.
git clone https://github.com/jaredpalmer/kev.git && cd kev
uv sync --extra serve
uv run --extra serve python -m kev.serve --run jaredpalmer/kev-4b --port 8009
# then point the TypeSafe SDK (or curl) at http://localhost:8009/v1/systemone
# 32 GB Apple Silicon handles kev-4b and kev-9b; kev-0.8b runs on less.
07 · Where it fits for us
Every place we currently ask a chat model a yes/no.
Same shape every time: a judgment on the left, the question type that replaces the prompt on the right.
Client inboxes
“Who should see this, and how fast?”
→ choice(team) · score(urgency)
Runs locally on the mail server. Nothing leaves.
Site forms
“Real lead, vendor pitch, or spam?”
→ choice(kind) · noul(spam)
Fast enough to decide before the thank‑you page.
Agent pipelines
“Is this draft actually done?”
→ score(acceptance rungs)
A check after every step, not one at the end.
Brand voice
“Does this copy sound like the client?”
→ score(off‑brand … on‑brand)
The tone half of the brand linter.
Moderation
“Can this comment go live?”
→ noul(safe) · choice(reason)
Auto‑approve above 0.95, queue the rest.
Search & CMS
“Which of these 50 posts answers the question?”
→ score(relevance) × 50, one call
Reranking for the price of a lookup.
-
1
Find one yes/no
Somewhere we already ask a chat model a question and parse the answer. That prompt is the spec.
-
2
Run Kev on a Mac
No waitlist, no data leaving the building. Ask the same question. Log the probability next to what a human said.
-
3
Set the threshold
A week of pairs tells you where to automate and where to escalate. That number is the whole product.
Jev's hosted API is behind a waitlist at console.typesafe.ai; Kev is open today. Engineers — for the request shape, the local install, and the exercise.
Exercise · 60 minutes, pairs
- Pick a real decision we make with an LLM today (a classifier prompt, an LLM‑as‑judge in an eval, a routing step in an agent). Pull 50 past inputs and the answer a human would give.
- Stand up Kev‑4B locally (beat 06). Write the question as noul, choice or score. Include an "other" option. Keep the instructions to one sentence.
- Run all 50 in one request each. Plot probability against the human label. Where does 0.9 actually sit?
- Pick thresholds: auto above, escalate the middle to the frontier model, auto below. Count how many frontier calls you just deleted.
- Stretch: try the same 50 against Jev's API if you're off the waitlist, and against Kev‑0.8B. Where does the small one fall apart? That boundary is your deployment note.
Sources
- TypeSafe AI — Introducing System One Models & Jev (claims, pricing, RLCD)
- Simon Willison — Jev introduces a new shape of LLM (API walkthrough, bias example)
- TechCrunch — A new kind of AI model from a ChatGPT inventor (founder, developer numbers)
- Anthony Maio — Jev: the language model that won't talk (eval caveats, composition)
- Kev on GitHub · Jev vs Laya (local numbers)