AI × Tools & Tips

Field guides · Jev

A frontier model that won't talk.

Jev can't write a sentence. It answers yes or no, picks from a list, or gives a score — with a probability attached — in the time a chat model takes to start typing. That turns out to be a very big deal, and not only for the cloud.

01 · The same question, twice

Is this email asking for a refund?

One customer message. A frontier chat model on the left, Jev on the right. Both get it right. Watch how they get there.

frontier chat model 0.00 s
jev · system one 0.00 s

What came back

A paragraph vs a number

Your code has to read the paragraph and hope the word “yes” is in it. The number goes straight into an if.

How long

Seconds vs a blink

TypeSafe quotes 70–500 ms end to end. Fast enough to run on every keystroke, not just at the end of a job.

What it cost

Output tokens: free

Input is $0.042 per million tokens. There is no output to bill — it never generates any.

02 · What Jev is

A frontier brain with three buttons.

Jev is TypeSafe AI's first “System One” model, out in early access since 15 September 2026. It reads anything you give it — an email, a form, a block of JSON — and answers only in one of three shapes.

?

Noul · yes or no

“Is this spam?” “Does it mention a competitor?” Back comes one number from 0 to 1: how likely the statement is true.

⋮

Choice · pick one

“Which team should get this?” You give the list. Back comes a pick plus a probability for every option — up to 255 of them.

▁▃▆

Score · how much

“How urgent, from calm to furious?” You describe the rungs. Back comes a position on the ladder, with a spread showing how sure it is.

That's the whole vocabulary. No prose, no code, no “Certainly! Here's a summary.” The name comes from William Stanley Jevons, the 19th‑century economist who noticed that making something cheaper makes people use far more of it. Keep that in mind for beat 05.

Maker
TypeSafe AI, San Francisco
Founder
Diogo Almeida, ex‑OpenAI (RLHF, ChatGPT)
Funding
$40M seed, led by DCVC
Weights
Closed. API only, waitlist.

03 · Why it's interesting

The number means what it says.

Ask a chat model how sure it is and you get vibes. Jev was trained so that when it says 80%, it is right about 80% of the time. That one property is what lets software act on the answer without a human reading it first.

100 emails · “is this a refund request?”

Each dot is one email. Colour is what Jev said; the bucket is how sure it was. A calibrated model fills each bucket to match its label.

So your code can say

if p > 0.90  → refund automatically
if p < 0.10  → route to sales
else         → a person looks (≈ 1 in 6)

You choose how much risk to automate. The model tells you, honestly, which cases it isn't sure about.

Why chat models can't do this

They were trained to be liked (RLHF). A confident tone gets rated up whether or not it's warranted. Jev was trained with RLCD — reinforcement learning for calibrated decisions — where the only reward is a probability that ages well.

04 · The claims, and the honest version

200× is the brochure. 7–25× is the job.

The launch numbers are real but measured on TypeSafe's own workflow evals. Here's what they say next to what people outside the company have found in the first week.

TypeSafe says

  • 40–200×faster than frontier LLMs on System One tasks; 193.6× on their best workflow eval.
  • 40–400×cheaper; 444.6× on that same eval. Input $0.042/M tokens, output free.
  • 70–500 msend to end, against 3–329 s for the models they compared (GPT‑5.6 Terra, GPT‑6 Astra, Fable 5.1).
  • 0%type errors. True by construction — and only about the shape of the answer.

Outside the company

  • 5–18×faster than OpenAI's Luna 5.6 on safety classification, with better accuracy — a Vercel engineer, via TechCrunch.
  • 10–20×cheaper than Gemini for email classification — Bryo AI's CTO.
  • 7–25×the realistic end-to-end gain once a decision sits inside a real pipeline with network and parsing around it.
  • 61.8%on an invoice-processing eval where the frontier reference scored 79.1%. It is not smarter; it is faster at the things it can do.

It's a black box with one number

No reasoning comes back. Simon Willison asked it to rank Bay Area cities and watched Cupertino land on top and East Palo Alto at the bottom. Don't point it at people without checking for that.

“Can't hallucinate” is narrow

It can't produce an option you didn't offer. It can absolutely pick the wrong one, or be 0.9 sure about something false. Calibration is the promise; a probability is not a guarantee.

The yardstick is other models

TypeSafe's eval labels are the averaged answers of GPT‑6 Astra and Fable 5.1 at high reasoning — not human ground truth. Good enough to show speed; not proof of correctness.

05 · What it does to frontier models

The big model stops doing the small jobs.

Look at what an agent actually does across a task. Most of the calls aren't writing. They're deciding: which tool, is this done, does this pass, which bucket. Today every one of those is a frontier call with a paragraph of reasoning you throw away.

One agent run · “triage today's 40 inbound emails”frontier: 0 calls

    1→2

    A second kind of model

    Kahneman's fast System 1 and slow System 2, as an architecture. The frontier model drafts, plans and explains. A decision model routes, scores, gates and checks. Expect every lab to ship one.

    ↓$

    Fewer, harder frontier calls

    The classifier, router and LLM‑as‑judge jobs peel off. What's left for the big model is the work only it can do — and the bill for a workflow falls by more than any price cut has.

    ×∞

    Jevons, on cue

    When a judgment costs a fraction of a cent and 150 ms, you stop rationing them. A check after every step. A score on every row. Far more AI decisions in total, most of them never touching a frontier model.

    06 · Local and private AI

    The first frontier trick that fits on a laptop.

    Jev itself is closed and cloud-only. But the shape of the task — read, then point at an option — is tiny compared with writing. Within a week, open re-implementations appeared that run on a Mac, and the good ones land within a point of Jev on the same tests.

    Kev

    Apache 2.0 · Jared Palmer

    A family of decision models on Qwen3.5 (0.8B, 4B, 9B, 27B). Speaks Jev's API, so TypeSafe SDK code points at it unchanged. Runs on CUDA, ROCm and Apple Silicon.

    Accuracy, new sources
    Kev‑27B 0.848 · Jev 0.857
    Kev‑4B on an M5, 32 GB
    136 ms cached · 721 ms new text
    Kev‑4B on one H100
    ~101 req/s · 18 ms model time
    Confident errors, 9B
    4.0% · Jev 3.7%

    Laya

    Apache 2.0 · Convai Innovations

    A 421M‑parameter encoder (ModernBERT‑large) with a 322M multilingual sibling. Small enough for a CPU box or an air‑gapped rack; fine‑tune notebooks included.

    Latency, local GPU
    7–33 ms
    Banking77 (77 options)
    0.425 · Jev 0.870
    Calibration error
    0.213 · Jev 0.144
    Sweet spot
    ≤ 20 options

    ⌂

    Private data never leaves

    Client inboxes, HR forms, contracts, health notes: classify, route and flag them on the machine they already live on. Only the hard, low‑confidence slice ever goes to a cloud model — and you can strip it first.

    4B

    Small models get a real job

    A 4B model writing prose is a toy. A 4B model answering typed questions is within a few points of the frontier. Deciding is the first task where local hardware is genuinely good enough.

    ⚠

    Small models lie small

    Laya's English model reported 0.945 confidence on Bengali text while scoring 0%. Local calibration only holds inside the distribution you tested. Route languages and domains you haven't checked to “unknown.”

    07 · Where it fits for us

    Every place we currently ask a chat model a yes/no.

    Same shape every time: a judgment on the left, the question type that replaces the prompt on the right.

    Client inboxes

    “Who should see this, and how fast?”

    → choice(team) · score(urgency)

    Runs locally on the mail server. Nothing leaves.

    Site forms

    “Real lead, vendor pitch, or spam?”

    → choice(kind) · noul(spam)

    Fast enough to decide before the thank‑you page.

    Agent pipelines

    “Is this draft actually done?”

    → score(acceptance rungs)

    A check after every step, not one at the end.

    Brand voice

    “Does this copy sound like the client?”

    → score(off‑brand … on‑brand)

    The tone half of the brand linter.

    Moderation

    “Can this comment go live?”

    → noul(safe) · choice(reason)

    Auto‑approve above 0.95, queue the rest.

    Search & CMS

    “Which of these 50 posts answers the question?”

    → score(relevance) × 50, one call

    Reranking for the price of a lookup.

    1. 1

      Find one yes/no

      Somewhere we already ask a chat model a question and parse the answer. That prompt is the spec.

    2. 2

      Run Kev on a Mac

      No waitlist, no data leaving the building. Ask the same question. Log the probability next to what a human said.

    3. 3

      Set the threshold

      A week of pairs tells you where to automate and where to escalate. That number is the whole product.

    Jev's hosted API is behind a waitlist at console.typesafe.ai; Kev is open today. Engineers — for the request shape, the local install, and the exercise.