Blog
AI & Machine Learning·8 min read

Jev: The AI Model That Can't Chat (and Why That's the Point)

TypeSafe's Jev doesn't write text. It returns typed answers with calibrated probabilities in ~100ms. Here's what it is, how it works, and where it fits.

Jo Vinkenroye·September 25, 2026
Jev: The AI Model That Can't Chat (and Why That's the Point)

Look at the last agent you built and count the LLM calls. Now count how many of them were really just asking a yes/no question.

"Is this email a refund request?" "Which team owns this ticket?" "Is this shell command destructive?" We send a frontier model a few thousand tokens and wait two seconds. Then we parse a paragraph, or a JSON blob we asked for nicely, and hope it didn't invent a fourth option.

That's the problem Jev is built for. It launched in early access on September 15, and it can't chat at all. That's on purpose.

What Jev Actually Is

Jev is the first model from TypeSafe AI, a San Francisco startup founded in 2024 by Diogo Almeida, Erik Gafni and Sasha Sheng. Almeida spent four years at OpenAI working on RLHF, InstructGPT, ChatGPT and GPT-4. The company spent about two years in stealth and came out with a $40M seed round led by DCVC.

The simplest description I've seen comes from Flavio Copes: Jev is a smart if statement.

You give it some state (text or structured data) and one or more typed questions. It gives you back typed answers with probabilities. It never returns prose, explanations or code, and it can't return a value outside the schema you defined.

TypeSafe calls this category "System One models", named after Kahneman's fast, intuitive System 1. LLMs are the slow, deliberate System 2. Jev makes the snap judgments.

The Three Question Types

Everything Jev does is built on three primitives:

  • Noul: A yes/no question. Returns the probability (0 to 1) that a statement is true. "Does this email request a refund?"
  • Choice: Pick one of up to 255 options. Returns the winner, the full probability distribution and a confidence score. "Which department handles this ticket?"
  • Score: Place something on an ordered scale of 2 to 10 levels. Returns a probability-weighted value, so you can get 1.43 instead of being forced into "low" or "medium". "How severe is this bug?"

That's the whole surface area. It sounds limiting, but most of the "AI decisions" in a production app turn out to be one of these three.

What a Call Looks Like

Here's the JavaScript SDK shape, based on published examples:

const { answers } = await client.systemOne({
state: { ticket: "Export crashes in Safari" },
questions: {
category: choice("What kind of ticket?", {
bug_report: "Something broken",
feature_request: "New feature",
}),
},
});

And here's a real response from a different call that asked two questions about an inbound email:

{
"model": "jev-1.13.0",
"answers": {
"is_sponsor_inquiry": { "type": "noul", "noul": 0.99 },
"product_category": {
"type": "choice",
"choice": "dev_tool",
"probabilities": { "dev_tool": 0.97, "course": 0.01, "unrelated": 0.02 },
"confidence": 0.95
}
},
"usage": { "input_tokens": 210, "output_tokens": 31 }
}

You don't need a JSON repair step or a retry for when the model gets chatty. You also don't need Zod to catch a made-up enum value, because Jev can't make one up.

There's a LangChain integration (langchain-typesafe) and a provider for the Vercel AI SDK too, so it plugs into the stacks most of us already use.

The Numbers (With Some Salt)

TypeSafe's headline claims:

  • Latency: 70 to 500ms end to end, usually around 100ms
  • Speed: 40 to 200x faster than frontier LLMs, peaking at 193.6x
  • Cost: 40 to 400x cheaper, peaking at 444.6x
  • Pricing: $0.042 per million input tokens. Output tokens are free.

To TypeSafe's credit, they say outright that those peak numbers come from workflows their own team built and probably sit at the high end of real-world results. Treat 193.6x as a ceiling, not a promise.

The more interesting claim is how it scales with questions. Jev samples all answers in parallel instead of generating tokens one by one. Adding questions to a call barely changes the latency. In Flavio's testing, 13 questions in one call came out 12.2x cheaper and 10x faster than 13 separate calls.

That changes how you design. With an LLM you ask the minimum. With Jev you ask everything you might need up front, in one shot, and throw away the answers you don't use.

Calibration Is the Real Feature

Speed gets the headlines, but the part I care about is calibration.

Jev is trained with what TypeSafe calls Reinforcement Learning for Calibrated Decisions (RLCD). The goal is that when Jev says 0.9, it's right about 90% of the time. LLM "confidence" is mostly vibes, so this is the difference.

Once probabilities mean something, you can build real policies on them:

  • Above 0.9: act automatically
  • 0.5 to 0.9: ask for confirmation
  • Below 0.5: send it to a human, or to a bigger model

That last option is where it gets good.

Jev + LLM, Not Jev vs LLM

Researchers at CMU published JEV-as-a-Judge, which tests exactly this cascade. Jev judges first. If it's confident, its verdict stands. If not, the question escalates to GPT-6 Astra.

The results:

  • On RewardBench, Jev alone scored 92.2% against GPT-6's 93.5%
  • On the harder JudgeBench, the gap widened to 14.5 points (78.6% vs 93.1%)
  • The cascade kept 99% of GPT-6's accuracy at about 57% of the cost
  • Jev on its own ran at 0.36% of GPT-6's fee

That JudgeBench gap is worth noticing. When the task needs actual reasoning, like checking a derivation, Jev falls well behind. That's not a flaw so much as the design. A System One model is supposed to know when to hand off.

So the right mental model isn't "replace your LLM". Jev takes the high-volume, low-reasoning decisions, and the LLM keeps the hard ones. LangChain frames it the same way:

  • Model routing: Jev judges how hard a request is and picks the model
  • Guardrails: Jev checks an agent's tool calls and blocks risky ones before they run

That second one caught my eye. Every coding agent I use runs shell commands. A 100ms "is this destructive?" check before each one is cheap enough to run on every call, and much faster than asking the main model to second-guess itself.

Where It Falls Short

Jev isn't magic, and the early write-ups are upfront about it:

  • It doesn't generate anything. No replies, no summaries, no code.
  • Math and counting are weak. Keep arithmetic in code where it belongs.
  • Date comparisons are shaky. Pull the dates out with a Choice, then compare them in code.
  • Noisy state hurts. Filter out irrelevant context before you send it.
  • Adversarial text can steer it. Like any classifier, it can be pushed around by input designed to fool it.
  • Thresholds don't transfer. The CMU paper stresses that you have to tune confidence cutoffs per task, on labeled data.

And it's still waitlist-only. You can't just swap it into production today.

The other caveat is about ROI. Flavio makes a good point: measure the whole pipeline, not the model call. If only a small slice of your LLM spend is classification-shaped, a 400x cheaper classifier doesn't move your bill much. And if a plain if already gets it right, keep the if. It costs nothing and it's never wrong.

Why This Matters

For three years the default answer to "I need a bit of intelligence here" has been "call a frontier model". That default is expensive and slow, and it hands back text you have to parse and distrust.

Jev is a bet that the default should split in two. Fast, typed, calibrated decisions go to one kind of model. Slow reasoning and generation go to another. Whether TypeSafe wins that category or not, I think the split itself is right. Most of what my agents do all day is closer to an if statement than to an essay.

Key Takeaways

  • Jev is a decision model, not a language model. It returns typed answers with probabilities and can't produce output outside your schema.
  • Three primitives cover most use cases: Noul (yes/no), Choice (pick one) and Score (rate on a scale).
  • It's fast and cheap: around 100ms and $0.042 per million input tokens, though the 200x/400x figures are best case.
  • Calibration is the killer feature. Trustworthy probabilities let you auto-act, confirm or escalate.
  • Use it alongside LLMs. Put Jev in front for routing, guardrails and triage, and escalate when it's unsure.
  • Know its limits. It's weak on math, dates, adversarial input and anything that needs real reasoning.

If you're on the waitlist, start by finding the LLM calls in your codebase that are secretly yes/no questions. There are probably more than you think.

Stay Updated

Get notified about new posts on automation, productivity tips, indie hacking, and web3.

No spam, ever. Unsubscribe anytime.

Comments

Related Posts