Jev costs $0.042 per million input tokens. Output tokens are free.
That is not a promotion. It is free because Jev does not write anything. You send it your application state and a typed question, and it returns one of your own options plus a probability. No prose, no JSON to parse, no tokens to pay for.
In one benchmark it classified 5.43 times faster than Claude Haiku, at 96% lower cost, and it was more accurate: 95.00% against 73.75%.
Here is what it is, when it beats an LLM, when it does not, and how to call it on OpenRouter today.
What Jev actually is
Jev is a decision model built by TypeSafe. It went into early access on 15 September 2026 and the current version is Jev 1.13.
TypeSafe calls the category a System One model, borrowing from Daniel Kahneman's Thinking, Fast and Slow. System One is the fast, automatic, pattern-matching kind of thinking. You do not reason your way to recognising a face or noticing that an email sounds angry. You just know.
That is the job Jev does. It takes a piece of state, and it returns a decision.
The word that matters is non-generative. It produces no text tokens at all, which is why output is free. It is not a language model with structured output bolted on, it is a different kind of model that only ever emits an answer from a set you defined.
A language model hides its uncertainty inside prose. Jev hands you the uncertainty as a number you can put an if-statement around.
What you send and what comes back
You send two things:
state: the text to evaluate. Strings, JSON objects or arrays of text. No images, audio or video.questions: one or more typed questions, with instructions and criteria.
You get back one of three primitives:
| Primitive | What it does | What it returns |
|---|---|---|
| Choice | Picks from your options | The selected option, a probability for every option, and a confidence value |
| Score | Rates on ordered levels, up to 0 to 10 | A probability-weighted number, probabilities per level, and confidence |
| Noul | Yes or no | The probability of yes, with no separate confidence field |
The part worth dwelling on is "a probability for every option". You do not simply learn that the ticket is urgent. You learn it is urgent at 0.87, normal at 0.11 and spam at 0.02. That distribution is something your code can threshold on, route with, or escalate to a human when it is too flat to trust.
And because the answer is always one of your options, there is no category to hallucinate. An LLM asked to classify can invent a label, misspell one, or wrap it in an apology. Jev structurally cannot.
The numbers
Independent benchmarking against Claude Haiku as a classifier:
| Jev | Claude Haiku | |
|---|---|---|
| Median latency | 126.81 ms | 688.40 ms |
| Speed | 5.43x faster | baseline |
| Classifier cost | 96% lower | baseline |
| Matched expected classification | 95.00% | 73.75% |
Faster, cheaper and more accurate at the same job is an unusual result, and it is worth understanding why it is possible. Haiku is a general model being asked to do a narrow task, so it carries the weight of everything else it can do. Jev only does this.
On support ticket triage, Jev completed the work at 2.5 cents per 1,000 tickets with a median of 194 milliseconds.
One honest counterexample. On LLMRouterBench with a 13-model pool, a Jev-difficulty plus retrieval router scored 62.4% against 60.3% for the best single model, at lower cost. A real improvement, but the researchers noted the gain came from the retrieval evidence rather than from Jev's difficulty signal. Jev is not magic at every task you point it at, and that write-up deserves credit for saying so.

What it costs, in real money
Input is $0.042 per million tokens. Output is $0.
TypeSafe's own worked example: a three-question call about a support ticket used 447 input tokens and cost $0.000019. That is about two thousandths of a cent.
A million tickets of that size costs roughly $19.
Run that against what a frontier model would charge to read the same million tickets and produce a structured answer for each, and the gap is the entire reason this category exists. If you are classifying, routing or gating at volume, the cost difference is not a percentage, it is an order of magnitude.
How to use it on OpenRouter
This is the practical part, and the good news is that it needs no new account.
Access. Jev is available to anyone with an OpenRouter API key. No waitlist, no separate TypeSafe signup.
Model IDs:
typesafe/jev-1.13for a pinned version~typesafe/jev-latestfor the latest alias
Pin the version in production. The alias moves, and a model that silently changes behaviour underneath a routing rule is a bad afternoon.
Endpoints. Two surfaces exist:
POST https://openrouter.ai/api/alpha/decisionsPOST https://openrouter.ai/api/v1/systemone
Context. 32,000 tokens for your state, and 64k tokens per request in total, which is the state plus the longest question. Budget for both halves rather than just the state.
SDKs. The OpenRouter TypeScript SDK works directly. TypeSafe also ships JavaScript and Python SDKs, and those can be pointed at OpenRouter with a one-line base URL change, so you are not locked into either route.
Costs come back in the response. Each response carries usage.cost in USD, so you can log actual spend per decision rather than estimating from token counts.
What Jev cannot do
This list is as important as the benchmarks, because most disappointment with it will come from expecting the wrong thing.
Jev does not give you:
- Text generation or explanations. It decides, it does not tell you why in words.
- Visible reasoning. No chain of thought to inspect.
- Tool calls, conversation or multi-step plans. It is one decision, not an agent.
- Exact arithmetic, date maths or threshold comparisons. Do these in your own code. It is a pattern matcher, and asking it to compute is asking the wrong organ.
- Images, audio or video. Text and structured text only.
It is also proprietary. No published weights, no paper. You are trusting a vendor's benchmarks and your own testing, which is a real consideration for anything you would struggle to replace.
When to use which
The honest framing from TypeSafe's own documentation is that these work together rather than competing: Jev routes and verifies, the LLM provides the language.
Reach for Jev when:
- Ticket routing and triage
- Classification and tagging at scale
- Gating an agent before a destructive action
- Checking an LLM's output against a policy before it ships
- Ranking or filtering candidates
Reach for an LLM when: you need the words. Replies, summaries, code, anything where prose is the product.
The gating use case deserves a mention on its own. If you are building agents, a fast and cheap check between "the agent wants to delete this" and actually deleting it is worth far more than its cost, and a probability you can threshold on is a much better gate than parsing an LLM's opinion out of a sentence.
Common mistakes
Reading confidence as correctness. This is the one that will catch people. Confidence measures how concentrated the probability distribution is, not how likely the answer is to be right. A confidently wrong answer is entirely possible and the number will look reassuring.
Trusting calibration on a single call. When Jev says 0.8, it is right about 80% of the time averaged across many answers. It says nothing about the specific answer in front of you.
Using vendor thresholds. TypeSafe's own advice is to pick your thresholds from your own labelled data. Your definition of "urgent" is not theirs.
Asking it to do arithmetic. It cannot do exact maths, date calculations or threshold comparisons. Those belong in your code, where they are also free and always correct.
Expecting it to replace your LLM. It has no language output at all. Almost every real system uses both.
Using the latest alias in production. Pin to typesafe/jev-1.13 so your routing logic does not change under you.
Key takeaways
- Jev is a non-generative decision model from TypeSafe, in early access since 15 September 2026, currently version 1.13. Output tokens are free because it emits no text.
- Input is $0.042 per million tokens. A three-question ticket call of 447 tokens cost $0.000019, so a million such tickets is about $19.
- Benchmarked against Claude Haiku as a classifier it was 5.43x faster (126.81 ms against 688.40 ms), 96% cheaper, and more accurate at 95.00% against 73.75%.
- It returns three primitives: Choice, Score and Noul, each with a full probability distribution you can threshold on in code.
- The answer is always one of your predefined options, so a hallucinated category is structurally impossible.
- Confidence measures how concentrated the distribution is, not whether the answer is correct. Calibration holds across many answers, not on any single one.
- It cannot generate text, show reasoning, call tools, do exact arithmetic, or read images, audio or video. It is proprietary with no published weights or paper.
- On OpenRouter use
typesafe/jev-1.13, via/api/alpha/decisionsor/api/v1/systemone, with 32k for state and 64k per request total. An OpenRouter key is all you need.
The gating use case is the one we find most immediately useful, because a cheap reliable check before an agent does something irreversible is worth more than it costs. That sits right next to the reliability testing problem, where the lesson is the same: verify the state, do not trust the model's own account of what it did. For the layer around all of this, we compared the agent frameworks, and if you want decisions like these wired into your own systems, that is the work we do.




