Jev is a decision model from TypeSafe AI. You give it text and a typed question, and about 200 milliseconds later it returns a yes/no as a probability, a pick from a list you defined, or a score against levels you wrote — each with a confidence number attached. It is not a better ChatGPT. It is a cheap filter you put in front of one.
Jev is the first model in a class TypeSafe AI calls System One models. The name is the whole idea: fast brain versus slow brain. Claude and GPT are the slow brain — they read, reason, then write. Jev is the fast brain. You look at a shirt and simply know it is blue, without deliberating.
That difference is the product. Every frontier model has been racing to be smarter, and the smarter a model gets, the longer it takes on a trivial call. Ask a reasoning model to make ten thousand tiny decisions and you pay for every token it thinks on the way to each one. Jev is built to be called by software, not read by a person, so it answers with a value your code can branch on instead of a paragraph you have to parse.
TypeSafe's own framing is that they took the opposite research direction to RLHF. Instead of training a model to follow human instructions, they trained for what they call RLCD — reinforcement learning for calibrated decisions. The claim on their site is 193.6x faster and 444.6x cheaper than an LLM on their test workflow: $0.000081 per call against $0.013880.
That gap matters more than it sounds. A decision costing a fraction of a cent is one you can afford to make on every single message that lands, without deciding in advance which ones deserve the attention. At the LLM figure on the same workflow, a thousand sorted messages would cost $13.88 — so you would not run it on all of them, and you would be back to triaging by hand. The price is what decides whether the job is worth automating at all.
Everything Jev does is built from three primitives. TypeSafe calls them Noul, Choice and Score.
Choice and Score come back with a confidence value, and that is the part that matters in production. You set a threshold. Above it, the system acts on its own. Below it, the item goes to a human. You cannot do that reliably with a chat model's self-reported certainty, because it is not calibrated.
You can mix all three in a single call, which is the real trick: one request, several questions, answered in parallel.
The short answer: create an account at console.typesafe.ai, take an API key, then POST your text and questions to api.typesafe.ai/v1/systemone with the model set to jev-latest. Because the output is typed, there is no prompt engineering for format and no parsing step after it.
Here is the shape of a real call — this is the inbox-triage example from TypeSafe's own documentation, asking three questions at once:
curl -X POST https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<'EOF'
{
"state": "Hi, I saw your channel and wanted to discuss a paid
partnership for our Series B launch. Can we jump on
a call this week?",
"model": "jev-latest",
"questions": {
"bucket": {
"type": "choice",
"instructions": "What kind of email is this",
"criteria": {
"receipt": "A receipt, invoice or order confirmation",
"brand_deal": "A sponsorship or paid partnership offer",
"scam": "Spam, phishing or a bad-faith pitch"
}
},
"fit": {
"type": "score",
"instructions": "How good a fit this is for a paid content partnership",
"criteria": ["Not a fit at all", "Possible, needs a human look", "Strong fit, reply today"]
},
"is_urgent": {
"type": "noul",
"instructions": "The sender is asking for a response this week"
}
}
}
EOF
The response is already in the shape your code wants, with the spread across options rather than a single word:
{
"answers": {
"bucket": { "type": "choice", "choice": "brand_deal", "confidence": 0.78,
"probabilities": { "brand_deal": 0.85, "scam": 0.15, "receipt": 0.0 } },
"fit": { "type": "score", "score": 2.0, "confidence": 1.0 },
"is_urgent": { "type": "noul", "noul": 1.0 }
}
}
From there it is ordinary code: if the bucket is a strong fit and the confidence clears your threshold, route it to the right person today. If it is below the threshold, queue it for a human. Nothing needs to be summarised, reformatted or re-asked.
The use case is any decision you make thousands of times, where the answer comes from a short list of options and getting it wrong is cheap because a human checks anyway.
The pattern behind all four: use Jev to make the cheap decision, then spend an expensive model call only on the items that earned it. That is the cost argument, and it is the same argument as putting a receptionist in front of a specialist.
Take an electrical contractor with two office staff and one shared inbox. Every day brings quote requests, existing customers chasing a date, supplier invoices, subcontractor paperwork, and a steady trickle of people asking whether a socket can be moved. Two people read all of it in the order it arrived, which means the four-thousand-pound rewire sits behind three invoices.
The Jev version of that business sends each message through one call with three questions: what is this, how good a fit is it, and does it want a reply today. Quote requests that score as a strong fit drop into the quoting list and get flagged. Invoices are filed. Everything the model is unsure about lands in one folder a human clears in ten minutes.
Nothing about that is clever. It is the sorting anyone would do, done consistently, on every message, without the two people having to choose what to read first. That is usually where the hours are — not in writing the replies, but in deciding which ones to write.
| Jev | A frontier LLM | |
|---|---|---|
| Output | Typed decision: yes/no, choice, score | Prose you have to parse or constrain |
| Speed | About 200ms | Seconds, more if it reasons first |
| Published cost | $0.000081 per call (TypeSafe's figure) | $0.013880 per call on the same workflow |
| Confidence | Calibrated probability you can threshold on | Self-reported, not reliable for routing |
| Can write a reply | No | Yes |
| Can reason through steps | No — that is not what it is for | Yes |
Read the table as a division of labour rather than a contest. Jev sorts; the LLM writes and reasons. Used together, the expensive model only ever sees traffic that already passed the filter.
No, and it is not trying to be. It cannot write, explain or reason through steps. It replaces the part of your stack where you were paying a reasoning model to make a small decision.
Want to see a cheap filter and an expensive writer working together in a real business? That is what an AI agent suite is — one system, several models, each doing the job it is good at.
TypeSafe's own FAQ addresses this directly and the honest answer here is that we could not verify it from outside. Assume it is probabilistic and write your thresholds accordingly — never assume the same input returns the same answer twice. Anything that must be exactly repeatable belongs in your own code, not in the model.
You need someone who can make an API call. If you have an agent or automation platform already, that is usually enough. If you do not, buying the finished automation is cheaper than learning to build it: our AI phone and booking agents are the same idea — cheap automatic decisions, human handoff when it matters. You can see what we build and run in the store.
Point it at one messy pile of text you already have and one question you already answer by hand. A shared inbox is the usual first win, and it is measurable within a week — often the same week it starts paying for itself. If you would rather not build it yourself, we run the finished version for UK businesses; it is in the store.
No — they do different jobs. A phone agent answers, speaks and books. Jev labels and routes text cheaply in the background. Plenty of businesses run both, and both are covered on the store page if you would rather not assemble it yourself.
At TypeSafe's published rate, a thousand triage decisions a month costs about eight cents in model spend — which is the whole reason the technique is worth considering. The meaningful cost is the setup and the human review time, not the calls. Our Q4 install window covers that part if you want it done for you.
Everything above describes a filter you have to assemble. We build and run the finished version for UK businesses — an agent that answers, qualifies and routes without anyone managing a stack: calls and enquiries handled 24/7, the strong ones flagged, the rest triaged quietly in the background.
One email: the decision layer we use to sort a business's inbox — the exact questions to ask, the thresholds to set and what to do with the ones it gets wrong. No spam, unsubscribe any time.
Honest note: AI Suite is not affiliated with TypeSafe AI and this is not a paid promotion — we earn nothing if you sign up to Jev. Speed, cost and accuracy figures are TypeSafe's published claims (plus one independent measurement, which was far more conservative than the headline claims) and are attributed as such, not measured by us. Public pricing was unavailable at the time of writing. Model names, endpoints and terms change quickly on new releases; TypeSafe's own documentation is the authority. Where we mention work we sell, it is labelled as ours.