15 min read
What Is Jev AI? A Guide to the New AI Decision Model
What Jev AI is, how it differs from a large language model, how it did in independent tests, and what a UK business should check before using it.

In this article
In my line of work, there are few things I find quite as satisfying as automating a complex workflow, taking a task that people find laborious, time-consuming and complicated, and turning it into a simple system that runs itself with next to no supervision, freeing up people's time to get on with the important things. So when a new kind of AI model becomes available, adding another tool to the automation toolkit, I am always keen to learn more.
TypeSafe AI says its new model, Jev, is 194 times faster and 445 times cheaper than large language models, in its own workflow tests. Independent tests agree it is cheaper and faster than the language models it was tested against, which cost anything from about 1.6 to 580 times as much (Good Start Labs, Forbes). Within a fortnight, dozens of similar models had appeared and OpenAI had announced a rival service. Judging by the traction Jev has been getting on social media since its launch, it's fair to say that the hype train has fully left the station (Bloomberg). Let's have a look at whether Jev lives up to the claims, and what a UK business should check before using it.
What is Jev?
Jev is an AI model built to make fast decisions that software can use directly, rather than to write text (TypeSafe docs). Large language models, like the ones behind ChatGPT, "are designed to produce text for humans to read", so a workflow has to dig the decision out of their reply (TypeSafe docs). Jev's answers come back ready to use, each with a probability, and TypeSafe says it built Jev "entirely focused on automation". It "is neither small nor an LLM", in TypeSafe's words (TypeSafe).
Jev launched on 15 September 2026. Its maker, TypeSafe, is a San Francisco start-up led by former OpenAI researcher Diogo Almeida, with $40 million of seed funding (SiliconANGLE). It is named after William Stanley Jevons, and TypeSafe calls it a "System One" model, after the fast thinking in Daniel Kahneman's Thinking, Fast and Slow (TypeSafe). TypeSafe has not published how it is built (Opper).
What Jev does, and what it cannot do
Jev works like a form. You give it some text, such as a support ticket or an invoice, and questions set in advance. It picks an option from a list, gives a score, or says how likely something is to be true, and every answer comes with a percentage (TypeSafe docs). That matters, because research suggests language models "tend to be overconfident, potentially imitating human patterns of expressing confidence" (Xiong et al.). Reasoning models do better (Yoon et al.), and asking several times helps (Xiong et al.), but each extra call costs time and money. The evidence on Jev's own percentages is mixed. Opper, which also hosts a rival model, found them more reliable than that rival's (Opper), but in a phishing test, asked a single question, they were less reliable than Claude Haiku's (jev-phishing-bench).
Here is a real example from Cloudflare's documentation, for the message "Help! My payouts have been failing for 3 days." (Cloudflare)

Because Jev picks from a list rather than writing, TypeSafe says it "can't hallucinate" (TypeSafe). That is true only in a narrow sense. Given billing, technical and sales, it cannot invent a fourth option, but it can still pick the wrong one. And it never explains itself. As one reviewer put it, "The biggest loss is the explanation" (O'Reilly).
Speed and cost in independent tests
Even TypeSafe says its headline figures are likely to be "on the higher end of real world gains" (TypeSafe). Independent testers found Jev cheaper than every language model it was tested against, and about 3 to 25 times faster against named models.
Cost per 1,000 decisions in independent tests
*Jev is shown once, at its highest cost, $0.16 in Good Start Labs' rubric test against five language models, of which DeepSeek was the cheapest and Claude Fable the dearest. It cost $0.043 in Near Here's event listing test against Gemini and Mistral, and $0.038 in the test emails against Claude Haiku. Each test used a different task. Near Here ran Gemini and Mistral on high reasoning settings and had them write a short explanation, which Jev did not. Good Start Labs costed the language models from its bills and Jev at TypeSafe's list price, per million checks, divided here by 1,000. Dollars are US dollars.
Time per decision in independent tests
*Jev is shown once, at its slowest, 0.59 seconds on average in Near Here's event listing test against Gemini and Mistral. It took a median of 0.239 seconds in the test emails against Claude Haiku's median of 0.687, measured from France and reported by The D AI LY Brief, a median of 0.35 seconds in a 12-passage review test reported by Forbes against Claude Fable 5.1 at high effort, and 0.087 seconds per decision in an n8n user's test against a language model they did not name. Each test used a different task.
Each test used a different task, so the fair comparison is within each test, and the saving depends mostly on the rival. In a test by Good Start Labs, which had early access and used a July version of Jev, DeepSeek V4.1 Flash cost about 1.6 times as much as Jev and Claude Fable 5.1 about 206 times as much (Good Start Labs). In a 12-passage review test reported by Forbes, Fable cost about 580 times as much (Forbes).
How accurate is Jev?
In one test, how the question was asked mattered most. Asked once whether emails were phishing, Jev did worse than Claude Haiku 4.5. Split into five narrower questions and tuned on labelled emails, it caught up, although a simple hand-written rule with no AI did almost as well (jev-phishing-bench).
Accuracy on test emails, half of them phishing
One-question scores are on all 2,000 emails, 1,000 of them phishing. The five-question scores and the hand-written rule were measured on 1,000 held-back emails, after the answers were weighted using the other 1,000. The tester found the gap between Jev and Haiku at five questions too small to be meaningful.
Other tests are mixed. Jev beat Gemini and Mistral on the main set of listings a UK events website used to tune its prompts, but Gemini edged ahead on fresh ones (Near Here). Jev agreed with five language models a little less often than they agreed with each other (Good Start Labs), yet scored about 96% with no training examples in a pre-registered test, against about 77% for hand-written keywords (PriorBench) and 95% to 97.5% on fresh data in Opper's tests (Opper).
Jev also cannot say "I don't know". It must pick one of your answers, even when none fit. One tester saw a cake recipe filed as a technical issue at 94% confidence (PriorBench). Giving it a way out helps. With an "other" option, Jev flagged 21 of 30 off-topic messages, against none without one (PriorBench).
Similar models
Jev is not the first AI built to pick an answer rather than write one. Researchers were benchmarking the idea in 2019 (Yin et al.). What followed the launch was quick. On the morning of 27 September a community leaderboard listed 67 similar models, 64 of them released since Jev, although 15 were new ways of running existing models (Jev Decision Index). On 29 September OpenAI announced its own Decisions API, built on its GPT-6 Luna model (OpenAI), which The New Stack called "likely a reaction to TypeSafe and Jev" (The New Stack). It opened as a public beta on 6 October (OpenAI). By 7 October Cloudflare also listed decision models of its own (Cloudflare), and the leaderboard listed 112 models, with an open model from Perplexity scoring 62.75 against Jev's 60.11 (Jev Decision Index).
Decision models before and after Jev
2019
Researchers benchmark AI that labels text it was never trained on
2023
GLiNER, a small model that picks out items in text in one pass
15 September 2026
TypeSafe launches Jev
16 September 2026
11 Jev-style models from at least 8 projects appear in a single day
17 September 2026
A community tracker of Jev-style models opens
22 September 2026
The tracker adds a scored leaderboard
27 September 2026
67 models on the leaderboard, 64 of them released since Jev
29 September 2026
OpenAI announces a Decisions API
6 October 2026
OpenAI opens its Decisions API as a public beta
7 October 2026Where it stands
112 models on the leaderboard, one from Perplexity scoring above Jev
Model dates are when each model's Hugging Face page was created, or its GitHub page where it has none, which can be before it was announced. Only models entered on the Jev Decision Index are counted, and Jev itself is not included. The count was 67 on the morning of 27 September 2026, and 15 of those 67 were new ways of running existing models rather than new models. It was 112 on 7 October 2026.
Practical advice for UK businesses
None of this is legal advice.
Test it first
Try Jev on your own data before switching anything on. TypeSafe suggests acting automatically only on high-confidence answers, starting with "conservative thresholds" (TypeSafe docs). Jev's answers can shift with the order the options are listed in (Archestra), or with one planted entry in the input (VentureBeat). Fix the model version too, because TypeSafe's default version "moves when a new release ships" (TypeSafe docs).
Where your data goes
Hosting matters if you send Jev personal data. TypeSafe's privacy policy says "The Services are hosted in the United States" (TypeSafe privacy policy). TypeSafe is not on the Data Privacy Framework list (Data Privacy Framework), so a UK firm would rely on the UK Addendum in TypeSafe's data processing terms (TypeSafe DPA). The ICO says a firm relying on a safeguard like this must also complete a "transfer risk assessment" (ICO).
The law on automated decisions
Since 5 February 2026, when a significant decision about a person is made solely by automation, the organisation must tell that person about it and let them make their case, get a human review and contest it (UK GDPR Article 22C). The ICO's draft guidance says organisations "must provide a concise explanation for the rationale behind the decision" (ICO). Jev gives a percentage and no reasons. Routing support requests is very different from rejecting a job applicant.
In summary
Jev is not the first AI of its kind, but its launch was followed by dozens of similar models within two weeks. OpenAI has since launched a rival service in beta, and by 7 October an open model from Perplexity scored above Jev on a community leaderboard. Independent tests show Jev is cheaper and faster than the language models it was tested against. Its accuracy depends on how the questions are set up, a simple hand-written rule did almost as well in one test, and its answers can shift with the order of the options or one planted entry. It cannot say "I don't know" or explain itself. For a UK business, TypeSafe processes the data in the United States, and significant decisions about people made solely by automation need a way to challenge them.
In my view, Jev looks like a genuinely useful tool for repetitive, high-volume decisions, from routing requests to sorting emails. This is not a question of whether it is better than a chatbot. It is about having the right tool for the job. But I would use it the way I think all AI should be used in a business, to take the routine work off people's desks, not to make the decisions that affect people. Test it on your own data, give it an "other" option, and keep a person in the loop for anything that matters.
Sources(33)
ShowHide
- Archestra, We tested Jev on 100 real agent calls
- Bloomberg, Jev, an AI Model That Can't Chat, Takes On Bigger Rivals, 25 September 2026
- Cloudflare, Jev (typesafe)
- Cloudflare, Models
- Data Privacy Framework, Participant list
- Forbes, Jev Cuts AI Decision Costs 100x And Vercel, Cloudflare Rushed To Add It, 19 September 2026
- Good Start Labs, Verification is the bottleneck
- ICO, A brief guide to international transfers
- ICO, Automated decision-making safeguards (draft guidance)
- Jev Decision Index
- jev-phishing-bench, 17 September 2026
- legislation.gov.uk, UK GDPR Article 22C
- n8n community, Jev in n8n, a node for TypeSafe's decision model, traced in Langfuse
- Near Here, TypeSafe Jev vs Mistral vs Gemini, Event Validation Test, 16 September 2026
- O'Reilly Radar, Will TypeSafe's Jev Change How We Build AI Applications?, 23 September 2026
- OpenAI, API changelog, 6 October 2026
- OpenAI, DevDay 2026 Recap, 29 September 2026
- Opper, Jev vs. Kev, an open-source Jev alternative, tested side by side, 25 September 2026
- PriorBench, Jev pre-registered evaluation, 20 September 2026
- SiliconANGLE, TypeSafe AI exits stealth with $40M, 16 September 2026
- The D AI LY Brief, TypeSafe's Jev Scores 62.6% Asked Once and 95% Split Five Ways, 20 September 2026
- The New Stack, OpenAI answers TypeSafe's Jev with a Decision API built on Luna, 29 September 2026
- TypeSafe, Confidence (documentation)
- TypeSafe, Data Processing Addendum, 24 April 2026
- TypeSafe, Introducing System One Models and Jev, 15 September 2026
- TypeSafe, Introduction (documentation)
- TypeSafe, Models (documentation)
- TypeSafe, Privacy Policy
- VentureBeat, Companies are putting Jev in charge of AI agent decisions, 21 September 2026
- Xiong et al., Can LLMs Express Their Uncertainty?
- Yin et al., Benchmarking Zero-shot Text Classification, 2019
- Yoon et al., Reasoning Models Better Express Their Confidence
- Zaratiana et al., GLiNER, Generalist Model for Named Entity Recognition using Bidirectional Transformer, 2023
