15 min read

What Is Jev AI? A Guide to the New AI Decision Model

What Jev AI is, how it differs from a large language model, how it did in independent tests, and what a UK business should check before using it.

Craig Moore

Founder, Apex Insights

Cartoon of two robots at desks under a sign reading Today's weather forecast. A big ChatGPT robot, buried in papers, maps and weather books, raises a finger to explain, while a small grey Jev robot at a tidy desk holds up a card showing a rain cloud and 100%

AI-generated illustration by Craig Moore

In this article

In my line of work, there are few things I find quite as satisfying as automating a complex workflow, taking a task that people find laborious, time-consuming and complicated, and turning it into a simple system that runs itself with next to no supervision, freeing up people's time to get on with the important things. So when a new kind of AI model becomes available, adding another tool to the automation toolkit, I am always keen to learn more.

TypeSafe AI says its new model, Jev, is 194 times faster and 445 times cheaper than large language models, in its own workflow tests. Independent tests agree it is cheaper and faster than the language models it was tested against, which cost anything from about 1.6 to 580 times as much (Good Start Labs, Forbes). Within a fortnight, dozens of similar models had appeared and OpenAI had announced a rival service. Judging by the traction Jev has been getting on social media since its launch, it's fair to say that the hype train has fully left the station (Bloomberg). Let's have a look at whether Jev lives up to the claims, and what a UK business should check before using it.

What is Jev?

Jev is an AI model built to make fast decisions that software can use directly, rather than to write text (TypeSafe docs). Large language models, like the ones behind ChatGPT, "are designed to produce text for humans to read", so a workflow has to dig the decision out of their reply (TypeSafe docs). Jev's answers come back ready to use, each with a probability, and TypeSafe says it built Jev "entirely focused on automation". It "is neither small nor an LLM", in TypeSafe's words (TypeSafe).

Jev launched on 15 September 2026. Its maker, TypeSafe, is a San Francisco start-up led by former OpenAI researcher Diogo Almeida, with $40 million of seed funding (SiliconANGLE). It is named after William Stanley Jevons, and TypeSafe calls it a "System One" model, after the fast thinking in Daniel Kahneman's Thinking, Fast and Slow (TypeSafe). TypeSafe has not published how it is built (Opper).

What Jev does, and what it cannot do

Jev works like a form. You give it some text, such as a support ticket or an invoice, and questions set in advance. It picks an option from a list, gives a score, or says how likely something is to be true, and every answer comes with a percentage (TypeSafe docs). That matters, because research suggests language models "tend to be overconfident, potentially imitating human patterns of expressing confidence" (Xiong et al.). Reasoning models do better (Yoon et al.), and asking several times helps (Xiong et al.), but each extra call costs time and money. The evidence on Jev's own percentages is mixed. Opper, which also hosts a rival model, found them more reliable than that rival's (Opper), but in a phishing test, asked a single question, they were less reliable than Claude Haiku's (jev-phishing-bench).

Here is a real example from Cloudflare's documentation, for the message "Help! My payouts have been failing for 3 days." (Cloudflare)

How Jev AI turns one customer message into three decisions

Because Jev picks from a list rather than writing, TypeSafe says it "can't hallucinate" (TypeSafe). That is true only in a narrow sense. Given billing, technical and sales, it cannot invent a fourth option, but it can still pick the wrong one. And it never explains itself. As one reviewer put it, "The biggest loss is the explanation" (O'Reilly).

Speed and cost in independent tests

Even TypeSafe says its headline figures are likely to be "on the higher end of real world gains" (TypeSafe). Independent testers found Jev cheaper than every language model it was tested against, and about 3 to 25 times faster against named models.

Cost per 1,000 decisions in independent tests

This article

*Jev is shown once, at its highest cost, $0.16 in Good Start Labs' rubric test against five language models, of which DeepSeek was the cheapest and Claude Fable the dearest. It cost $0.043 in Near Here's event listing test against Gemini and Mistral, and $0.038 in the test emails against Claude Haiku. Each test used a different task. Near Here ran Gemini and Mistral on high reasoning settings and had them write a short explanation, which Jev did not. Good Start Labs costed the language models from its bills and Jev at TypeSafe's list price, per million checks, divided here by 1,000. Dollars are US dollars.

Time per decision in independent tests

This article

*Jev is shown once, at its slowest, 0.59 seconds on average in Near Here's event listing test against Gemini and Mistral. It took a median of 0.239 seconds in the test emails against Claude Haiku's median of 0.687, measured from France and reported by The D AI LY Brief, a median of 0.35 seconds in a 12-passage review test reported by Forbes against Claude Fable 5.1 at high effort, and 0.087 seconds per decision in an n8n user's test against a language model they did not name. Each test used a different task.

Each test used a different task, so the fair comparison is within each test, and the saving depends mostly on the rival. In a test by Good Start Labs, which had early access and used a July version of Jev, DeepSeek V4.1 Flash cost about 1.6 times as much as Jev and Claude Fable 5.1 about 206 times as much (Good Start Labs). In a 12-passage review test reported by Forbes, Fable cost about 580 times as much (Forbes).

How accurate is Jev?

In one test, how the question was asked mattered most. Asked once whether emails were phishing, Jev did worse than Claude Haiku 4.5. Split into five narrower questions and tuned on labelled emails, it caught up, although a simple hand-written rule with no AI did almost as well (jev-phishing-bench).

Accuracy on test emails, half of them phishing

Five questionsNo AIOne questionThis article

One-question scores are on all 2,000 emails, 1,000 of them phishing. The five-question scores and the hand-written rule were measured on 1,000 held-back emails, after the answers were weighted using the other 1,000. The tester found the gap between Jev and Haiku at five questions too small to be meaningful.

Other tests are mixed. Jev beat Gemini and Mistral on the main set of listings a UK events website used to tune its prompts, but Gemini edged ahead on fresh ones (Near Here). Jev agreed with five language models a little less often than they agreed with each other (Good Start Labs), yet scored about 96% with no training examples in a pre-registered test, against about 77% for hand-written keywords (PriorBench) and 95% to 97.5% on fresh data in Opper's tests (Opper).

Jev also cannot say "I don't know". It must pick one of your answers, even when none fit. One tester saw a cake recipe filed as a technical issue at 94% confidence (PriorBench). Giving it a way out helps. With an "other" option, Jev flagged 21 of 30 off-topic messages, against none without one (PriorBench).

Similar models

Jev is not the first AI built to pick an answer rather than write one. Researchers were benchmarking the idea in 2019 (Yin et al.). What followed the launch was quick. On the morning of 27 September a community leaderboard listed 67 similar models, 64 of them released since Jev, although 15 were new ways of running existing models (Jev Decision Index). On 29 September OpenAI announced its own Decisions API, built on its GPT-6 Luna model (OpenAI), which The New Stack called "likely a reaction to TypeSafe and Jev" (The New Stack). It opened as a public beta on 6 October (OpenAI). By 7 October Cloudflare also listed decision models of its own (Cloudflare), and the leaderboard listed 112 models, with an open model from Perplexity scoring 62.75 against Jev's 60.11 (Jev Decision Index).

Decision models before and after Jev

  1. 2019

    Researchers benchmark AI that labels text it was never trained on

  2. 2023

    GLiNER, a small model that picks out items in text in one pass

  3. 15 September 2026

    TypeSafe launches Jev

  4. 16 September 2026

    11 Jev-style models from at least 8 projects appear in a single day

  5. 17 September 2026

    A community tracker of Jev-style models opens

  6. 22 September 2026

    The tracker adds a scored leaderboard

  7. 27 September 2026

    67 models on the leaderboard, 64 of them released since Jev

  8. 29 September 2026

    OpenAI announces a Decisions API

  9. 6 October 2026

    OpenAI opens its Decisions API as a public beta

  10. 7 October 2026Where it stands

    112 models on the leaderboard, one from Perplexity scoring above Jev

Model dates are when each model's Hugging Face page was created, or its GitHub page where it has none, which can be before it was announced. Only models entered on the Jev Decision Index are counted, and Jev itself is not included. The count was 67 on the morning of 27 September 2026, and 15 of those 67 were new ways of running existing models rather than new models. It was 112 on 7 October 2026.

Practical advice for UK businesses

None of this is legal advice.

Test it first

Try Jev on your own data before switching anything on. TypeSafe suggests acting automatically only on high-confidence answers, starting with "conservative thresholds" (TypeSafe docs). Jev's answers can shift with the order the options are listed in (Archestra), or with one planted entry in the input (VentureBeat). Fix the model version too, because TypeSafe's default version "moves when a new release ships" (TypeSafe docs).

Where your data goes

Hosting matters if you send Jev personal data. TypeSafe's privacy policy says "The Services are hosted in the United States" (TypeSafe privacy policy). TypeSafe is not on the Data Privacy Framework list (Data Privacy Framework), so a UK firm would rely on the UK Addendum in TypeSafe's data processing terms (TypeSafe DPA). The ICO says a firm relying on a safeguard like this must also complete a "transfer risk assessment" (ICO).

The law on automated decisions

Since 5 February 2026, when a significant decision about a person is made solely by automation, the organisation must tell that person about it and let them make their case, get a human review and contest it (UK GDPR Article 22C). The ICO's draft guidance says organisations "must provide a concise explanation for the rationale behind the decision" (ICO). Jev gives a percentage and no reasons. Routing support requests is very different from rejecting a job applicant.

In summary

Jev is not the first AI of its kind, but its launch was followed by dozens of similar models within two weeks. OpenAI has since launched a rival service in beta, and by 7 October an open model from Perplexity scored above Jev on a community leaderboard. Independent tests show Jev is cheaper and faster than the language models it was tested against. Its accuracy depends on how the questions are set up, a simple hand-written rule did almost as well in one test, and its answers can shift with the order of the options or one planted entry. It cannot say "I don't know" or explain itself. For a UK business, TypeSafe processes the data in the United States, and significant decisions about people made solely by automation need a way to challenge them.

In my view, Jev looks like a genuinely useful tool for repetitive, high-volume decisions, from routing requests to sorting emails. This is not a question of whether it is better than a chatbot. It is about having the right tool for the job. But I would use it the way I think all AI should be used in a business, to take the routine work off people's desks, not to make the decisions that affect people. Test it on your own data, give it an "other" option, and keep a person in the loop for anything that matters.

Sources(33)

Show

Responsible use of AI

Every article here is researched, written and edited by me. I use AI tools to help me find sources, to structure the article, and check spelling and grammar. The ideas and final wording are my own. Where possible, I follow the ten commitments of the Responsible AI Use Campaign.

The views and opinions in these articles are my own. They do not represent any organisation I work for or am affiliated with.

See my public declaration

You stay anonymous on this site.

We count visits only to help us improve it. No tracking or advertising cookies, and we never identify you or follow you to other sites. Cookie Policy