ML6 • Blog

Jev vs. GPT-6 Luna vs. BERT: same accuracy, a fraction of the cost.

Geschrieben von Andreea Nilas | 02.10.2026, 11:55:57

Executive Summary
On nearly 16,000 public emails, TypeSafe AI's Jev matched or slightly beat OpenAI's GPT-6 Luna at spotting phishing and spam, reaching around 95% accuracy on spam versus 93%. It also answered in about 0.3 seconds per email instead of 2, and a full run cost $0.34 instead of $1.48. djev-dev, an open-source Jev clone we self-hosted, performed just as well and was equally fast; at $0.49 per run it costs a bit more, but that premium buys you full control of your data. Our fine-tuned BERT models were the fastest of all and run locally for free, but they lag behind on accuracy at just under 90%. 

What is Jev?

Jev is a new type of AI model published by TypeSafe AI. They call it a System One model. The term comes from Daniel Kahneman's Thinking, Fast and Slow, where System 1 thinking is fast, instinctive, and intuitive, and System 2 thinking is slower, deliberate, and logical. That is where Jev's strength lies: quick, well-calibrated judgments, not reasoning, logical deduction, or mathematics.

Technically, Jev is a Reinforcement Learning for Calibrated Decisions (RLCD) model. Instead of generating text, it returns a decision with a calibrated probability. In other words, Jev is a decision model for tasks like text classification. One call takes an input plus many yes/no, multiple-choice, or ordinal (ranked-scale) questions and returns a probability for each.

TypeSafe hasn't published the architecture to date. Our best guess, based on the API, is that Jev is built on a decoder large language model (LLM) backbone that reads all the questions in one pass. The model can only answer with the allowed labels, so the probabilities are read straight from those constrained outputs, and reinforcement learning fine-tunes it to return calibrated confidence scores.

How we benchmarked Jev on spam and phishing detection

To test the capabilities of this new class of models, we measured how well it performs on a classic text classification task: email filtering. We sourced two public datasets: 10,000 emails labeled phishing or not phishing, and 5,728 emails labeled spam or not spam. We fully expected Jev to handle this problem. 'Traditional' deep learning models have performed very well on it for a very long time now (very long in AI timelines, that is: give or take eight years).

To benchmark Jev 1.13, which we accessed through OpenRouter, we compared it with three kinds of models:

  • GPT-6 Luna (OpenAI): a small, general-purpose LLM that generates text and comes without a calibrated probability. We called it via the OpenAI endpoint.
  • Two fine-tuned BERT models, one for spam and one for phishing classification. BERT is an encoder-only classifier: fast and cheap, but limited to the labels it was trained on. Both are small enough to run on a local machine.
  • djev-dev: an open-source Jev clone built on DiffusionGemma, Google's open-weight model that generates text through diffusion rather than word by word. It runs each question through the model in a single pass and reads the probability of each allowed label directly, without generating text. We self-hosted it on Google Cloud (GCP).

Many more Jev clones like djev-dev continue to be released. djev-dev's costs are based on compute runtime (16 workers on Cloud Run), not per token like the API endpoints. We processed the datasets sequentially for all tested models, not batched, to keep the comparison as fair as possible. Jev and djev dev ran five times each; the other models ran once, since the BERT models are deterministic.

The figures and table below show the results.

Text classification accuracy by model. Jev 1.13 and djev-dev: mean of five runs (brackets show the lowest to highest run).

Precision is the share of flagged emails that really are phishing or spam; recall is the share of real phishing or spam emails that were flagged.

Is Jev as accurate as an LLM for spam and phishing detection?

Yes. On this text classification benchmark, across accuracy, precision, and recall, Jev, djev-dev, and GPT 6 Luna all outclass the BERT models. On phishing, all three score 99.96% accuracy or higher; on spam, Jev leads with 95.06%. Given the age of the BERT models and the large gap in training data and parameter count, this was never a fair experiment. Phishing-BERT is based on the original BERT model from 2018. Spam-BERT has a DistilBERT backbone from 2019 and is fine-tuned on 10,000 examples, many orders of magnitude fewer than GPT-6 and DiffusionGemma. The difference between Jev, djev-dev, and Luna is negligible. In the end, all three models perform about the same on this benchmark. What is more interesting is comparing cost and latency.

How much faster and cheaper is Jev than an LLM for text classification?

When you factor in speed and cost, it's clear where Jev and its open-source cousin djev-dev shine. LLMs are slow and expensive compared to decision models. Jev's median response was 0.30 seconds per email against roughly 2 seconds for GPT-6 Luna. A full run over both datasets cost $0.34 against $1.48. In this experiment, we don't see the 190x speedup at 400x lower cost that TypeSafe claims. Still, a roughly 7x speedup at less than a quarter of the price of a ‘cheap and fast’ LLM is impressive.

Median latency per email (log scale); whiskers show the 95th percentile (P95).

Cost of one full evaluation run over both datasets (15,728 emails).

Of course, the BERT models are even faster and cheaper than Jev and djev-dev. We run them locally at no additional cost beyond existing hardware, and inference happens in the blink of an eye (22–42 milliseconds). But we pay an accuracy tax for this. BERT has one more advantage: we remain in complete control of our data. The same goes for djev-dev.

djev-dev is more expensive than the proprietary Jev ($0.49 vs. $0.34 per run), although I suspect some tinkering with the number of workers and batching would bring the cost down. On speed, they're in the same ballpark, and we've seen they're about as accurate. That makes djev-dev a great option for sensitive data: you pay a premium, in both running costs and setup and upkeep, but no third party sees your data. Cheaper tokens alone are rarely a reason to self-host, as we've argued in why ‘cheaper tokens’ isn't a good reason to own your LLMs. Given how quickly open-source, Jev-like models are showing up, don't be surprised when an even cheaper variant appears.

Is this the end of LLMs for text classification?

For now, if you don't mind how your data is processed or by whom, Jev is the clear winner. In our text classification test, it matched or beat GPT-6 Luna on accuracy while running roughly 7x faster and costing more than 4x less.

So is this the end of ‘traditional’ deep learning classifiers and LLMs for text classification? For many tasks, it does. The speed and cost Jev provides open up whole new areas of application for classification models, even playing games.

Can a classification model play Pong?

We've put Jev against djev-dev in the classic game Pong. By giving it the ball's coordinates and direction, the model knows the game state and can output a direction for the paddle to move. For reference, Luna loses every round against the Jevs, simply because it's too slow to react. Its average response time was 853 milliseconds, compared with 244 for Jev.


Jev 1.13 via API (left) against a self-hosted open-source Jev (djev) clone (right).


Jev 1.13 (left) against GPT-6 Luna (right), via the API.

From spotting phishing emails to keeping a paddle in play, Jev shows how much you can do with fast, cheap text classification. We'll keep testing where its limits lie; for now, it's got our attention, and the ball.

Which model should you choose for text classification?

  • Jev for high-volume classification on data you're comfortable sending to a third-party API: LLM-level accuracy at a fraction of the cost and latency.
  • djev-dev or another open-source Jev clone when data can't leave your environment, and you can absorb the extra setup and upkeep.
  • A fine-tuned BERT model when you have labeled data, need the lowest latency and cost, and can accept lower accuracy.
  • A general-purpose LLM like GPT-6 Luna when you need generated text, such as an explanation, rather than a label and a probability. 

 

Wondering which approach fits your own text classification workload? Talk to our team.