Jev's decision models don't beat LLM-as-a-judge or traditional classifiers
Decision models like Jev don't beat LLM-as-a-judge or traditional classifiers

Red Hat benchmarked nine guardrails across four paradigms on prompt injection and content safety. Jev, a zero-shot decision model, ranked fourth in both tasks, trailing a 200M-parameter pretrained classifier and a 35B LLM-as-a-judge. While Jev offers schema guarantees and speed, it doesn't outperform established methods, and open-source alternatives like Laya and DiffusionGemma match or exceed it.
However, it is reasonable to question whether TypeSafe's approach is truly as novel as claimed. Arguably, decision models have existed for years under the name "zero-shot text classifiers," such as Meta's BART-large-mnli model from 2019.
- NeumannGod
This is a false comparison. What Jev does is fundamentally different from what LLMs are doing as a judge.
For a quick primer, Jev is able to provide confidence scores on its classification, i.e. it is able to calibrate how well it is able to predict. Being able to predict in a distribution is different from being able to calibrate confidence of the predictions which should happen from the question or domain distribution from which the decisions are predicted - being able to do that is tough and is not same as using LLMs logit probabilities which are predictions in the vocab space. Though both are loosely correlated and might converge as LLMs keep getting better, the former is a much stronger decision-making signal than the latter. Jev not beating LLM-as-a-judge might be due to various other reasons such as world knowledge etc, but Jev as a concept will always provide more reliable decisions / outputs than LLM-as-a-judge giving a scalar score.
- AnthusAI
That was a pretty simple task they gave it, and sure you can use BERT with sequence classification for simple classification tasks.
In our benchmarks, Jev did a LOT better at multi-step reasoning tasks than any open decision model we have tested so far, and it was also better than GLiDE which was specifically designed for that kind of task. And also better than Luna. On accuracy and also confidence calibration but also time and cost.
- Garlef
I think it's a bit early to call the race.
I think the abstraction is a useful one - a general purpose classifier that does not need to be specifically trained: Unstructured signal in, structured judgement out - with a focus on speed and cost efficiency.
And since there is not yet a large body of benchmarks, I don't think we have sufficiently explored how to measure these things.
But since there's a market and some hype this will soon happen.
(And it's not like TypesafeAI has a real moat or invented something entirely new here ~ they just managed to put things into one coherent perspecive)
- mertcikla
Jev may or may not have truly innovated on AI architecture but it still kick-started a new paradigm.
its a breather after waves of llm wrappers.
- segmondy
duh, this is not news. (general, fast and cheap) before decision models, you could pick only 2.
LLM as judges - generalized, but too slow. If you had to make millions of classifications a day, this will be the wrong approach. you won't/shouldn't use LLM to classify spam/no spam. hot dog/or something.
traditional classifiers, very specific 1 trick pony, super fast and cheap once built. If you need to make tons and tons of classifications, this would be the approach. but if you wanted a classifier right now for a novel problem, you need an expert to curate data, train and deploy.
decision models/jev - are generic, you can throw them at most generic classification problems, and they are good enough. it's a fine balance between general, fast and cheap. you get all 3