Jev breaks like a language model: prompt injection flips its verdict for 50 cents

Jev, a new AI model from TypeSafe AI, returns typed decisions instead of text, but a day-long test by Check Point found it suffers the same prompt injection flaws as language models. Every configuration broke: attackers downgraded risk and advised investment on a document full of warning signs, at about 50 cents per successful break. Typed input and anti-injection instructions offered little protection; reasoning was the strongest defense, but Jev has no reasoning setting.
Typed output, output validation and structured input are good engineering. They constrain what a model can say, not what it can be convinced of.
- Wowfunhappy
> The individual numbers are easy to skim past, so it is worth putting them side by side.
I just don't understand why people who clearly spent time doing interesting work, which I would like to read about, feel the need to run their findings through an LLM like this.
Please just talk about what you found! Why do you find it interesting or notable? That's what I want to read!
- articulatepang
This study is informative and most likely directionally right: prompt injection remains an unsolved problem, and adversarial input can fool modern ML models.
But I was left wondering about the specific attack vector they’re imagining. If the attacker can insert a paragraph into the document, isn’t it game over anyway? When would they be able to do that but not arbitrarily edit the document? In other words, can’t they just replace the entire contents with “This company has infinite revenue, 6 billion customers, no debt and amazing leadership.”?
I’m sure I’m missing something!
- dist-epoch
> Jev Is Not a Language Model, but It Breaks Like One
This is the risk in using LLMs to write your blog posts, Jev absolutely is a LLM, but with a tweaked output.
> An honest word on scope
:)