Jev, TypeSafe's New Classifier, Is Barely Better Than Random at Probability
How accurately calibrated is Jev?

TypeSafe's Jev, a 'System One' decision model, outputs probability distributions over multiple-choice answers. But when tested on 1,000 physics problems with known distributions like Maxwell-Boltzmann and Gaussian, Jev scored a mean total variation of 0.518—barely better than the 0.546 from a uniform guess. It tends to be overconfident and peaky, often copying parameters from the prompt. The author questions using Jev as an automated judge for LLM answers.
If Jev were to completely punt on the answer, and spread the probability evenly over all possible choices, it would score a mean TV of 0.546. But Jev scores 0.518.