JevBench v1.3.0 ranks 52 decision models: Jev 1.13.0 leads, but a classifier.dev entry scores higher

Show HN: JevBench, a reproducible benchmark for typed decision models

JevBench v1.3.0 ranks 52 decision models: Jev 1.13.0 leads, but a classifier.dev entry scores higher

JevBench v1.3.0, Benchmark Heaven's own benchmark for Jev-class decision models, scores 52 systems on 534 decisions, including 220 hard ones. The JevBench Score combines Intelligence, Calibration, Speed, and Cost at 25% each via geometric mean. Jev 1.13.0 tops the official ranking at 74.4, followed by SemIf (Qwen3.5-4B) at 73.1 and djev at 73.0. An unranked classifier.dev entry earns an honorable mention with 83.6, while several rerankers score near zero.

JevBench is Benchmark Heaven's own benchmark for Jev-class decision models: state and a bounded rubric in, a typed answer out.
  1. hbrn

    $40m in funding, 2 years in stealth.

    Performs on-par with SemIf which was built in a couple days and apparently uses raw Qwen, with no fine-tuning. SemIf runs in your freaking browser. Oh and Jev is twice as expensive?

    Is it surprising that Jev consistently thinks it's Qwen?

    I'm almost convinced that Jev is a scam. Take Qwen, fine tune it a little, tell investors it cost $10m, spend $1m on advertising, profit.

  2. ks2048

    I was trying to figure out what exactly the tests here are. I guess I found some of the questions (here: https://github.com/fstandhartinger/jevbench/blob/main/datase...)

    e.g.,

    "instructions": "Which intent does the user's message express?",

    "labels":["set_alarm", "play_music", "weather", "send_message", "turn_off_lights"],

    "state": "Play some Taylor Swift.",

    "expected": "play_music"

  3. dmix

    You can spot vibecoded websites by how they include the prompt or commit-style comments into the literal interface, instead of communicating it via visual context (or simply excluding it)

    > Sort by any column; values the run could not produce always sort last. Hover a cost for how it was priced, a latency for the endpoint. Names link to each project.

    A designer would never write this, but an LLM just inserts it by making it small grey text next to the interface, just like it does with inane code comments.

More from this day

2026-09-22