Jevman - Pac-Man benchmark for AI decision models
Show HN: Jevman – AI decision models play Pac-Man

Jevman is an open-source Pac-Man benchmark that evaluates AI decision models in real time. Six models, including jev 1.13, GPT-6 Luna, and Clef Flash, each played 100 games against classic ghosts, with scores, survival times, latency, and costs tracked. The benchmark runs any model behind an HTTP endpoint: at every junction, the game sends maze state as JSON and asks for a direction, with a 2-second deadline. Results are replayable and verifiable, and models can join the leaderboard via pull request. It offers a fun, practical way to compare decision-making under pressure.
Every game is recorded and replays exactly, so any result can be checked.
- nico
Very cool, are you also testing local, cpu-runnable/trainable models/classifiers?
I trained some to do some interesting things, including playing doom: https://github.com/nicobrenner/jeffy
I’ll try training one for this benchmark, seems like fun
- bayarearefugee
Paying 2 cents per game to avoid playing it.
What even is this reality.
- NichoPaolucci
This is neat. At least I'm still better than the models at pac-man.
Also love the idea of a shared pool for users to try things out. I was considering more of a crowdfunded approach for one of my toy projects, something like... Giving it $10 in credits to begin and somehow allowing users to feed a buck in if they wanted.