Jevman - Pac-Man benchmark for AI decision models

Show HN: Jevman – AI decision models play Pac-Man

Jevman - Pac-Man benchmark for AI decision models

Jevman is an open-source Pac-Man benchmark that evaluates AI decision models in real time. Six models, including jev 1.13, GPT-6 Luna, and Clef Flash, each played 100 games against classic ghosts, with scores, survival times, latency, and costs tracked. The benchmark runs any model behind an HTTP endpoint: at every junction, the game sends maze state as JSON and asks for a direction, with a 2-second deadline. Results are replayable and verifiable, and models can join the leaderboard via pull request. It offers a fun, practical way to compare decision-making under pressure.

Every game is recorded and replays exactly, so any result can be checked.
  1. nico

    Very cool, are you also testing local, cpu-runnable/trainable models/classifiers?

    I trained some to do some interesting things, including playing doom: https://github.com/nicobrenner/jeffy

    I’ll try training one for this benchmark, seems like fun

  2. bayarearefugee

    Paying 2 cents per game to avoid playing it.

    What even is this reality.

  3. NichoPaolucci

    This is neat. At least I'm still better than the models at pac-man.

    Also love the idea of a shared pool for users to try things out. I was considering more of a crowdfunded approach for one of my toy projects, something like... Giving it $10 in credits to begin and somehow allowing users to feed a buck in if they wanted.

More from this day

2026-10-09