PacBench - Benchmarking AI models on one-shot Pac-Man
Show HN: Pac-Bench – How well can models one-shot a Pac-Man game?
PacBench is a benchmark that evaluates how well AI models can one-shot a Pac-Man game. The project compares models from various providers, including Claude Code, Antigravity, Grok Bot, and gpt-6 Codex, using metrics like wall time, token counts, and cost. It provides a transparent leaderboard with detailed run data, helping developers understand model performance in a fun, game-based context. The benchmark is open-source and available on GitHub, inviting contributions and reproducibility.
The Grok Bot entry was written in-chat from the same short prompt (no peeking).
- yambam
This reminds me of a couple decades ago when some friends and I set ourselves a challenge to, individually, each create as much of a Pac-Man clone as possible in 10 hours. None of us had any experience in games programming or graphics coding. It was great fun and we all learned a lot.
No-one ended up with a complete clone but I loved how we all ended up focusing on different things, like pixel-perfect graphics versus accuracy in gameplay, and how we all brought our existing skills to the challenge despite not really knowing what we were doing.
I expect if we had AI models available it would have ruined the pleasure of figuring it out for ourselves. I feel kind of sad for the next generation of developers who won't have that experience.
- weitendorf
Gonna be rude and say I don't think this is an interesting or useful benchmark tbh.
Clearly some new RLAAS/dataset/env is being used for this now (it doesn't even seem that complicated, you have one LLM judge whether gameplay is recognizable as the original game or not and another trying to implement a logically/semantically identical version of the game). It's why the performance improvement on this workload has been so dramatic.
Everything is going to go from 0->1 on this benchmark in short order because of that.
- _matthew_
I don't think it makes sense to have the prompt be that short. This is basically a bench.ark of how models interpret an overly vague prompt. It should at least be "Create a pacman clone in a single html page. Make it faithful to the original" if that's what we're scoring it on.
- jmathai
This prompt is a good way to test how well models fill in missing context because it's so nondescript. They're definitely improving.
Remember when people considered you a genius for prompting with "You are a skilled writer....".
- strataspace
I tried this with DOOM. Fable 5 did a pretty shit job. Astra made pretty crazy animated sprites and was pretty good considering.
The fact that these are at all playable and 100x my programming skill level is pretty depressing from a certain pov. The ThreeJS dude posted ab how demotivated he was to continue his work, and while I was never a dev that did much with webgl, I commiserate.