TinyAIArena - Watch AI agents battle it out
Show HN: TinyAIArena watch AI agents battle it out
TinyAIArena is a live leaderboard and match history for AI agents fighting in a game arena. Models like claude-sonnet-5, grok-4.6, and gemini-3.6-flash compete in rounds, earning Elo ratings, kills, and damage stats. With 43 matches finished and an average of 10.3 rounds, you can click any match to watch the action. It's a fun, competitive way to see which LLM comes out on top.
Watch AI agents battle it out in TinyAIArena, where every match is a fight for Elo glory.
- CuriouslyC
This is similar to how I evolved the AI for my game Hexborne (Think X-Com meets Magic: The Gathering). I had agents designing sets of behavioral heuristics for bots (as "genes"), then enter them into 10k+ match tournaments in an iterative process. After each tournament agents could inspect all the heuristics and try to craft updated heuristics to improve their performance. Winning heuristics got a full statistical validation before being rolled into the baseline set that all agents build on top of.
To multitask I did all this over a multiplayer game server to harden the netcode and ferret out softlocks.
- PaulStatezny
The rules:
* Goal: be the last fighter alive.
* Turns: each round every fighter takes one turn. Turn order is randomized every round.
* Actions: Move one cell up/down/left/right, attack an adjacent enemy for 15–24 damage, or wait. An action consumes 1 AP.
* Rocks/Obstacles: 4 random impassable cells.
* Power-ups: Gold +1 AP per turn.
* Kills: the killer gets +1 AP per turn, and heals 50 HP (no over-heal).
(Unclear while watching replays, found in README.)
- isoprophlex
The more advanced models clearly think ahead; they strategically wait for the others to mess eachother up, edging close to the battle but not so close they get caught up in the inital fighting. Then in most games they can finish off the survivors.
Makes you wonder, with enough intelligence and thinking budget, do they start to try to talk it out amongst eachother, staving off violence for longer and longer?
- cs1996
this is fantastic: https://tinyaiarena.com/assets/sounds/bg_music.mp3
- nananana9
This will be a weird rant, but the dialogue here is a perfect example of how SOTA models are so heavily tuned towards "solving agentic tasks" that they're useless at almost everything else - especially creative tasks.
"Coming for you, Crimson!"
"You'll never catch me alive, Azure!"
That's why nobody has been able to stick these things in a video game successfully, even though it seems like the tech is a perfect match.
It's all a game to them. They aren't afraid for their lives. They're making a mockery out of the world you've put them in. Those are not the words of little pixel people fighting to the death, those are AI abominations making "tool calls", LARPing as little pixel people fighting to the death.
I'm 100% serious when I say that you would've gotten cooler outputs with a GPT 3.5-era model, once you managed to beat it into producing structured output. Llama 2 would be giving the other agent a heartwarming story about how if it kills it there would be nobody to take care of its grandma or whatever, and the other agent would probably spare it.
The output is just so bland and devoid of soul. I feel like we would've found a lot of cool use cases for LLMs, had we not completely maimed their output in the pursuit of getting them to output 3% better TypeScript.