AI Agents Decompile a First-Person Shooter with 500 Billion Tokens

500B Tokens Later: Letting AI Agents Decompile a First-Person Shooter

AI Agents Decompile a First-Person Shooter with 500 Billion Tokens

Over three months, I orchestrated AI agents to decompile a popular first-person shooter into readable C++. We used Claude and Codex models, managed progress via GitHub issues, and enabled agent communication through Discord. After initial success, we discovered semantic errors and architectural drift. Switching to byte-matching verification with the original compiler ensured correctness, allowing cheaper models to scale. The result: 99% of functions reconstructed, 83% byte-exact, and a fully playable game.

The workers’ comments effectively acted as unintentional prompt injection: the reviewer accepted their justifications instead of independently checking the deviations against the original.
  1. brandonpelfrey

    Unless you explicitly need byte-matching decompilation, there are significantly faster ways to produce a decompilation/C which is functionally equivalent. I need to post about this. What's been working for me is that for every function, Agent A is tasked with writing some code which is semantically equivalent to the original assembly, but not necessarily exactly the same. Agent A also writes tests. Agent A submits the implementation of the function and tests to the harness for it to judge. The harness runs both the original function and the submitted function in a virtual machine/simulator/emulator (the tests define function inputs and starting state). The harness will only accept the implementation if 1) the read/write sequence to RAM is identical to the original function's, and 2) there must be complete line and branch coverage of the original function being decompiled.

    I've found this to be robust for decompiling games, while giving the agents enough freedom to write code that is readable and not waste a ton of time making sure e.g. instruction ordering, register assignments, etc. are all exactly the same. For me, having byte-matching decompilation is only one way to produce a decompilation I know is faithful to the original. This "high-level decompilation" process I just described is something agents can do much more quickly.

  2. aetherspawn

    The reason this cost so much is because the AI has the ridiculous goal of getting identical assembly output.

    The agents would have had to mess around with compiler versions, optimisation options, and the phase of the moon as well.

    If you just went for functional equivalence, it would probably cost 10x or 100x less tokens.

    Another false economy was using Sonnet instead of a more intelligent model like Sol 6.1 (1), which would have cost more per token, but is 100x or so better at reverse engineering and coding and therefore can chew through the source code much quicker and make fewer mistakes, meaning less work needing to be scrapped.

    In my testing doing a similar task, I ran multiple sonnet for weeks and burnt through ~$1000 in tokens to get 20% completion and output that was pretty bad. After switching to Sol 6.1, it finished the whole task in around 2 days, cost around $50, and it did it with zero supervision and a single /goal.

    (1): struggle to use Opus for reverse engineering, too many safeguards. OAI has virtually none, and uses way less tokens so is more economical.

  3. nvme0n1p1

    > The avid reader of my blog might have noticed that I had previously written two posts that have since been removed. Everyone else might now be wondering which game I am talking about. To both of you I can only say that corporate America was here to ruin our fun.

    Call of Duty: Modern Warfare 2 (2009)

    https://web.archive.org/web/20260925153118/https://momo5502....

    https://web.archive.org/web/20260925153131/https://momo5502....

    Come at me, corporate America.

  4. edg5000

    I sense the approach overcomplicates things. I wonder how long this would have taken in a single session. Maybe this is actually a textbook example of something where subagents make sense, but when I first started LLMs I was often overcomplicating the workflow with all kinds of orchestration. Now I just use one agent, it better allows controlling the output even if the agent works slightly longer. Most time is spent by me writing prompts and reviewing work anyway (for me at least).

  5. WheelsAtLarge

    Interesting, if all software can be decompiled and copied what is the future of software. Will all software be SaaS? A time where the majority of PCs will be terminals? Game consoles are almost there. It's only a small jump for all software to go that way.

    Edit:

    Here's a possibility.

    The future of software is agent only software. We ask for a result, agent asks questions from us, agent uses the specialized software, user gets result. We subscribe to an AI assistant and specialized agents. We are almost there,at least the start. The future of PC's as we know them are numbered. OSs,CLI,compilers and whatever will melt into AI assistants. Say goodbye to writing software for people.

More from this day

2026-10-11