Grok 4.5, GPT-5.5, and Claude Build-Off: Who Codes Best?
We made Grok 4.5, GPT-5.5, and Claude build the same apps
We pitted Grok 4.5, GPT-5.5, Claude Opus 4.8, and Claude Fable 5 against each other to build three distinct apps in a single HTML file. While the Claude models nailed the complex 3D Rubik's Cube first try, GPT-5.5 created the most mesmerizing gravity sandbox, and everyone delivered a playable Breakout game. Ultimately, Grok 4.5 emerged as the speed and value champion, streaming twice as fast as competitors at half the cost, despite stumbling on its initial cube attempt.
That is not just drawing, that is comedy.
- dadoum
I tried to one-shot the first test (the Rubik's Cube test) with LucidQuery's Swift model, to test it, as there are not much benchmarks about it and that they brag a lot about it, and I was pleasantly surprised to see it achieving a result similar to Grok 4.5 but in one shot (there is the same issue that if you scramble twice the solve button does not work anymore, but it got it in one shot).
Though it crunched most of the free quota, 47111 tokens, so I couldn't make multiple attempts.
- wwind123
Half year ago I tried to use Codex, Claude and Gemini build the same scripts to automate various things on my machine. Claude was the clear winner back then, making the most reasonable assumptions, presenting results in the easiest-to-read format, writing runnable script with minimum dependency. Half year later I think Codex and Claude models have both advanced a lot, but Gemini is still lackluster. Gemini could catch problems when reviewing Claude/Codex's design plans and code, but it's hard to make Gemini make complex plans or implement complex code by itself.
- mlmonkey
I am 99% sure the post was written by AI
- jeffgreco
So strange to write a whole post with Claude giving the best results and Grok consistently the worst, but awarding Grok the winner because at least it did the worst fastest?
- steve-atx-7600
I have not used grok 4.5 yet, but the other pictures match my experience doing anything graphical with the other models that it cracks me up. gpt 5.5 has no design sense whatsoever. It cannot even make terminal output not look terrible. I've asked it to use colors and formatting in various ways and got goofy randomly colored output. opus 4.7 and later seemed to have an inuitive design sense by comparison - 2d or 3d. Fabel 5 is just rock solid.
Yes, subjective. But it matches my repeated experiences with these models for what it is worth.