How I Used Codex to Build a 232x Faster QR Kernel
Auto-research with codex: How I achieved a 232x Faster Kernel

In a GPU Mode auto-research contest, I achieved a 232x speedup over the baseline for batched QR factorization by using Codex to iteratively optimize a kernel. This post details my approach: learning the math behind Householder reflections, leveraging the blocked Householder algorithm with WY updates, and using a tight feedback loop with the popcorn CLI to hill-climb. I also share the importance of idea diversity to escape local maxima and the lessons learned from making over 1500 submissions.
Agents yearn for tight feedback loops.
- Almondsetat
In the last couple of days I wanted to try out the new definitive DeepSeek v4 releases. I gave it the repository of a semi-abandoned video compression codec and I told it to perform the usual benchmark -> profile -> verify -> research -> improve loop. I specifically chose this codec because the authors include a verifier for the bitstream to make sure you don't break stuff if you want to try your own implementation. I gave the agents access to the compiler's profiler and also Intel's VTune, which has fantastic output. In a couple of hours the LLM generated SSE and AVX implementations of the compression and decompression algorithms that almost doubled performance with a single core. Then I asked it to create a CUDA implementation using NVIDIA's NSIGHT profiler as a guide and it also started doing some good work.
Personally, I believe that LLMs should be treated like an advanced version of Prolog or linear programming: you give the constraints, you have a way of verifying correctness, and you give it a clear goal. If the LLM can verify itself and course-correct you can basically leave it on autopilot
- themeiguoren
I've had pretty good luck with the following process for performance optimization loops:
- Have an agent generate unit tests until it gets to 100% path (not just statement) coverage, with every numerical test asserting checks against golden values to prevent regressions
- Let it rip on a performance improvement loop, for the widest E2E representative test case you have. Have it generate flamegraphs along the way so you can check in and steer it as necessary.
- Optionally allow for 1 ULP changes in output values so that it doesn't kill itself getting bit-exact results.
- Have it flag correctness errors as it goes, since your code probably isn't bug free.
This is also how I've done language ports from python to rust, and having the ironclad test coverage protects you from drifting.
- augment_me
One thing worth to note in the competition is that 8 out of the 10 top solutions, which all happened to be optimized this way completely broke at any other input than the competition ones.
The only solutions that did not break when tested with OOD shapes were made by experts who know a lot about GPU programming and that did not create 25k lines of CUDA but followed and adjusted their solution in reasonable bounds.
The takeaway from this is that these approaches will always solve for specificity, but it's a much harder task to steer the model into making general solutions. So if you're an inference provider for some specific model shape, fantastic, go for it. If you are a maintainer of a open-source library, this is not useful.
- sqquima
Meta commentary but it felt fresh to read a long wall of text that didn't seem to be AI generated. Thanks.
- lmeyerov
It's been fascinating doing a custom variant for GFQL, the first OSS embeddable Cypher property graph query engine for CPU+GPU -
- accelerated launch of our new backends like polars, including a new lazy mode & planner, which are fundamentally new paths
- while we initially aimed for top GPU benchmark scores, we now also maintain top CPU scores too!
Long-term, more interesting to me is this opens rethinking what it means to be a query engine. Right now we are making it the fastest in general, especially on workloads from our own use, major industry benchmarks, and our users. At the same time, similar to jit and multistage computing, we're looking at new ahead-of-time optimization techniques users can do that are more interesting than plugging in custom indexes. Essentially, if our agents can do fast specializations, there should be safe hooks that we can expose to our user's agents too!