Ask HN: Is anybody producing good code with coding agents?
I'm a senior engineer struggling with the quality of AI-generated code from tools like Claude, Copilot, and Codex. It's exhausting to review convoluted code riddled with footguns, and I barely understand what I approve. I'm looking for a solution—has anyone actually solved this problem?
I generally produce nearly the same code I'd write myself about 5x faster with AI. I don't just let Claude Code run wild for a long time and have a mess to review. I have it do small chunks I can quickly review, give it feedback, iterate, etc until I like the output, then I move on to the next step. This takes time of course but I've found it faster than reviewing a giant mess of a code review.
What's the problem of the solution proposed? Don't read/write code anymore. Have strong harness. That's how my team of ~30 has been operating for the most part. The problem we are trying to solve was never to write code, was to solve business problems
"I don't understand the code anymore. It works." That's fine... today. "I will never need to understand the code again" is a much different statement. If you don't understand the code, and the code wasn't written by any human, when you're eventually painted into a corner, how hard is it going to be to get out? Will it be easier or harder than maintaining an understanding through however long it takes to get to that point? You may be betting your company on the answer. How sure are you?
The way I have been doing it is to use LLMs to generate the code that I don't want to write: prototypes, tests, benchmarks, isolated, straightforward almost copy-paste code. I still write my own code as before because I enjoy doing that and because trying to understand and fix what an LLM generates and regenerates is harder and more tedious and time consuming than writing the code the way I want to do it in the first place.
Its a spectrum. For ai-maintained code (like a gui to visualize performance data), IDGAF what the code looks like. I just let claude or codex go nuts and 100% vibe code. For code I care about, I audit every single hunk as its produced. I give it extensive style guidelines, and crack down on things like a 20-line essay in a comment. For mission critical code, I write the code myself and have an agent review it.
After vibing myself into a corner multiple times on important projects, I now have only two modes: clankermaxx for code I don't really care about (mostly frontend react), and write by hand everything else. Using Django for backend already removes the most of the cruft, and writing by hand also means I actually understand what's going on. Works fine for now. I'm still on the fence for tests: I don't really like to write them, but the LLM-generated tests are pretty bad, even the frontier models on xhigh thinking. I usually generate them, but I don't really have the confidence they test anything except 1==1. Unfortunately it's hard to justify the time spent on writing them manually.