Code Review Is More Than Bug Detection—and AI Can't Replace the Human Judgment It Requires
There is more to code review than (automatable) detection
A new paper argues coding agents can replace human code review because they match its stated functions: defect detection, style enforcement, knowledge transfer, and awareness. This response pushes back, calling that a 'substitution myth.' It points to what can't be decomposed into functions: a reviewer's genuine confusion as a signal of poor design, skepticism about whether a change is even necessary, noticing what's missing, calibrated attention based on who wrote the code, bidirectional learning, operational context that lives outside repos, and personal accountability. Code review is also coordination, sensemaking, and governance.
The most fundamental issue I have with the article is that it assumes code review is a first and foremost a detection process: you find defects, style violations, security issues, etc., and the assumption is that detecting these faster and cheaper is universally better. But code review is also a coordination process, a sensemaking process, and a governance process.
- n4r9
There's been a lot of talk about the purpose of code review recently. It makes sense in the face of AI. Heres a link that was submitted a little while ago: https://mathstodon.xyz/@mjd/115096720350507897
And in response I wrote a non-exhaustive checklist of things that a code review can look for:
- Does it functionally achieve what it sets out to (as per tacker issue or PR description)?
- Does it have extraneous code? Leftover debug prints, private API keys etc...
- Does it have any obvious defects? Memory leaks, un-handled edge cases, security flaws, obsolete API calls, etc...
- Could it be more understandable? Add/remove abstractions, better variable/method names, more/less functional etc...
- Is the style consistent with the codebase and/or style guidelines?
- Are there obvious performance improvements? Hashset instead of list, lazy evaluations, etc...
- Is it sufficiently well tested?
I think LLMs are okay at most of these, and worst at the first.
- LunicLynx
Unfortunately this often represents the only feedback given by the people in those „higher“ positions.
„The indent is wrong here“
„Comments should end with a period“
Because this kind of feedback is and was always easy.
- hazard
Pangram check on the article: 94% of this text is AI