Coding Agents Are Still Dumb: A Developer's Frustration
Why Are Coding Agents So Dumb?

A developer who has used coding agents since early 2025 argues that while underlying models have improved, the agents themselves remain the bottleneck. He details specific failures: agents execute embarrassingly parallel tasks sequentially, never delegate to cheaper models, don't know their own features, write unreadable plans, stop working for trivial input, and lack real sandboxing. He calls for basic improvements like task decomposition, model routing, and OS-level security.
If I had a human employee tell me they sat idle their whole shift because they wanted my input on some superficial detail, I'd quickly fire them.
- johnfn
All these problems are solved.
> Agents can’t manage tasks
Say "use subagents".
> Agents can’t delegate
Say "use subagents that are Haiku/Sonnet"
> Agents have never heard of agents
This isn't even a problem: the author is just confused why Anthropic didn't bake all of Claude Code documentation into the harness, but of course it makes no sense to pollute context like that.
> Agents suck at communicating plans
True, communication could be better.
> Agents take any excuse to stop working
Just use /loop.
> Agents are only useful when they take unnecessary risks
Just use a sandbox. OK, technically, this one isn't literally solved by Anthropic/OpenAI, but it is solved by a million agent sandbox startups.
- Neywiny
I keep running into this. It's nice seeing others here struggle. I guess when all your doing is one-shot simple trivial tasks who cares. I've found they're great for that. But once I need to do real work, everybody makes their own esoteric abandonware that kinda works but kinda doesn't. Stars aren't a perfect indicator, but I haven't seen anything over a few hundred for these things I'm finding on GitHub. Same with downloads of plugins. It was very isolating feeling like I'm the only one not enamored by the state of this.
- polyterative
I see your point, but even the fact that I can speak to my computer and anything useful happens is still a miracle to me.I don't think I will ever get accustomed to how good the new models are.I just can't keep up.And I do this for a living.Ten hours a day.
A lot can be improved, but this is already so much speed.
- mstank
I used to relate to this article quite a bit. In the last 3-4 months, not so much. I've found that the latest models -- Opus 5.5, Astra, etc. juggle multiple tasks, delegate exceptionally well and are very good at working independently.
I still occasionally have issues with open-weight models, but the frontier labs have solved the above for most use cases.
- ilamont
the agent never stops and says, “Wait, this is something another model could do cheaper and faster.” It just plows on with the slow, expensive model. Conversely, the agent never says, “This model is too dumb for this task. Let me tag in a smarter one.”
This is a pretty big failing, which is compounded by the fact that most humans don't know which model to pick, or make assumptions based on Anthropic's hierarchy or "effort" involved.
Like Fable: your toughest challenges. You mean, like Fields Medal toughest challenges? Or analyzing and updating three monster spreadsheet toughest challenges? Or writing a new novel in the style of William Gibson toughest challenges?