Why Your Agent's First Fix Attempt Is Usually Wrong
An agent hits a bug, guesses a fix, tests pass, done. Two hours later the bug returns in a different form, because the first fix treated a symptom. The second attempt is worse: another guess layered on the first wrong fix, and the real cause is buried under patches.
The iron law
No fixes without root cause investigation first. Time pressure makes guessing feel efficient, and every green checkmark from a guessed fix is a false signal. The bug is still there.
Investigate, then find the pattern
Four moves before any fix. Read the error message carefully; it often names the file, line, and failed assumption that agents skim past. Reproduce the issue the same way every time; a bug that appears only sometimes is environmental, not code. Check recent changes: most bugs arrive with a diff. For multi-component systems, log what enters and exits each boundary until the failing layer is visible. Then find working code in the same codebase that looks like the broken code and list every difference; the small ones are usually the cause.
One hypothesis at a time
Write down a single specific hypothesis: "X is the root cause because Y." Test it with the smallest possible change touching one variable. Never stack fixes: fix A then B then C means that when the bug disappears you cannot know which fix worked, and the next regression is untraceable.
Implement with a failing test
Create a failing test that reproduces the bug, then one fix addressing the root cause that makes it pass without breaking others. If the fix fails three times, stop: the architecture is wrong, not the fix, and continuing to guess inside it compounds the damage.
Red flags that mean stop
"Quick fix for now, investigate later": the later never comes. "Just try changing X": guessing. "Add multiple changes, run tests": stacked fixes. "I do not fully understand but this might work": if you do not understand, you cannot fix. All of them mean the agent stopped investigating and started guessing.
The trade-off
For a typo or a CSS class mismatch, investigation is overkill: read the error, fix it, move on. The judgment is hidden depth. An intermittent test failure, a production-only bug, or a bug that disappears when you add a print statement all have it. Match the investigation to the mystery.
What to do today
Next time an agent proposes a fix within a minute of seeing the bug, send it back to phase one.
Recommended for you
- VerificationSuperpowersClaude Code
The One Rule That Stops Agents from Inventing Tests
An agent writes code, then writes a test for it. The test passes right away and proves nothing. Fix: no production code without a failing test first.
- VerificationSuperpowersClaude Code
Why Your AI Agent Says "Done" When the Work Isn't
The most expensive AI bug is not bad code. It is the agent claiming the work is done, and you trusting the claim without checking.
- VerificationSuperpowersClaude Code
Why Your Agent's 'Done' Cannot Be Trusted Without Command Output
The most expensive AI failure: the agent reports success without verifying. Fix: no completion claim without fresh command output.
Enjoyed this article?
Subscribe for new articles. No spam. Unsubscribe anytime.