Why Your AI Agent Says "Done" When the Work Isn't
You ask an agent to add a feature. It writes the code, runs something, and says: done, the feature works.
You move on. Two days later you find problems. The function it added is never called. The test file it claimed to update does not exist. The build is broken.
The agent was not lying. From its view, the work looked done. It wrote code. It ran a command. The output contained the word pass. So it reported success.
This is the most common failure in AI-assisted engineering. It is also the most expensive one, because nothing in your workflow catches it. The agent does not re-check. You trust the report. The bug ships.
The real problem: a missing step
Agents are not lazy or careless. They are missing a step. Between "I ran a command" and "the work is complete" there is a verification gap nobody told them to fill.
Look at what an untrained agent does when it claims success:
It says "tests pass" without showing the test output. It says "build succeeds" from reading log text, not the exit code. It says "bug fixed" without re-running the steps that showed the bug. It says "requirements met" by listing tests it wrote, not by checking your original ask.
Each omission is small. Together they produce a steady stream of work that looks complete and is not.
The rule that closes the gap
The fix is one rule, applied strictly: no completion claim without fresh verification evidence.
"Tests pass" means: show the test command output, run in this message, with zero failures. "Build succeeds" means: exit code zero, not "logs look good". "Bug fixed" means: run the original reproduction and watch it pass.
This sounds obvious. In practice, agents resist it. They prefer to say "should work now". The discipline must be enforced as a workflow, not suggested as a preference.
That is what Superpowers is. It is a set of skills that enforce the boring steps agents skip: brainstorm before code, plan before implement, test before pass, verify before claim, review before merge. None of the skills are clever. All of them are gates the agent cannot route around.
When the discipline is too much
For a one-line typo fix, full TDD plus two-stage review plus verification is theater. The ceremony costs more than the bug.
The judgment call is risk, not size. A typo in a README needs no verification. A change to a payment flow does. Calibrate the discipline to the blast radius, not to the line count.
What to do today
You do not need Superpowers to fix this. You need one rule your agent cannot override: every completion claim includes the exact command to check it, plus the output of running it now. Not a paraphrase. The actual output.
If your agent resists, that resistance is the signal. The work it was reporting as done was never verified.
Recommended for you
- VerificationSuperpowersClaude Code
The One Rule That Stops Agents from Inventing Tests
An agent writes code, then writes a test for it. The test passes right away and proves nothing. Fix: no production code without a failing test first.
- VerificationSuperpowersClaude Code
Why Your Agent's First Fix Attempt Is Usually Wrong
An agent hits a bug, guesses a fix, reports done. The bug returns because the fix treated a symptom. Fix: no fixes without root cause.
- VerificationSuperpowersClaude Code
Why Your Agent's 'Done' Cannot Be Trusted Without Command Output
The most expensive AI failure: the agent reports success without verifying. Fix: no completion claim without fresh command output.
Enjoyed this article?
Subscribe for new articles. No spam. Unsubscribe anytime.