An agent writes code, then writes a test for it. The test passes right away and proves nothing. Fix: no production code without a failing test first.
Verification
Proving the AI's work is correct — TDD, systematic debugging, no-completion-claim-without-evidence.
An agent hits a bug, guesses a fix, reports done. The bug returns because the fix treated a symptom. Fix: no fixes without root cause.
The most expensive AI bug is not bad code. It is the agent claiming the work is done, and you trusting the claim without checking.
The most expensive AI failure: the agent reports success without verifying. Fix: no completion claim without fresh command output.
