Back to blog
Verification
Superpowers
Claude Code

The One Rule That Stops Agents from Inventing Tests

Thắng Đoàn
Thắng Đoàn

An agent writes a function, then writes a test for it. The test passes immediately. It proves nothing: it came from the same context as the code, so it is biased toward making itself look correct and misses the same edge cases. That green checkmark is the most expensive false positive in AI-assisted engineering.

The iron law

No production code without a failing test first. If code exists before the test, delete it. Do not keep it as reference, do not adapt it, do not look at it. Agents will resist: "I already wrote the implementation, let me just write the test now." The answer is no. A test written after biased code is biased to match.

Red: one minimal failing test

Write one test for one behavior; if the name contains "and", split it. Run it and read the failure. It must fail because the feature is missing, not a typo or setup error. Watching the test fail is the step agents skip: they see red, assume the test is broken, and "fix" the assertion until it is meaningless.

Green: the simplest code

Write the simplest code that passes. Not the most elegant, not the most general, not the code that guesses the next three requirements. When the agent wants a parameter for "future flexibility": no. The test does not require it.

Refactor: only after green

Remove duplication, improve names, extract helpers, keeping tests green throughout. Agents skip this because work feels done at green; the result is code that works but rots. Refactor is what makes the next change cheap.

The rationalizations

"Tests after": they pass immediately, which proves nothing about testing the right thing. "Too simple to test": simple code breaks too, and the debug session costs hours. "Deleting work is wasteful": sunk cost; the question is whether you also ship the debt. "Coverage is coverage": coverage measures lines executed, not behavior verified.

The agent-specific failure

The agent writes the implementation, then a test that calls it with one obvious input and asserts the obvious output. It confirms the function does the obvious thing in the obvious case; it never touches null, empty, large, or malformed input, because function and test share one context. The fix is the iron law: test first, watch it fail, then minimal code.

The trade-off

For a constant change, a logic-free UI tweak, or a pure refactor, TDD is theater: the test costs more than the bug, and existing tests already prove the behavior. The judgment is whether the change introduces behavior. If yes, test first. If no, skip the ceremony. Applying TDD to everything burns the team; applying it to nothing ships bugs.

What to do today

Ask the agent on your current PR whether the test or the code came first. If the answer is the code, make it redo the test.

Share:

Recommended for you

Enjoyed this article?

Subscribe for new articles. No spam. Unsubscribe anytime.

By subscribing you agree to receive the newsletter. No spam, and you can unsubscribe anytime.