The Real Problem
Picture this: an AI agent generates 100 tests for a new checkout flow. All 100 pass. Coverage reports say 91 percent. The team ships. Two weeks later, a discount code stacking bug reaches production, one that a human tester would have caught by asking "what if someone applies two codes?" None of the 100 generated tests asked that question, because nothing in the code or the prompt suggested it should be asked.
This isn't a hypothetical edge case. It's the default failure mode of AI-generated tests: they are extremely good at covering what the code does, and structurally weak at covering what the code should do when something unexpected happens. A test suite full of assertions that mirror the implementation will pass even when the implementation is wrong, because the test and the bug were generated from the same flawed understanding of the requirement.