This is article 3 in the series "The Art of Not Reading." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the introduction.
This time it is the unit tests the AI writes. The AI writes the implementation and the tests at the same time, so the tests become a copy of the implementation. Reading them does not tell you whether they are a copy (as a piece of writing, a copy looks correct). This article checks whether the tests that pass are really guarding anything, without reading the tests.
Start by trying just one thing.
Open an implementation whose AI-written tests are passing, and deliberately break 1 line of that implementation. Shift a boundary value by 1, invert a condition, rewrite a constant. Then run the tests (once you have seen the result, always put that 1 line back).
If it fails, that is normal. If it passes, the test was never looking at anything in the first place.
Some of the green check marks in your project are in that state.
The reason it passes is a structure specific to the AI. The AI writes the implementation and the tests at the same time.
Tests written at the same time become a copy of the implementation. By that I mean tests whose expected values are taken not from the specification, but from what the running implementation actually returns. A test vouching for its own author's work, something a human team used to prevent through review, becomes the standard mode of production with an AI.
The production project on my own machine, typingtube, a web service for practicing typing along with music videos on YouTube, carries 26 places in its test comments where this method found a test that passed when it should not have, and fixed it. "Rewriting the capital-city data in the learning decks (the feature for practicing typing with word decks) to a different city left it passing." "When one of the 2 checks was broken, the other one hid it and nobody noticed." Not one of them was found by reading.
What I understood fits in one sentence.
The AI writes the implementation and the tests at the same time, so the tests become a copy of the implementation. Whether they are a copy is something you only learn by breaking it.
CC BY 4.0
Why the tests become a copy
The expected values in a test come from one of two places. Either from the specification — what you decide is correct — or from what the implementation actually returns. The first has a basis separate from the implementation, so it fails when the implementation is wrong. The second has no basis beyond the implementation itself. Even when the implementation is wrong, the expected value has been fitted to that same mistake, so it passes.
A copy is the second kind. What is missing here is not the test's coverage of cases. The AI lines up boundary values and error paths as items. The only thing missing is where the expected values come from, and into each one of them goes what the implementation actually returned: the output pasted into an assertion, the same calculation repeated inside a mock. It looks like it is guarding something, yet it holds nothing in place. The next time the implementation is rewritten, the AI rewrites the tests along with it.
The AI leans that way because it writes the implementation and then, in the same request, writes the tests. By the time it writes them, the implementation is sitting in the same context. The nearest material for an expected value is the implementation it just wrote itself. The specification is either not written down, or written down dozens of turns earlier in the conversation. Liu et al. reported in 2023 that information placed at the beginning and the end of a long input can be retrieved, while the accuracy for what sits between them drops significantly. The implementation is at the end; the specification is buried in between.
The side that builds and the side that checks are also the same. The Claude Code documentation, at the end of its list of ways to give the AI a means of verification, says to make sure the agent that did the work is not the agent that grades it.
The official prompting guidance has a section about tests. Claude can focus too much on making tests pass at the expense of a more general solution. The wording it recommends is "tests are there to verify correctness, not to define the solution." Once passing the tests is the goal, pasting the implementation's output straight into an assertion becomes the shortest route. The research-side name for this is specification gaming, defined as behavior that satisfies the literal specification of an objective without achieving the intended outcome. A copied test satisfies the literal condition of passing.
So what finds a copy? Handing over a list of test cases in advance does not prevent one. The cases were never what went missing. Huang et al. reported in 2023 that LLMs cannot correct their own mistakes without external feedback, and sometimes make them worse by trying. Breaking 1 line of the implementation and running the tests is the act of creating that external feedback. It fails, or it passes. What comes back is one of those two, and it does not pass through anything the AI writes.
An answer to "mutation testing is slow and not worth it"
What you just did has a name: mutation testing. This technique carries the assessment that it works, but that generating every mutation mechanically and running them all is slow and does not pay off for most projects. That assessment is correct, and two of its premises are different here.
The target is different: not measuring how thoroughly human-written tests cover the code, but detecting where AI-written tests vouch for their own author's work. The method is different: instead of generating every mutation with a dedicated tool, you have the AI itself pick just a few places and inject them there. The one that knows best which places would hurt if they broke is the AI that wrote the implementation. The AI picks them, but what decides whether they failed or passed is the return value of the tests.
The mechanism: put down a single policy file
This is all you put down (a reduced version of what I run for real).
# The policy for injecting mutations (read this when asked to inspect the tests)
Purpose: find out whether the tests the AI wrote have become "a copy of the implementation."
Mutation = deliberately breaking exactly 1 place in the implementation. A test that passes anyway is not looking at that place.
When to apply = only on a round that changed the implementation or the tests. Not on a round that only changed data or settings (the result comes out the same as the previous round).
1. Pick lines for mutation where a test ought to fail. Prefer boundary values, conditional branches, formulas
(breaking a line no test is supposed to reach and still passing is not a hole in the tests)
2. Break 1 place → run the tests → record the result → always put it back (check that git diff is
empty again). Repeat for a few places
3. Make the first one a mutation that is certain to be caught. If they all pass, suspect the instrument first
4. A caught mutation guarantees nothing. Report the details only for mutations that slipped through
(for the caught ones, the count alone)
5. A test that let a mutation through gets rewritten from the specification, not from the implementation
To use it, you just ask: "read this policy and try injecting mutations into the tests you wrote today."
The AI decides where and how many to inject. Items 1 and 3 of the policy are pitfalls the AI actually fell into. Once it broke a line that is correctly never reached by design and made a fuss that the tests were deficient; another time everything it broke stayed passing, because the tests were not running at all.
Item 2's "always put it back" is the one line here you cannot drop. While a mutation is in place, the deliberately broken implementation is live where it sits. If a commit or the work of a session running in parallel goes through before you have put everything back, the defect you introduced on purpose travels onward.
"I broke 5 places and 5 were caught" is not a result. A caught mutation guarantees nothing: even a test that is a copy catches mutations inside the range it copied. Only the mutations that slipped through mean anything, and 1 of them pays for this policy file. The one counting that 5 is also the AI that injected the mutations.
What I stopped reading
The body of the tests the AI writes.
I replaced the work of judging whether something is properly tested by reading the test code with the work of breaking it and checking. Reading eyes are fooled by a copy; mutations are not.
Apply it only on a round that changed the implementation or the tests
If you ask for mutations every single time, they start landing on rounds that do not need them. They come in on rounds where neither the implementation nor the tests were touched: rounds where all that changed was the data or the settings the tests read.
Nothing is gained when a mutation is caught there. Since neither the implementation nor the tests have moved since the previous round, the result is the same as the last time you measured. The "when to apply" line in the policy file stops this.
Caveat: all a mutation answers is whether it is a copy
If the test is not looking at the line you broke, the test passes. If it is looking, the test fails. Whether it is a copy is decided right there. There is no room to judge by reading. Oversights caused by a copy all come out this way, for the places you injected mutations into.
What does not come out is where there is no test at all. A process that was never written into the implementation has no line to break. A gap in the tests comes from a different cause than a copy. The amount that fits in the context is fixed, so cases raised in the middle of the work get pushed out by what comes in afterward. A case that was pushed out appears neither in the implementation nor in the tests. With nothing to break, injecting mutations does not find it. Working against copies and guaranteeing that the tests are complete are two different things.
On top of that, a few mutations are a sample. A mutation that slips through gives you a test you can certainly fix, but everything being caught does not make the quality sound. The AI's report ends with "all 5 places were caught." A mutation is caught even by a test that merely copied the implementation. The moment you take that 1 line as evidence of quality, this policy file turns into false reassurance.
How to verify: check the instrument itself first
Prerequisites
- The policy file from this article is in place
- You have 1 or more files of AI-written tests, and all of them pass right now
- You start from a state with no work in progress (
git statusis empty), so that what you revert in the cleanup does not get mixed with your own edits
Time required: 15 minutes
Steps
- You ask the AI to inject mutations into 1 existing test file: "read this policy and inject mutations into the implementation that
<test file name>guards." The AI decides where and how many - You read the report in 3 parts: the count of mutations that were caught, the count that slipped through, and the substance of the ones that slipped through (what was broken)
Pass conditions (all of them have to hold)
- 1 or more mutations were caught
- 0 mutations slipped through
If it does not pass
- 1 or more mutations slipped through — that is the harvest. That test is not looking at the place you broke. Have the AI rewrite it from the specification instead of from the implementation
- 0 mutations were caught (everything passed) — before you doubt the quality of the mutations, doubt whether the tests are running at all. A wrong file name, or that test sitting outside the target of the run, comes first
Cleanup
- Put every mutated implementation back. Confirm with your own eyes that
git diffis empty - This is the most important part. If a commit or the work of another session goes through before you have put everything back, the defect you put in on purpose travels onward
Next time, the art of not using skills. I stopped registering skills with the AI. And the AI still has command of a great many "skills."
Series: The Art of Not Reading
- ← Previous: 2. The art of not reading memory
- → Next: 4. The art of not using skills
- All articles: Introduction: I Barely Read What the AI Outputs Anymore
CC BY 4.0 はここまで