This is article 5 in the series "The Art of Not Checking." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the introduction.
This time it is thinning out the tests. Thinning out means keeping the heavy tests out of the everyday run, and running them only on the day you touch that area. One day I received a report that said "the tests are passing." How many tests had not run at that point was written nowhere in the report. Work that makes things faster hides the blind spot — what you thinned out disappears from the report as well.
In article 1 of the first series, "The Art of Not Reading," I turned the test output into a summary of a few lines holding nothing but the counts and the failures. This time I add to that summary what is not being run.
Start by measuring just one thing.
How many seconds does a full run of your tests take?
And in a session with an AI, how many times does that full run happen?
On my machine the full run is 277 seconds and 13,950 tests.
Inside it are image processing that writes out real images, rate limiting that repeats real requests up to the threshold, an accuracy check that loads 260,000 rows of data — heavy tests, in quantity, that are not worth running on a day you have not touched them. So thinning out is itself the right call, and even left to the AI, it starts thinning out on its own: "I will run only the related tests."
The problem is born at the moment you thin them out. The existence of what did not run disappears from the report. The AI's report is "the tests are passing." It does not say what is not running. The side that reads takes a pass to mean "all of it passed." Work that makes things faster turns, just as it is, into work that hides the range you are not looking at.
What I understood fits in one sentence.
Work that makes things faster hides "what you are not looking at." In exchange for not running them, put what is not being run in front of you every time.
CC BY 4.0
The mechanism: in exchange for thinning out, always print the list of what was thinned out
Two things together make a single mechanism. First you give the heavy tests an area name, and make them run only when that area name is specified.
# on the side of the heavy test (example: the E2E for the payment flow)
setup do
skip "runs only with E2E=payment" unless ENV["E2E"].to_s.split(",").include?("payment")
end
And the wrapper from series 1, article 1 prints the list of areas that did not run at the end of the summary, every time. All it does is pick up the area names from the test code and match them against the ones specified this time.
# add at the end of the summary in scripts/test.sh (a shortened version: it assumes a single area is specified. the code on my machine handles several areas and a description of each area)
echo "[E2E] areas run: ${E2E:-none}"
echo "[E2E] the following areas were not run. Re-run with E2E='<area>' only when you touched them:"
grep -rhoE 'E2E=[a-z_]+' test/ | sed 's/E2E=//' | sort -u | grep -vx "${E2E:-__none__}" | sed 's/^/ - /'
The output on my machine looks like this.
[E2E] areas run: payment
[E2E] the following areas were not run. Re-run with E2E='<area>' only when you touched them:
- thumbnail generation (writes real images)
- rate limiting (repeats real requests up to the threshold)
- score posting → propagation to aggregates
- batch idempotency (run twice, same values)
…
On my machine, 22 entries line up in this "did not run" list. Every time. It may feel noisy, but these 22 lines are the body of this mechanism. The fact that things were thinned out does not appear in the report of the one who did the thinning (the AI). So I put it into a shape that always comes out in the tool's output, so that it cannot be hidden. It is the same thinking as "the line with the counts is the one line I do not cut," which I wrote in series 1, article 1, made bigger.
22 lines every time is waste, if you look at the tokens alone. Even so, it works differently from adding "watch out for forgetting to run the E2E" to your instruction file. A rule asks you to remember it and then keep it in force across all of your judgments — as the introduction showed, that is why it dilutes the more you add. What this list asks for is only that you match it up. If the area you touched is not on it, the AI and the human alike skim past and that is that. The only lines that need attention are the ones whose condition has come up. What you may leave in place permanently is only what you do not have to remember.
The effect shows up in speed. The everyday run goes round with all the heavy areas taken out, and only on the day I touch payments do I attach E2E=payment and let it run. I cannot go back to a life of waiting 277 seconds for the full run every time.
What I stopped running
The full run for peace of mind. The ritual of "I have not done anything, but let me run it all and see a pass" is gone. What my peace of mind rests on has moved from the total number of passes to a read-through of the list: "today I have not touched any of the 22 that did not run."
And this read-through is over in a few seconds, because the list comes out in front of me every time.
Caveat: "you touched it and forgot to run it" remains
This list makes a forgotten run visible, but it does not prevent one. On a day payments were touched, the AI can forget to attach E2E=payment.
On my machine, as the last net, I have put a check into the pre-commit hook for "did the E2E for the area you fixed run?" It is one instance of "put a guardrail at the entrance where a mistake costs a lot," which I did in series 1, article 4. Make it visible with the list, and catch it at the entrance. Defend it twice over, and a forgotten run no longer slips quietly through.
How to verify: three points, once each
Prerequisites
- You have picked 1 or more heavy tests, given them an area name, and put them into a shape where they run only when that area name is specified
- You have finished adding, at the end of the summary from the wrapper of series 1, article 1, the part that prints the list of areas that did not run
Time required: 10 minutes
Steps
- You run the tests as you normally do (without specifying an area name)
- You run them with an area name specified (like
E2E=<area name>)
Pass conditions (all of them have to hold)
- In step 1, that area name is on the list at the end of the summary
- In step 2, that heavy test actually runs (the number of runs goes up from step 1)
- In step 2, that area name is gone from the list at the end
If it does not pass
- The list stays empty and does not change — the summary is not picking up the area names from the test code. What is inside the list is "the rest, with the areas you specified subtracted," so please take an empty list as more likely to mean "it is not picking them up" than "everything ran"
Cleanup
- None needed. This check breaks nothing
What you confirmed is only that "what is not being run is visible." Being visible and not forgetting to run something are different things. The accident of forgetting to specify an area you touched is not something this list can prevent (the side that prevents it goes at the entrance to the commit).
Next time, the finale, "The art of not deciding on the spot."
End of CC BY 4.0