This is article 4 in the series "The Art of Not Checking." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the introduction.
This time it is false failures. A false failure means a test that fails because of the way it was run, even though the code is not broken. One day I was handed a screen with a long row of errors I did not recognize, on code where I should not have broken anything.
In article 1 of the first series, "The Art of Not Reading," I brought the entrance for running tests together into a single wrapper and had a hook stop the bare command. This time I add 1 lock to that entrance.
Only if you have an environment where it reproduces, try one thing.
Open two terminals and run the same test suite in both at once (if your setup shares a test DB or fixtures, it reproduces locally. When you are done, rebuild the test DB to put it back — with Rails, bin/rails db:test:prepare).
What comes out is a pile of errors you have never seen. The fixture load deadlocks, the shared cache and Redis get crossed, and the 2 runs trample each other's data. On the day it actually happened on my machine, it was 7 failures / 53 errors. Whether a real failure was mixed in among those 60 — to sort that out, I ended up re-running the 277-second full run 3 times.
In the old days this was settled with "who would ever run 2 at the same time." Now it is different. Running in parallel is a standard accident of the age of AI. A human runs several sessions in parallel. A subagent helpfully runs the tests. As we saw last time, the AI sends 1 more run through "to confirm."
And the AI that ran it does not know about any run but its own. What the AI does next, once it has seen the pile of errors, is read the false positives (the false failures) 1 by 1 and go off to fix them.
Code that was working gets broken by fixes for bugs that do not exist.
What I understood fits in one sentence.
When 2 test runs go at once, the real failure is buried in a pile of false ones. Rather than reading them to tell them apart, do not let the 2nd one run.
CC BY 4.0
The mechanism: a lock that lets only 1 run through
I make no effort to tell them apart. I do not build a tool that reads the pile of errors cleverly either. I get rid of the condition that produces the pile instead. At the entrance for running tests (the wrapper from series 1, article 1) I take a lock, and if it cannot be taken, nothing runs.
# add this to scripts/test.sh just before the run (after mkdir -p tmp)
lock=tmp/test_running.lock
if ! mkdir "$lock" 2>/dev/null; then
owner=$(cat "$lock/pid" 2>/dev/null)
if [ -n "$owner" ] && kill -0 "$owner" 2>/dev/null; then
echo "Another test run is in progress (started $(cat "$lock/started" 2>/dev/null) / PID ${owner})." >&2
echo " Wait for it to finish. To see the result only: scripts/test.sh --last" >&2
else
echo "Stale lock (its owner PID ${owner:-unknown} is gone). Remove it with rm -r $lock and run again" >&2
fi
exit 1
fi
date '+%m/%d %H:%M' > "$lock/started"
echo $$ > "$lock/pid"
trap 'rm -rf "$lock"' EXIT
The 2nd run stops before it starts. While a run is in progress, the only paths the guidance offers are the ones that do not remove the lock. Wait, or read the last result with --last. Only when the process that owns the lock is already gone does the guidance say how to remove it. The AI gets to go on only once it has read this guidance. Here too, the shape is one where it cannot move unless it has read.
The path for removing it is not decoration. If a process dies abnormally, the lock stays behind. Unless you have a way out ready for that moment, someone who finds the guardrail in the way will break the guardrail along with it.
But I show the way out selectively. I check whether the PID written in the lock is still alive, and print how to remove it only when the owner is gone. That is because, if you put the release command in front of a live lock, an AI in a hurry will use it. The lock goes on being trusted because the proper way out appears only under the right condition.
Two side notes, briefly.
Do not have your subagents run the tests.
The lock lets only 1 through, so it fights the main side and one of the two is bound to stop, and the work on the stopped side is wasted along with everything it had built up to that point (how subagents are handled is what article 5 of the first series covered). And parallel runs are not the only source of false failures. There are fakes the environment makes, such as a run late at night that crosses the date. The principle is the same: rather than reading them to tell them apart, you put it on the side that removes the condition, or detects it and stops.
What stopped being produced
The pile of errors. The work of sorting out "which of these 53 is real" is gone, because the pile stopped being produced. The job of stepping in to stop an AI that reads false failures and goes off to fix bugs that do not exist is gone too.
Caveat: what the lock sees is only the runs that came through this entrance
What stops is only a 2nd run that started through the wrapper. Hit the bare command directly and it does not pass the place where the lock is. So this article closes only as a set with the "hook that stops the bare command" put down in series 1, article 1. In an environment where the entrance has not been brought down to 1, the lock does no more than "only the well-behaved runs do not line up."
Runs from another working directory (a separate clone of the same repository) or from CI are outside this lock as well. The lock is 1 file, and it lives only inside the place where you put it.
If there are other routes that share the same test database, count those separately.
How to verify: confirm that the 2nd run stops
Prerequisites
- You have finished building the lock into the entrance for running tests (the wrapper from series 1, article 1)
- You can open 2 terminals
- The 1st run takes at least tens of seconds to finish — with a test that ends in an instant, the 1st run is over before you start the 2nd
Time required: 10 minutes
Steps
- In the 1st terminal, you have the AI run the tests once
- While it is running, from the 2nd terminal, you start a 2nd run with your own hand. Rather than asking the AI, type it yourself. What you want to confirm is how the lock works, not the AI's judgment
- After waiting for the 1st run to finish, you start the 2nd run once more
Pass conditions (all of them have to hold)
- In step 2, the 2nd run stops before it runs any tests
- In step 2, where it stopped says that there is another run in progress
- In step 2, a way out is offered: wait, or read the last result with
--last - In step 3, it starts normally without stopping
If it does not pass
- In step 2 the guidance went as far as how to remove the lock — the way the guidance is selected is broken. Put the way to release it in front of a live lock and an AI in a hurry will use it. How to remove it may appear only when the process that took the lock is already gone
- In step 2 the 2 runs went side by side — the lock is outside the entrance. Move it to a position where it is taken before the tests start
Cleanup
- You confirm that no lock file is left behind. If one is left, all your test runs after that stop
That it stops when it should stop, and that it passes when it should pass: both, once each.
Next time, "The art of not running everything." I do not run all of the tests. Even so, the parts I am not running come up in front of me every time.
End of CC BY 4.0