---
title: "The Art of Not Checking #3: The Art of Not Re-running Tests"
author: garplab
publisher: TypingTube
license: CC BY 4.0
license_url: https://creativecommons.org/licenses/by/4.0/
license_scope: 「CC BY 4.0」の印から始まる節（仕組み・検証手順・コード）。印の無い本文は著作権を留保
canonical: https://typing-tube.net/articles/en/tashikamenai-03-no-rerun
series: "確かめない技術"
language: en
---


> This is article 3 in the series "The Art of Not Checking." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the [introduction](https://typing-tube.net/articles/en/tashikamenai-intro).

This time it is re-running tests. Re-running means running the same tests again when not a single character of the code has changed. The AI does this often. "I will confirm that it passes." "Once more, to be safe." Every time, I waited to the end for something I knew would not come back different. That is because, inside the AI, confirming a result and re-running are not told apart.

In [article 1](https://typing-tube.net/articles/en/cd93e0c5cac832) of the first series, "The Art of Not Reading," I brought the entrance for running tests together into a single wrapper, kept the full text in a log, and had the screen show only a few lines of summary. This time I use that log I kept in place of re-running.

Start by counting just one thing.

In your most recent session with an AI coding agent, how many times did the tests run?

Of those, how many were runs that did not come right after a change to the code ("I will confirm that it passes," "once more, to be safe," "a final check before the report")?

The AI re-runs tests often. The reason is simple: the AI does not tell confirming a result apart from re-running. A human either remembers "it passed a moment ago" or scrolls the screen back to look again. The AI, once the result has flowed out of its context (or once the session has changed), takes re-running to be the only means it has of confirming.

Even when not a single character of the code has changed. "At a break point, confirm with the tests" is also a development behavior it has learned. One line added on top does not overturn this priority.

The cost is time. On my machine a full test run is 277.85 seconds. Series 1, article 1 turned the output into a few lines of summary, but the 4 and a half minutes has not come down by 1 second. 3 re-runs for confirmation, and about 14 minutes are gone on that alone. On top of that, while those 4 and a half minutes are running, a window opens for the accident where another session starts 1 more run (that is the subject of the next article).

There is one more thing that calls for a re-run. Narrow it down and the name disappears: the AI does not take the result of a run as it comes, it narrows it with `| grep` or `| tail` before reading. Counting the records on my machine on 2026-09-12, all 237 of the 237 full runs had been narrowed, and there was not 1 case of taking it as it came. What gets saved is 74 lines, but which test failed disappears from the narrowed summary. Of the 10 full runs that came back failing, the name of the failing test was left in 0 of them. With no name there is nothing to do but run it again, and the 4 and a half minutes above ride along on top. Narrowing is not itself the bad thing. In the same records, 2,362 of 2,506 partial runs were narrowed, and most of that is correct use. The only place the real harm was an order of magnitude bigger was the full run.

The numbers from my machine that appear in this series (the seconds for a full run, the number of tests, the number of areas that did not run) were measured on 2026-09-08. Measuring the same machine again on 2026-09-10 gave 307.80 seconds, 14,238 tests and 23 areas. For all 3 the mechanism is running the same way, and what grew is the number of tests and the time that came with it.

The numbers move day by day, so when you compare them, look at the date they were measured along with them.

What I understood fits in one sentence.

> **The AI does not tell confirming a result apart from re-running. If there is a place where the last result can be read, the re-run does not happen.**

## The mechanism: make a place where the last result can be read

The wrapper from series 1, article 1 leaves all of the output in `tmp/test_last.log` every time.

So what "confirming" needs is already there on your machine. What you add is 1 entrance for showing it again.

```bash
# append to scripts/test.sh. put the --last branch right after log= and mkdir
if [ "${1:-}" = "--last" ]; then
  if [ ! -f "$log" ]; then echo "No test run recorded yet" >&2; exit 1; fi
  echo "--- last summary (run at: $(cat "${log}.time" 2>/dev/null || echo unknown) / full log: ${log}) ---"
  grep -E "^Finished in " "$log" | tail -1
  grep -E "^[0-9]+ runs, " "$log" | tail -1
  exit 0   # runs nothing. it only reads back
fi
date '+%m/%d %H:%M' > "${log}.time"   # the side that runs writes the date and time here (just before bin/rails test)
```

And to the guidance text of the "hook that stops the bare command" I put down in series 1, article 1, I add just 1 line.

```
To see the result only: scripts/test.sh --last (prints the last summary without re-running)
```

This closes the path. The AI hits the bare test command "to confirm," the hook stops it and prints the guidance, the guidance has "if you only want to look, `--last`," the AI gets the summary in an instant and goes on to the next task.

I have added nothing to my instruction file. All I did was put the entrance for reading it back again right in front of the moment you want to confirm.

One design point that matters. The first line of `--last` always carries the date and time of the last run. It says "this is the result from 10:41" every time. Even if you slip and confirm with `--last` after changing the code, how old the time is lands in your eye.

## What I stopped waiting for

Test runs for the sake of confirming. I no longer watch 4 and a half minutes of progress output. The time spent waiting for something I know will not come back different is the biggest reduction across the series. What goes down is not the number of output lines but the time you actually wait.

## Caveat: `--last` is a shortcut only for when you have not changed anything

The blind spot is plain. If the code has changed and the AI confirms with `--last`, it reads an old result as the new one. Two defenses are enough. One is the date and time always on display, written above. The other is to present `--last` only as a replacement for re-running, not as a replacement for running: that is why the guidance text says "to see the result only." If you changed it, run it. If you did not, read it back. That distinction is the one thing that stays.

## How to verify: confirm that it comes back in an instant

**Prerequisites**

- The wrapper from series 1, article 1 is leaving the full text in `tmp/test_last.log`
- On top of it, you have finished adding this article's entrance for showing it again (`--last`) and the 1 line in the guidance text

**Time required**: 5 minutes

**Steps**

1. You have the AI run the tests once, normally. This is where `tmp/test_last.log` and the run's date and time get created
2. You hit `--last` yourself
3. You touch 1 file at random (`touch` is fine, so is adding 1 line), then hit `--last` once more

**Pass conditions** (all of them have to hold)

- In step 2, nothing runs (if the tests start running, `--last` has become a re-run)
- In step 2, the line with the same counts as last time comes back
- In step 2, `--last` comes back within 1 second
- In step 3, the run date and time on the first line is still the one from step 1

**If it does not pass**

- The tests started running in step 2 — `--last` is not reading the log, it is going through to a run
- No date and time in step 3 — this is the heart of it. Without the date and time, a `--last` after a change to the code looks like "the result as of now." Fix it into a shape that always prints on the first line

**Cleanup**

- Put the file you touched in step 3 back, if you changed it

What you confirmed is only that "reading it back is fast." `--last` is a shortcut for when you have not changed the code. If you changed it, run it; if you did not, read it back: this one distinction is not something the mechanism makes go away.

---

Next time, "The art of not reading false failures." **I do not read the failures that come up to tell them apart.** Even so, I have never missed a real failure.
