This is article 11 in the series "The Art of Not Telling." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the introduction.

An apology first. I could not write the article this title promises. I have the AI review, every single time. Since article 12 of "The Art of Not Reading" I have written again and again about the one sentence I always type at the end of a session, "Are there any gaps so far?", and that sentence is a review in itself. It was the same when I finished writing the previous series, "The Art of Not Listening to the AI's Opinions": I had the AI read the whole thing and bring out the gaps. The record is still there. The reason this article exists all the same is that I found the one thing a review does not show: whether the automated tests I had put down are spinning idle.

And still I am writing this article. When I laid out the points the review came back with, something caught my eye. Half of those points were of a kind that the automated tests already in place would have dropped, had the same thing been inside the repository. They were not dropped because those automated tests were not looking at the code inside an article's code blocks.

What I was short of was not a review. What I did not know was what the automated tests I had put down were not looking at. Once that was clear, what I had to write changed from "do not have the AI review" to "who inspects what you have put down."

That is what this article is about. Have the AI review and any number of points come back. How many of them the mechanisms already in place could have picked up, though, is something you do not know until you count.

In article 6 of the first series, "The Art of Not Reading," I wrote that you replace a rule with an automated test that fails when it is broken. It is the same article I drew on in article 1. Back then I looked at the instruction file's side. This time I put the automated tests I have put down under suspicion.

Start by recalling just one thing.

The points that came out the last time you had the AI review.

Of those, how many could an automated test already in your repository have picked up?

CC BY 4.0

What came back from the review of the previous series

When I had finished writing the previous series, I started up 5 subagents. 4 of them for the reader's experience, the technical side, resistance to criticism and consistency with the existing articles, and 1 that read the whole thing through with an editor's eye. The points that came back went like this.

  • 5 pieces of code in the articles did not run when copied and pasted (the shell reads angle brackets as a redirect / a call to a function that is not defined / a regular expression that goes straight past its target)
  • Someone else's argument was attributed wrongly. 1 of the arguments I quoted was a different claim by a different person, not the one I had named
  • Inconsistent notation (two spellings of the same word mixed in together), places where the numbering no longer lined up, names used inconsistently

Lay them out and sort them and they come to 2 kinds. The first, the 5 pieces of code in the articles that did not run, are of the kind the automated tests already in place would have dropped, had the same thing been inside app/. They were not dropped because the 26 automated tests look at the repository and not at the code inside an article's code blocks. The same string of characters fails if it sits in app/ and passes if it sits inside an article's code block. What lies outside the range the mechanisms look at was this close at hand.

The remaining 2 are different in nature. Neither the attribution of someone else's argument nor inconsistent notation is something a program can decide the truth of. It is the same line I drew in article 5 of "The Art of Not Listening to the AI's Opinions," where I wrote that less than half of the prohibitions can be handed to a program.

The repository's side gets reviewed first

I counted up the production project on my own machine, typingtube, a web service for practicing typing along with music videos on YouTube.

What is lookingHow many
Automated tests that check against the specifications and the principles26
Hooks that cut in on a command or a launch6
Checks in the pre-commit hook8
Ratchets watching how things grow8 entries
The full test suite14,149 cases (0 failures at the time I wrote this article)

There is not much a review can point out in code that has been through all of that. Naming, a prohibited column, a missing translation, a departure from the design: every one of them fails before the commit. Before you get to "have it review," the project has already become the reviewer.

The more it grows, the further this goes. The instruction file says "Ruby 3.3.x," but open .ruby-version and it is 3.3.6. The file is the more accurate of the two, and it already was on the day I wrote that line. The official command for reviewing an instruction file draws the same line: it says to cut what the model can derive from the codebase and to keep the pitfalls and the reasons. The more the code grows, the more falls on the "can be derived" side, and the less there is to say.

So it is a matter of order

"Not having the AI review" is not a matter of telling you not to have it review. It is a matter of order. What the mechanisms can drop, the mechanisms drop, and the review goes on what is left. Stop reviewing where there is no mechanism and you are looking at nothing at all.

The problem was that I had not grasped that range. I knew there were 26 of them. What those 26 were not looking at, I did not know until I counted.

What I understood fits in one sentence.

What range the mechanisms look at is something you do not know until you doubt it and try breaking it. So instead of asking for an inspection, you put down a mechanism that looks at whether the automated tests are spinning idle.

The mechanism: count the automated tests that are spinning idle, from outside

What you put down is 1 automated test. What it targets, though, is not the code but the automated tests themselves.

# test/reference/zero_target_guard_test.rb (a shortened version of the code on my machine)
SCANNING = /Dir\.glob|Dir\[|\.glob\(/          # it is scanning a set of files
MINIMUM  = /assert_operator[^\n]*:>=|MIN_[A-Z_]*\s*=|refute_empty/  # states that it fails on 0 targets
SCANNED_MIN = 20                                # ⚠️ fails if this automated test's own reading place breaks

test "checks that scan a set of files never go green on zero targets" do
  files = Dir.glob(DIR.join("*_test.rb")).reject { |p| p == __FILE__ }
  assert_operator files.size, :>=, SCANNED_MIN, "this check is looking at the wrong place"

  unguarded = files.filter_map do |path|
    body = File.read(path)
    next unless body.match?(SCANNING)
    next if body.match?(MINIMUM)
    File.basename(path)
  end

  assert_empty unguarded, "these checks scan files but never state a minimum. " \
               "they stay green on zero targets, so nobody notices when they break\n  - " + unguarded.join("\n  - ")
end

An automated test that scans returns the same pass for "there are no violations" and for "it is not looking at anything in the first place." Move a file and its targets go empty, and it spins idle in silence. On the first run after I put it down, it found 1.

There is an automated test that scans but never states a minimum count.

  • e2_pixel_art_not_in_core_jobs_test.rb

It was checking only that the directory existed, so it passed even when there was not 1 file inside. I added 1 line and closed the hole.

This shape carries one more meaning. Of the 37 articles across 3 series, 30 carry a "How to verify." 13 of those were this same shape.

ArticleHow to verify
Article 4 of the first seriesTake it away, and confirm that it stops
Article 6 of the first seriesBreak it, and confirm that it fails
Article 1Add a request with no machinery behind it, and confirm that it fails
Article 6Drop a heading from the translation, and confirm that it fails
Article 9Name a file that exists but carries no marker, and confirm that it fails

What those 13 were doing was breaking the mechanism you had put down. I had simply never given it a name. The automated test in this article puts that into a shape where it comes into view once, without a human having to do it.

What I stopped having it do

"Review the whole thing." I have not stopped reviewing — for the previous series I had 5 subagents read it. What I stopped was asking for a range that reaches as far as what the mechanisms can drop. Now I cut the range down before I ask.

There were 3 titles I could not write

As I apologized at the top, this article could not be written as it stood. There are 3 titles I put up as candidates for the series and dropped.

CandidateIn or outWhat dropped it / let it through
The art of not writing plans❌docs/plan/ held 200 plan documents
The art of not asking for a plan✅"Write a plan" appeared 0 times on the instruction file's side (article 4)
The art of not having the AI review❌ → ✅The record of the previous series still held a record of me having it review. I changed what was inside and let it through (this article)

For all 3, what decided it was the record and not my memory.

And on 2 of them I had believed the opposite. "Surely I am not writing plans." "Surely the review is just that one sentence every time." What got proved was not that "the mechanisms work."

When there is a record, it is the assumption that fails.

Caveat: an automated test that has never failed has proved nothing yet

One, this automated test looks only at how things are written. All it does is look statically at whether a minimum count is stated, so write >= 0 and it passes. Whether the number is a reasonable one is not something it can measure (the same blind spot for the 4th time, counting from article 1).

Two, the number 26 is not a boast. On a small project 2 will do. What I want you to look at is whether you can say what those 2 are not looking at.

Three, an automated test that has never failed has proved nothing yet. As article 6 of the first series said, an automated test that is guarding something and an automated test that is broken cannot be told apart until you have seen it fail once. The automated test in this article, too, was only known to be working once it had found 1.

How to verify: break the automated tests' side, from 2 directions

Prerequisites

  • You have several automated tests that scan, and each of them states that the scan found N targets or more
  • You have finished putting down the meta test that counts them from outside (an automated test that looks at automated tests)
  • You start from a state with no work in progress (git status is empty)

Time required: 15 minutes

The automated test in this article can break in the same way as the ones it looks at, so there are 2 directions to confirm from.

Steps

  1. You break the side being looked at. From any 1 of the automated tests that scan, you delete the 1 line that states the minimum count (the line that corresponds to assert_operator files.size, :>=, 1). You run the meta test and note the result
  2. You put back the 1 line you deleted and confirm that the meta test goes back to passing
  3. This time you break the meta test itself. You change the name of the directory the meta test scans to a name that does not exist. You run the meta test and note the result

Pass conditions (all of them have to hold)

  • In step 1, the meta test fails
  • The message in step 1 shows the name of the file whose statement you deleted
  • In step 3, the meta test fails on itself (something of the form Expected 0 to be >= 20)

If it does not pass

  • In step 1 all that comes out is "one of them is missing," with no name — whoever fixes it (the AI, or you) ends up hunting through the several automated tests
  • It does not fail in step 3 — this is the heart of it. The automated test that hunts for automated tests that pass on 0 targets is itself passing on 0 targets. This meta test has been looking at nothing since the day it was put down

Cleanup

  • You put the directory name back and confirm that it goes back to passing and that git diff is empty

Next time is the finale, "The art of doing nothing." Over 11 articles I have cut down the words I hand over. In this article alone, I have miscounted my own machine 2 times.


Series: The Art of Not Telling

CC BY 4.0 はここまで