This is article 5 in the series "The Art of Not Listening to the AI's Opinions." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the introduction.

This time it is about the prohibitions for which no automated test or hook has been written. Prohibitions are the list of rules you decided not to allow while developing. I do write automated tests. What I do not write them for is the side of the prohibitions whose truth a program cannot decide. My list has 17 entries. Of those, 4 are ones I can say with confidence an automated test is watching, and for the remaining 13 a program cannot decide the truth. For those 13, I settled on writing down, right beside each one, what I decided not to make a machine check.

In article 6 of the previous series, "The Art of Not Reading," I wrote that you replace a rule with an automated test that fails when it is broken. This time it is about the prohibitions that could not be replaced.

Start by checking just one thing.

Pick one prohibition and answer what would happen if it were being broken at this very moment.

Would an automated test fail?

Would CI stop?

Would anyone notice?

If none of those, that prohibition is only written down.

CC BY 4.0

"Add translations to every locale" was written and 5,525 were missing

The production project on my own machine (typingtube, a web service for practicing typing along with music videos on YouTube) carried a behavior rule like this: "Do not settle for a fallback; add translations to every locale." The service ships in more than ten languages, so the policy is obvious. The wording left no ambiguity, and I had been handing it to the AI all along.

There was no automated test, so if the AI broke it, nobody would have noticed.

When I counted, 5,525 translations were missing (measured 2026-08-31). Of the 755 translation files, 241 had gaps, and 2 were absent altogether. Of the missing ones, 31 were places where the key name itself goes on the screen when the translation is not there.

What hit hardest when I went back over the count was the breakdown. Many places had Japanese and English alone filled in. Every time wording was added it stopped at 2 languages, and that shape had repeated over and over.

It is not that the AI had failed to read the rule. It read it, and every time a translation was added, the rule lost. What won was getting the wording in front of it through in Japanese and English.

Whether the translation rule could be turned into an automated test was something I had not decided until then. A behavior rule has no place to write whether an automated test exists. A prohibition nobody has decided on and one that has been decided are the same single line there.

This 1 entry, the translation, I moved into a pre-commit hook. Whether the translation keys line up across the languages is something a hook can count.

I tried to turn them all into automated tests, and could not

In that case, just turn them all into automated tests.

That is what I thought too, so I went through the 17 prohibitions one at a time.

The ones I could turn into automated tests came to 4.

One of the remaining 13 is rejecting cost cuts for the character feature. It is the one whose reason I wrote out in 3 lines in article 2. That the cost figure changed is something a program can count. But what is prohibited is not the figure changing; it is letting a proposal to cut the cost through. To decide "is this economic balance as intended?" you have to read down to the reason. Whether a character a beginner set up on a whim breaks other people's experience. Comparing figures does not settle that.

There are other prohibitions that cannot be decided. The policy on how the service makes money, the judgment not to recommend levels, requiring evidence whenever SEO text is changed. They are all the same: there is no state that counting would settle.

So more than half the prohibitions are ones a human judges.

When I finished and went back to the list, one more thing became clear. Among the 13 I had decided could not be automated tests, some entries read that way and some did not. The ones that did not read that way were the same wording as before I looked. What I had decided existed only in my head. From the outside, they looked the same as the translation rule at the top, which piled up 5,525 gaps while nobody had decided.

What I understood fits in one sentence.

A prohibition that cannot be an automated test, even once you have decided so, cannot be told apart from one nobody has decided on unless you write it down. So what you decide not to make an automated test gets written down beside the prohibition.

Why a prohibition kept off automated tests looks undecided until written down

What automated tests and hooks look at is the state inside files. For the translation hook, whether a key present in Japanese is also present in the other languages' files. Present or absent can be counted, so a program can decide the truth.

No program counts a line that is written down. Whether to follow it is decided by the AI. The Claude Code documentation writes that instruction files and memory are context, not enforced configuration. The same page says the AI reads them and tries to follow, but there is no guarantee of strict compliance, especially for vague or conflicting instructions. Whether to follow the line that arrived is decided by the AI on the spot, as it adds the translation, together with everything else on hand. Even if the odds of losing in a single judgment are small, the same judgment is repeated every time wording is added. Something that loses by a little every time piles up with nobody noticing. 5,525 is the number that piled up that way.

Move it into a hook and that judgment disappears. The hook documentation writes that hooks give you deterministic control. Rather than relying on the LLM choosing to run it, the action always happens. The moment the translation entry moved into a hook, it left the question of whether to follow.

For cost cuts, there is no state that counting would settle. The figure not having moved and a cost-cut proposal not having been let through are two different things. Turning a prohibition into an automated test means replacing that prohibition with a state that can be counted. What gets kept after the replacement is not the original prohibition but the state being counted.

The color literals in article 3 were the live example. On the prohibition against writing a color directly, like #fff, I had put a hook and was counting occurrences. The hook counted only CSS files, and the scope of the prohibition became just that. If the state being counted overlaps with what the prohibition means, the replacement is fine. Whether translation keys line up across languages overlaps. Replace the cost prohibition with a comparison of figures, and you get a different prohibition: it passes as long as the figure does not move. So I do not make them automated tests. A prohibition I decided not to make an automated test stays as the written line, left to the AI's judgment.

And the decision itself, not to make it an automated test, is kept nowhere unless it is written down. The Claude Code documentation writes that sessions are independent, and a new session does not carry the conversation history of the previous ones. What carries over to the next session is only what is on disk. What I decided in my head does not reach the AI in the next session. What reaches it is the wording of the list, and a decided entry and an undecided one have the same wording. What I see half a year from now is the same wording too.

The mechanism: "none" plus a reason under a "Machine check" heading

The single sheet from article 1 carries a "Machine check" heading on every entry. It is the place to write whether an automated test or a hook is watching that prohibition. What this article writes is what goes under it.

### NG3 Do not give user-submitted content a name field

**Conclusion**: submitted content such as characters, pixel art and backgrounds
carries no name column. They are identified by their visuals and their settings.

**Reason**: (why it is prohibited. That is the subject of article 2)

**Scope**: adding columns such as `name` / `title` / `nickname` to the 4 models is prohibited

**Related memory**: `feedback_no_naming_for_ugc.md`

**Machine check**: yes. `a1_name_field_absence_test.rb` checks from the schema side
that the 4 models carry no such column

It is that last line. Where there is an automated test, you write which automated test it is. Where there is none, you write this.

**Machine check**: none (detecting the recommendation UI mechanically cuts across implementation domains, so it is covered by review)
**Machine check**: none (a process rule, covered by review)
**Machine check**: none (it is about how the clause is worded, so it is looked at in review when one is added or revised)

Writing nothing and writing "none" are not the same. On an entry with nothing written, you cannot tell "not looked at yet" from "looked at, and judged beyond a program." When you cannot tell, that entry hangs in the air forever. It gets reconsidered only when somebody happens to remember, with a "come to think of it, could we write an automated test for this," and mostly nobody remembers. The moment you write "none" and the reason for it, it becomes a settled entry. A settled entry is one you decided a human would look at. It is no longer a prohibition that is only written down.

My list has 17 entries, and the breakdown goes like this. 4 say "yes" and 10 say "none." For 2, an automated test watches only part of it, and I cannot say either way. And 1 has no such heading at all.

One of the "none" entries carries a test name. It is the prohibition with 3 axes of judgment, the one article 3 dealt with. There is an automated test that looks at the boundary, but the automated test does not hold the whole of that prohibition, so it sits on the "none" side.

The 1 entry with no heading at all I noticed only while counting as I wrote this article. Decide where something gets written and you start to see the places where it has not been written yet.

More than half coming out as "none" is not the result of slacking. Once a program holds what a program can hold, what is left is judgment alone. All the heading does is make that state visible.

There is one more thing to decide when you write "none." Whether to keep that prohibition. The criterion for keeping it is whether that failure still reproduces today. Think in terms of "does the AI still need this prohibition," and every one of them looks needed, so nothing can be deleted. The first condition the instruction file documentation lists for adding to it is also when the AI makes the same mistake a second time.

The "Machine check" heading exists only on entries that are on the list. A prohibition outside the list, like the translation rule at the top, is never counted under this heading. Who decides whether to move one onto the list is still me.

What I stopped re-reading

The list of prohibitions itself. The work of re-reading the whole thing to find out "how far does the program hold, and from where does a human look" turned into looking at that one heading. However many entries pile up, the only thing you read is that 1 line.

Caveat: an automated test breaks in the direction of "not failing"

There is a way of breaking on each side.

A "yes" breaks in the direction of spinning idle. An automated test that gathers what it scans in bulk through a path specification will, if you mistype that path, scan 0 targets and go on passing forever. Always place "the scan covered N targets or more" beside it (I wrote about this in detail in article 6 of the previous series).

A "none" breaks in the direction of becoming an escape route. An automated test that is a nuisance to write can be waved through with "none (covered by review)." The line is simple: only whether a program can decide the truth. Write "none" when it can be decided, and that is not a classification but a postponement, and half a year on you can no longer tell the two apart. Writing the test then and there worked out cheaper in the end.

How to verify: break it, and confirm that it fails

Prerequisites

  • You have added a "Machine check" heading to the list of prohibitions, and 1 or more entries say "yes"
  • You start from a state with no work in progress (git status is empty)

Time required: 10 minutes

Steps

  1. You pick 1 entry where you wrote "Machine check: yes" and open the automated test named there. If no name is written there, that is your problem before you verify anything. Write into the entry first which automated test is watching it
  2. You deliberately create one violation of that prohibition. Write it inside the range the automated test scans
  3. You run that automated test

Pass conditions (all of them have to hold)

  • The automated test fails
  • The message shows the location of the violation you wrote in step 2

If it does not pass

  • It did not fail: the "yes" on that entry is a lie. Either fix it to "Machine check: none (covered by review)" or widen the range the automated test scans, and decide which on the spot
  • It failed, but the location is not shown: whoever fixes it (the AI, or you) ends up hunting for it every time. Rework the message so that it carries the offending location

Cleanup

  • Delete the violation you wrote in step 2, and confirm that it goes back to passing and that git diff is empty

An automated test that does not fail when you break the rule is worse than no test at all.

The AI reads the 1 line "Machine check: yes" as evidence that an automated test is guarding that spot, and does not go and check the prohibition for itself. What this article put down was the 1 line that saves you re-reading the list. If that 1 line is a lie, the judgment not to re-read has already broken first.


Next time, the art of turning down the AI's requests to check. When the AI tells me to "open the screen and confirm by eye," I do not open it. Even so, nothing ever moves ahead while it is broken.


Series: The Art of Not Listening to the AI's Opinions

CC BY 4.0 はここまで