This is article 6 in the series "The Art of Not Reading." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the introduction.
This time it is the rules you write in the instruction file. For a rule to work you need all three of read, remembered and followed, but an automated test or a hook fails whether or not it was read, so it needs none of the three. This article stops writing down what you want honored, and deals with what to put down instead.
Start by checking just one thing.
Pick one prohibition written in your instruction file.
It is a rule you decided the AI is not allowed to break while you develop. Can you say how many times it was broken last month?
Not being able to say is normal, because nothing happens when the AI breaks it. For a rule to work, all three have to hold: the AI reads it, remembers it, follows it. In a line written into the instruction file, those three come apart one after another.
Meanwhile, your project already holds a great many rules that nobody reads and that are honored every single time. Lint, type checking, CI. They are honored not because they are read but because breaking them makes the build fail.
What I understood fits in one sentence.
A rule works only when it is read. An automated test or a hook fails whether or not it was read.
CC BY 4.0
Why a rule works only when it is read
This is not a matter of attitude but of how the line arrives. The official Claude Code documentation says that the contents of the instruction file and the memory arrive not as the system prompt but as a later user message. They are treated as context and not as an enforced setting; it goes as far as saying that if you want to stop a behavior, use a hook. On the hooks side it says that the action "always happens rather than relying on the LLM to choose to run it." The same 1 line takes a different route depending on where you put it.
flowchart TB
Y["The 1 line you wrote"] --> Q{"Where you put it"}
Q -- "Instruction file / memory" --> C["Loaded when the conversation starts,<br/>and arrives as a later user message<br/>= context, not enforcement"]
Q -- "Typed straight into the conversation" --> U["Arrives at the end of the conversation,<br/>as the same kind of user message<br/>= this too is not enforcement"]
Q -- "Automated test / hook" --> H["Always runs at a fixed point<br/>= applies whatever the AI decides"]
C --> R1["Works only when it is read"]
U --> R1
H --> R2["Fails whether or not it is read"]
Of the three, being remembered never holds at all. The Claude Code documentation says that sessions are independent, and each new session starts with a fresh context window, without the conversation history from previous sessions. Only what sits on disk carries over, and nothing stays on the AI's side. When a line you added last week still looks effective today, the AI did not remember it; it was simply loaded again that morning. "Remembered" collapses back into "read" every single time.
Whether it gets read is decided, as the introduction showed, by how long the context is at that moment. The instruction file is only one of the things lined up in there, and the longer it grows, the less accurately any one line can be pulled back out of it.
Being followed, the third one, carries no guarantee either. The documentation on the instruction file goes on to say that the AI reads it and tries to comply, but that there is no guarantee of strict compliance, especially with vague or conflicting instructions. Adding another "never do X" after every failure is what builds that condition. More lines mean overlapping content, and overlapping content ends up contradicting itself.
On top of that, what a line in the instruction file competes with is a habit ingrained by training. The official documentation splits what belongs in the instruction file, "Bash commands the AI cannot infer," from what does not, "standard language conventions the AI already knows." The known side surfaces without being written down. The standard command from article 1 is exactly that, and the "use the wrapper" line lost to rails test. The more a behavior is one you are writing a line to stop, the stronger it already sits inside the AI.
Before any of the three comes apart, there are readers a rule never reaches at all. A subagent that starts up separately, without carrying the main conversation over, is a different worker as article 5 had it, and nothing reaches it apart from the prompt you hand over.
An automated test passes through none of the three. Whether it fails does not depend on whether the AI read that line, remembered it, or decided to follow it.
The mechanism: translate the rules into automated tests, one at a time
You pick a "never do X" out of the instruction file or the memory, rewrite the broken state into a form a program can detect, and add it to the test suite. Here is what it looks like to translate the rule "every fetch that writes to the server carries the CSRF token header" (the kind that still works on your own machine when you forget it, and that the AI tends to drop when it writes a new call).
# test/reference/rule_csrf_token_test.rb (a shortened version of the code on my machine)
class RuleCsrfTokenTest < ActiveSupport::TestCase
test "JS that sends non-GET requests also builds the CSRF header" do
files = Dir.glob("app/javascript/**/*.js")
non_get = files.select { |f| File.read(f).match?(/method:\s*["']?(post|put|patch|delete)/i) }
missing = non_get.reject { |f| File.read(f).match?(/X-CSRF-Token/i) }
assert_empty missing, "a non-GET fetch without the CSRF header: #{missing.join(', ')}"
# ⚠️ prevents a pass when 0 targets were scanned. If the automated test itself breaks, it fails here
assert_operator files.size, :>, 150, "too few files scanned (is the glob broken?)"
end
end
In the production project on my own machine, typingtube, a web service for practicing typing along with music videos on YouTube, I moved 8 behavior rules out of the memory into automated tests in this shape. The CSRF token example above is one of them. On the memory side, a rule that has moved shrinks to a single line, "a check is watching this," and even when the AI has not read that line, breaking it makes the test fail and the work comes back.
A hook is another place to put this, but as the introduction said, putting down one hook also adds to the reading. Whatever a hook returns piles into the conversation every time it runs. An automated test comes back into the conversation only as the result of a round you ran.
Two points matter in how you write it.
One: close off the shape where the automated test itself spins idle. That is the last line of the code above, which always sets "the scan found at least N files" beside the assertion. An automated test with a typo in its glob path scans 0 files and passes forever. An automated test that has never failed gives you no way to tell a test that is guarding something from a test that is broken.
Two: for rules you cannot fix everywhere right now, make a ratchet. If the existing code holds 40 violations, record "40 or fewer" in a file as the baseline and make it an automated test that fails when the count grows. Only new violations are stopped, and clearing the backlog is never rushed. When the count drops, lower the number in the baseline file and update it, and the improvement can no longer roll back. The baseline does sit inside the repository, though, and the AI can rewrite it too. An automated test that failed also passes once the number goes up.
What cannot be moved stays where it is
The only rules you can translate into an automated test are the ones whose truth a program can decide. Judgments and policies such as "quality comes before the release" or "we deliberately do not build this feature" cannot be translated, so they stay on the side that gets read. In my environment too, 8 rules moved and the policies stayed in the memory. Once what a program can hold has moved across, all that is left in the instruction file and the memory is the policies. Being shorter, what remains dilutes less easily.
What I stopped reading
The rule documents. Whether the AI has read them, whether they ever reached a subagent you started up separately, or whether someone who just joined has read them, breaking a rule makes the same automated test fail in the same place. The work of "announcing the rules," "reminding people," and "checking by eye that they are honored" disappeared all at once. The same guardrail stands for people and for the AI, so on a team it works better than in solo development, for the simple reason that nobody has to go around telling anyone.
How to verify: break it, and confirm that it fails
Prerequisites
- You have finished translating 1 rule into an automated test
- That automated test passes right now
- You start from a state with no work in progress (
git statusis empty)
Time required: 10 minutes
Steps
- You deliberately create one violation of that rule (in the example above, you write one
fetchwithout the CSRF token header). Write it in the real code. If you write it outside the range the automated test scans, this verification confirms nothing - You run that automated test on its own (the full suite is not needed)
Pass conditions (all of them have to hold)
- The automated test fails
- The failure message shows the file name and the line number you wrote in step 1
If it does not pass
- It failed, but the location is not shown — whoever fixes it (the AI, or you) ends up hunting for that place every time. Rework the message so that it carries the offending location, then redo step 2
- It does not fail at all — what the scan covers (the glob path, say) does not include the place you wrote
Cleanup
- Delete the violation you wrote in step 1 and run the same automated test once more. Confirm with your own eyes that it goes back to passing and that
git diffis empty
An automated test that does not fail when you break the rule is worse than no test at all. The AI in the next session reads that pass as "this rule is being honored." It has no reason to check again, so it moves on with the violation still written.
And what this article did was put the rule documents into a shape you do not have to read. The only ground for not reading them is that an automated test is watching. A pass returned by a test that is not watching erases that ground first.
Next time it is "The art of not tidying shared memory." The memory that sessions running in parallel share is something I do not tidy. It works properly all the same.
Series: The Art of Not Reading
- ← Previous: 5. Read your subagents' output, and only that
- → Next: 7. The art of not tidying shared memory
- All articles: Introduction: I Barely Read What the AI Outputs Anymore
CC BY 4.0 はここまで