I barely read what the AI outputs anymore. Not the test results it produces, not the notes it appends to its own memory, not the work logs of the subagents (the separate AI instances it starts up without carrying the main conversation over). And it is not only me: the AI itself is set up so that it does not have to read its own output either. Reading spends the attention budget, and what you add dilutes what is already there, so adding one thing makes everything else a little weaker.
Did quality drop because of it? It did not. This series is about building that state one mechanism at a time.
Near the end of each article there is a section called "How to verify." You do not need to read it first. Come back to it when you want to see for yourself why the mechanism you put down is needed.
Learning the remedy from the symptom
Each article starts with a symptom, something that is not working. The remedy comes with the relevant behavior and specifications of the AI, and an explanation of why that remedy fits.
You do not need to learn why it fits before you use it. Put one mechanism down and the symptom is handled.
Even when you know the behavior and the specifications, the AI can still get it wrong depending on how they combine. Ordinary design and development is unlikely to catch these pitfalls, where several behaviors and specifications are tangled together. You learn them from the symptom, one at a time.
Start by checking what you recognize
Start by recalling just one thing.
You hand your AI coding agent an instruction file named CLAUDE.md or AGENTS.md, the place your rules live. How many lines is it now, and can you remember which of those lines you added last week? Was that line honored in today's session, one continuous stretch of work with the AI?
If you are not building alongside an AI yet, of course you cannot remember. From here on we follow what becomes of that instruction file over three weeks.
If you have a terminal at hand, counting the lines makes the numbers in the next section yours.
wc -l CLAUDE.md # substitute the file name you actually use, such as AGENTS.md
It stops working in three weeks
Here is the usual course of events. Week 1: the instruction file is 10 lines, and the AI works well. Week 2: you add one "never do X" line after every failure and it reaches 30 lines. It still works, but days start to appear where a rule you added earlier is broken. Week 3: at 60 lines, what the AI could do in week 1, it can no longer do.
So you reorganize the rules, give them priorities, and move the important ones to the top in bold. 80 lines. They are honored even less.
The effort to improve is itself what pushes the state further down. That is the pit.
The cause is not the AI's ability but how attention gets distributed. The AI takes in the whole of the working memory it loads on every response, the context, and decides how much attention goes where. The instruction file, the accumulated memory, the conversation for the task at hand, tool output: all of it sits in one and the same context. The longer that context grows, the less accurately the AI pulls out the single line it needs from inside it. One irrelevant line is enough to lower that accuracy.
Rules dilute as you add them. Adding one is making every other one a little weaker. So "add a rule to make it obey" is a structurally losing move.
Anthropic calls this a finite resource that diminishes. An LLM has an attention budget, and the more tokens it holds, the less accurately it recalls from them.
The Claude Code docs are blunter: aim for under 200 lines per instruction file, because longer files consume more context and reduce adherence.
That reordering does not help has been measured too. There are tasks that ask a model to find the line it needs from a long input. In them, documents put in a shuffled order scored better than documents laid out in a coherent order. Putting things in order does not act on the accuracy of pulling them back out. All that grew was the line count.
That said, no measurement behind the figure of 200 lines has been published. The docs themselves call it a target. What the measurements quoted here measured was whether a model can find the line it needs from a long input, not the rate at which it follows instructions. How many lines it takes before your own instruction file stops working is something you have to count yourself.
flowchart TB
subgraph K["Loaded every session, on every response"]
direction TB
k1["The instruction file"]
k2["The memory index"]
k3["The conversation for the task at hand"]
k4["Tool output (test results, diffs, ...)"]
end
K --> B["The attention budget does not grow"]
B --> R["Adding one line is making<br/>every other one a little weaker"]
The production project on my own machine went through this pit as well. It is typingtube, a web service for practicing typing along with music videos on YouTube, which I run on my own.
The instruction file reached 101 lines and the accumulated memory 89 files, about 1.1MB. Merely injecting the list of rules through a hook, a mechanism that cuts in when the AI runs a command, put about 26KB into the context every single time.
A hook you install is loaded every time as well, so adding one more mechanism is also adding to what gets read. This series is the record of walking back from there.
The art of not reading
What lay on the other side of walking back was the opposite of "make it read more."
Test results arrive as failures only. Memory does not grow on its own. Subagents stop where they should stop without being watched. One by one, the things you were being made to read disappear, and that much of the load goes with them.
Why a mechanism rather than a rule? A rule only works when it is read. Whether it gets read depends on how long the context is at that moment. A mechanism works even when it is not read. Unless the condition is met, the AI cannot get past it. However long the context grows, then, the effect does not change.
The Claude Code docs draw the same line. Of instruction files they write that there is no guarantee of strict compliance. And they go on: if you want something to run at a particular point without fail, put it in a hook.
A mechanism works simply by being placed, and understanding is soon enough once you have started using it. There is no need to read it through before you begin.
Many of the articles use a hook mechanism that interrupts the agent (hooks, if you are on Claude Code), but if the tool you use does not have one, build it. What you need is a structure that cannot proceed unless something has been read; a hook is only one way of implementing that.
You do not need to put all of this in place today. Putting a mechanism in place ahead of the failure it prevents tends not to work out. Depending on your scale and structure, there are failures that never occur at all (small projects are strangers to memory bloat).
Most of the articles open with a single operation for checking the symptom in your own environment. If the symptom is there, that is the moment to place that article's mechanism.
A mechanism you place after failing settles in well, because you have just witnessed why it is needed.
How the series is laid out
The basics are for people who use a mechanism that has been handed to them as it is. Place it, confirm it, and it works.
- 1. The art of not reading test output
- 2. The art of not reading memory
- 3. The art of not reading unit tests
- 4. The art of not using skills
The advanced articles are for people who build mechanisms and hand them over. They put things into a shape that does not fall apart as the scale grows.
- 5. Read your subagents' output, and only that
- 6. The art of not reading rules (for teams)
- 7. The art of not tidying shared memory (for teams)
- 8. The art of not reading fix history
- 9. The art of not reading handoffs
- 10. The art of not asking AI to summarize
- 11. The art of not cutting in
- 12. The art of not reading work results
The final article puts together how the parts up to that point connect into something that runs without my reading it.
When, how many, and by what means to run the tests is not covered in this series.
Next time, the nearest thing to hand. The 26,147 lines that pour in every time the AI runs the tests are lines I do not read. I have yet to miss anything because of it.
Series: The Art of Not Reading