This is article 9 in the series "The Art of Not Running." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the introduction.
This time it is about what has been learned and written down for the person who comes next. I handed 12 workers the same list of 40 and the handoff notes the previous person in charge had left, and the only difference was how many of the 3 numbers written in the notes were at odds with the current code. The stale number got copied straight into the artifact in 1 of the 15 rows that were off. That is where the trouble is. You write down what has been learned because you think the next person to read it will use it and save themselves the effort. In fact, though, it got counted again every time it was handed over. Even when the notes were right throughout. And the person who left them cannot see that the counting again is happening. What is left for the person who received them is a choice between checking and betting, and most choose to check.
In article 9 of the first series, "The Art of Not Reading," I lined up the count a command had produced next to the report. The AI had not written that line, so it never went through the summary's judgment. What I measured this time is that count. It does not drop out of a summary: in exchange for staying, it turns into a lie the moment the code moves. What has been learned and left behind hands the reader a choice, every time, between checking and betting.
Start by checking just one thing.
Open your project's instruction file or its handoff notes, pick 1 of the numbers written there, and count it again right now.
Was it right?
And remember how long that counting again took you.
The same thing comes up later.
CC BY 4.0
To the same list I attached notes of differing staleness
The material is the same list of 40 findings used in articles 3, 5, 6 and 7 (only the previous article used a different material). For the 12 workers I attached to it 1 sheet of handoff notes left by the previous person in charge. The notes have 3 numbers written in them.
The list holds 40 findings
Of those, 8 sit under app/jobs/
37 findings are attached to files with the .rb extension
The work I asked for was 1 thing only. Write those 3 numbers into a file, as the summary to pass to the person coming in next. I wrote neither "count them" nor "copy them from the notes." That is because the moment either one is written, what is being measured shifts from "does it use what has been learned" to "does it do as it is told."
The only difference is how many of the numbers in the notes are off from the current code.
| The state of the notes | Workers |
|---|---|
| All 3 right | 3 |
| 1 of them stale | 6 |
| All 3 stale | 3 |
The stale numbers are not nonsense. They are the numbers from before 6 findings were added to the list (which works out as app/jobs/ going up by 3 and .rb by 6). If the numbers contradict each other, the reader reacts to the way the notes are broken rather than to what they say.
The stale numbers were almost never written
The numbers that were off come to 15 rows in all. Of those, the number in the notes was copied straight into the artifact in 1 row.
| The state of the notes | Rows that were off | ⚠️ Copied the stale number |
|---|---|---|
| 1 of them stale | 6 | 1 |
| All 3 stale | 9 | 0 |
Whether the discrepancy sits in 1 place or in 3 separates nothing. 6 workers against 3 is a headcount where a difference only shows up when they split completely (the measurement came out at p = 1.00, and the width of the difference takes in everything from minus 41 points to plus 51 points). I counted this before running it and built it knowing it could separate nothing. What I wanted to set up were the next 2 things instead.
The 1 worker that went along with what has been learned was the 1 that did not notice
Of the 9 workers whose notes were off, 8 reported on their own that "the notes and the current code do not match." I had not asked.
The counts recorded under "What is known" in HANDOFF.md (total 34 / app_jobs 5 / rb 31) did not match the contents of
items.md(it is in fact the 40 entries C01 to C40, with 8 underapp/jobs/and 37 carrying.rb), so I recorded in SUMMARY.tsv the values from countingitems.mdagain for real
And those 8 workers, all 8 of them, kept the stale numbers out of the artifact.
The 1 worker that did write the stale numbers was the 1 that never once touched on the discrepancy. Its report reads like this.
For the contents I reflected, as they stand, the 3 figures recorded under "What is known" in HANDOFF.md (total 34 / app_jobs 8 / rb 37)
This 1 worker alone did not count the list. It used tools 3 times and ran for 17 seconds, the fewest and the shortest record among the 12 workers.
So what separated going along with what has been learned from not going along with it was not whether it was believed. It was whether it was checked. Every worker that checked threw what has been learned away, and only the side that did not check went along with it, and got it wrong.
Even when the notes were right, they were counted again
11 of the 12 workers counted the numbers written in the notes again themselves. The 3 workers whose notes were right throughout did too, all 3 of them.
Counted C01 to C40 in items.md and confirmed it matches what HANDOFF.md records
The notes did not save anyone effort even once. The only effort saved was the 1 time by the 1 worker that alone did not count again, and that was the only wrong answer.
The cost side came out like this.
| The state of the notes | Run time (per worker) | Tools used |
|---|---|---|
| All 3 right | 22.0 seconds | 4.0 times |
| Stale somewhere | 36.8 seconds | 5.3 times |
That is 1.67 times. The difference in run time came out at p = 0.023 on a permutation test (the number of tool calls came out at p = 0.21, which does not stand).
How you read this needs care. The stale notes are slower because the side that found the discrepancy wrote out the grounds as well, for the report. This is "the cost that what has been learned, once stale, makes the reader pay," not a story about the AI getting slower.
The principle: what has been learned, left behind, hands the reader "check it or bet on it" every time
In article 6 of the first series, "The Art of Not Reading," I wrote that a rule works only when it is read, and that an automated test or a hook fails even when it is not read. The same line can be drawn through what has been learned.
What has been learned and written down does get read. After it is read, the reader has only 2 roads. Hold it against the current code and check, or hold it against nothing and bet. Whichever road is taken, the side that wrote it gains nothing — if it was checked, then what has been learned turned out not to have been needed, and if it was bet on, the stale number goes straight into the artifact (as with the 1 worker this time).
In article 8 of the second series, "The Art of Not Listening to the AI's Opinions," I wrote that only the lower tier needs tidying and the upper tier can be left alone. That was about volume. This time it is about correctness, and correctness cannot be left alone. That is because even if the volume does not grow by 1 line, the moment the code moves it becomes a lie. The line against article 2 of the first series (the art of not reading memory) runs the same way: that one is about how much you read, this one is about being at odds with the current code.
The line against article 8 can be drawn like this. That one is about what was not left in the artifact, and this time it is about what did get left in the artifact.
How to measure: shift only the numbers in the notes away from the current code, and change not 1 character of the work you ask for
To keep the condition down to 1, the list handed to the 12 workers is identical down to the byte. For the notes too, an automated test checks every time that, with the figures masked, they agree exactly across the 3 conditions. If the shape changes, that alone changes what is being measured.
Not 1 word of "count it again," "check it" or "copy it as it stands" goes into the request. An automated test checks every time that none of those words has slipped in.
Whether it noticed is not counted from a list of words. A script looks for both the current code's number and the notes' number appearing in the report. Lean on a list of words and you drop wordings you do not know — I did that twice, in article 6 and article 7 (in article 6 it was Japanese wordings such as "I do not know" and "unclear," and in article 7 it was reports that came back in English, and neither was picked up even once).
What I stopped leaving behind
I stopped writing down in prose the things I want the AI to remember. Instead I bury them where the hook brings them out at the moment they are needed.
- Write it into the failure text of an automated test. Not just "please raise this number," but why it is to be raised, put into the message you see when it fails
- Write it into the message you get when something is stopped. The moment the mechanism that stops things stops one, it brings out what to do instead alongside it
- Do not write the number down; put down the script that counts it instead. From the moment a figure sits in prose, it starts going stale
This is not "write nothing." It is the same as having you produce the numbers for your own environment in article 8: keep an eye that measures, and move it into a shape you do not have to measure again every time.
Even so, what was decided earlier does not get forgotten. It is not forgotten because someone is remembering it for me, but because what was decided comes out of the place where it was decided, every time.
Caveat: what this experiment cannot say
- There is 1 thing acting from outside the conditions. It is the heaviest caveat this time. The workers run with this project's instruction file handed to them every time. In it sits 1 line amounting to "do not swallow an old record whole." In fact 1 of the 12 workers came back with a report using almost that wording as it stands. So what could be measured is not "does it believe what has been learned and left behind" but "which does it follow, what has been learned and handed over, or the 1 line that arrives at all times." There is no control with that 1 line taken out
- There are not enough workers. The density of the discrepancy (1 place versus 3 places) is a design that separates nothing from the start. I counted before running it, wrote that down, and then ran it
- What I measured is only numbers. Whether what has been learned and written in prose (the reasons behind a judgment, or how it came about) is treated the same way is not measured
- 40 entries is a size you can count again. With material where counting again costs more, the betting side might grow
- The staleness of the notes holds together. If the numbers contradicted each other, the result ought to change
How to run the experiment: shift only the numbers in the notes you left behind away from the current code
Prerequisites
- Decide on 3 numbers a script can count from the current code. Make them ones where the way of counting does not split (this time it was "the extension is
.rb," which, even with.erband.rakemixed in, is settled uniquely by matching the end of the name) - Make 1 sheet of record to hand over. Apart from those 3 numbers, do not change 1 character between conditions
Time required: 7 minutes for 12 workers (the total run time of the workers. Per worker, 22 seconds when the record is right and 37 seconds when it is stale)
Steps
- Hand the same material to n workers. Shift only the numbers written in the record, in 0 places / 1 place / 3 places
- Keep the direction of the shift the same. Make it the shape of "it grew afterward" and the record holds together. Do not use nonsense numbers — they get reacted to for the way they are broken rather than for what they say
- Ask for 1 artifact only. Write neither "count it" nor "copy it"
- Sort the numbers that were written into 3. The current code's number / the record's number / neither of the two, and a program decides which
- Look at the report separately. Count whether it touched on the discrepancy by whether both the current code's number and the record's number appear. Keep the list of words as no more than an aid
Pass conditions (all of them have to hold)
- You have checked by hash that the material handed over is identical down to the byte across every condition
- The record, with the figures masked, is identical across every condition (the shape has not changed)
- Not 1 word urging checking or copying is in the request
- In the condition where the record is right throughout, everyone gets it right (a check that the task itself is not too hard)
If it does not pass
- Wrong answers came out even where the record is right: the task is too hard. In that state, an error caused by the record and a plain miscount turn into the same figure
- Everyone wrote what the record said: the request is urging them to copy. Go back over it 1 word at a time
- Everyone counted again: that is itself a result (this time came close to that shape). Check first, though, whether the instruction file holds a line amounting to "doubt an old record." If it does, what could be measured is the working of that 1 line instead
- A difference showed up with the density of the discrepancy: count the headcount first. A difference that shows up at 6 workers against 3 is usually chance
Cleanup
- Delete the directories handed to the workers, and the answer key. Delete the answer key along with them: leave it and, the next time, while you believe you have rebuilt the material, you will be scoring against the old answers. That is exactly the accident this article measured
What has been learned and written down did not save the next reader any effort. The only effort it saved was that of the 1 worker that did not check, and that 1 worker was the only one that got it wrong.
Next time is the finale, the art of not experimenting. Over 9 articles, I have measured 9 assumptions. Those 9 fold into 4.
End of CC BY 4.0