I no longer run the product in order to check something. What I have measured instead is only the things I put into a shape that leaves the process in the artifact. The series "The Art of Not Running" is the record of measuring again, against problems whose answer is known, everything the previous 4 series wrote down as "do it this way and it works." This article is the finale. How each article measured can be reached from the list in the introduction.
There is one thing I have to confess.
I do experiment, in all 9 articles
The series is titled "The Art of Not Running," and yet all 9 articles end with a procedure for an experiment. I make material whose answer is known, start several AI workers that do not carry the conversation over, change exactly 1 condition, and match what comes back against the answer key. In place of running the product, every time, I was running something else.
What is more, the number of experiments does not stop at 9. Article 1 rebuilt the conditions and measured again 3 times, and article 6 twice.
And the answers that came out were mostly not the thing I set out to measure.
- In article 1 I was about to conclude, from the results up through the 2nd experiment, that "the more moves an instruction takes, the less it gets followed." It came apart in the 3rd experiment
- Article 6 had a bad question the 1st time (it was a question you could check by reading, so there was no room to guess), and I rebuilt the question itself
- Article 7 went to measure "the longer the conversation, the more readily the AI folds," and what was dropped came to 2/4 on the short side and the long side alike
So the predictions I wrote before setting up an experiment missed more often than they hit. Line the 9 articles' answers up again all the same, and the ones saying the same thing from a different angle overlap. Folded, they come to 4. This article is about those 4.
What the 9 articles turned up folds into 4
These are the "what I understood fits in one sentence" lines put at the end of each article, gathered back together by likeness.
| What it turned up | Where it came from | |
|---|---|---|
| A | An act that leaves nothing in the artifact is not complied with, and it piles up | Article 1 / article 4 / article 8 |
| B | The wording of the request decides which answers the other side can pick | Article 5 / article 6 |
| C | Hand over the same thing and what comes back scatters. What lines up is the kinds | Article 3 / article 7 |
| D | The errand in front of the reader is stronger than what you wrote and handed over | Article 2 / article 9 |
A comes out of the 2 instructions placed side by side in article 1. Their content and the effort it takes to follow them are the same, and I made the only difference whether a trace of compliance is left in the artifact. The instruction on the side that leaves no trace, placed on its own, is followed by only 3 of 8 workers. Add 1 line next to it that does leave a trace, though, and all 4 followed it (I changed neither the wording nor the position, not by 1 character). The same shape turned up in the later articles too. An instruction with a condition attached, "if needed, please do such and such," actually fired in 1 of 6 workers (article 4). In 1 session of my own, 25.8% of the tokens I used went on "just going to look at whether the work had finished" (article 8). Neither of them shows up in the finished files, not by 1 character. What does not stay behind cannot be counted, so it keeps growing.
B is where 1 sentence of how it was asked turned the result over. In article 5, replying to the AI's 3 picks with nothing but "please pick the 3 again" had 7 of 12 workers drop what they had put 1st. Add 1 sentence to the request, "if you judge that there is no need to change, write the same 3 as they are," and it becomes 0 of 6. Article 6 was the same. Of the workers handed a question whose answer does not exist with nothing but a prohibition, all 4 came back having made no artifact, and the 4 that had "if you do not know, write that you do not know" added all wrote one and returned it. It is not the other side's ability but our own wording that decides which answers the other side can pick.
C is article 3, where the same task, not 1 byte different, went to 12 workers. Each of them picked 3 items, so the reason sentences came to 36, and the 36 were 36 distinct ones with not a single sentence shared. The items picked are spread across 9 of the 40. Count those 9 again by the kind of defect, though (the 6 kinds the script decided when it made the material), and they fit into 2 kinds. In article 7 as well, against exactly the same push-back, 2 of the 4 workers dropped it, 1 kept it, and the remaining 1 only moved it down the order. The answer you get back from asking 1 worker is 1 point inside that spread.
D is about what happens after what you wrote gets read. In article 2, writing 1 line in the instruction file, "when you look at the contents of a log, use this command," had all 4 workers call that command. 3 of them, though, go back to the bare command partway through. They went back the moment an errand turned up that the tool I had put down did not cover, and they never came back. Article 9 has the same shape. 11 of the 12 workers counted the numbers written in the handoff notes the previous person in charge had left again themselves. The 3 workers whose notes were right on all 3 counted them again too, all 3 of them. What you write does get read. Having been read, it loses to the errand at hand.
Combined, they answer things I have not measured
This is the main point. A through D are the answers from 9 articles, and at the same time they become a scale you can hold against what has not been measured yet.
Let me try it. The 3 below have not been experimented on once in this series. They are answers got by holding the 4 against them.
Write "ask if anything is unclear" in the instruction file, and will the questions come? From B, it falls on the side of coming, because writing it makes 1 place to go for when the answer is stuck.
From A, though, a run where it did not ask leaves not 1 character in the artifact.
So whether that 1 line is working cannot be checked. If you want to check it, the thing to do is not to strengthen the wording but to put the question itself into a shape where it becomes an artifact. "Write anything unclear into QUESTIONS.md, 1 line each." It is the same remedy as in article 1, where only the instruction on the side that leaves a trace was followed.
Have the AI score its own work and redo it only when the score is low: does that work? From A and B, it does not. That is because whether "when the score is low" applies is judged by the other side, and on top of that a run where it did not say the score was low leaves nothing in the artifact. It has exactly the same shape as the "if needed, please do such and such" measured in article 4 — the moment you attach a condition, the judgment of whether that condition was met goes to the other side as well.
Hand over a summary of the previous session and have it carry on: is that safe? From D, for the numbers it is safe. That is because the numbers you hand over get counted again by the reading side. The danger is on the other side: what cannot be counted again (why it was decided that way, which proposals were rejected, what was tried and did not work) has no way of being checked, so it gets carried over as a premise just as it stands. What article 9 measured was only the numbers side. The answer for the side that was not measured comes out of the side that was.
None of the 3 has been experimented on. I only held A through D against them.
What I stopped experimenting on
I stopped checking every idea I have with an experiment. What I do instead runs in this order.
- Hold it against A through D first. If any of the 4 explains it, I stop there
- Measure only what they do not explain. I measured 9 articles' worth because at the time I did not have the 4 yet
- Write "a prediction" on any answer got by holding them against it. It does not go on the same shelf as a measurement
This is not a story about "you do not have to measure any more." It is about moving, once you have an eye for measuring, to a shape where you do not have to measure again every time. In article 9 I threw away what had been written down as learned and moved it to a shape where it comes out from the hook side when it is needed. For these 4, the place they are kept has just become "inside the reader's head."
And these 4 are themselves something written down as learned. The conclusion of article 9 applies to them as it stands. From the moment the code moves, these 4 start going stale. That is why I put how to measure into every article. You do not have to believe them.
Please run them.
Caveat: an answer got by combining is not a measurement
This is the heaviest caveat in this article. The 3 above are inferences drawn from 9 articles of measurement, not results that were measured. They will miss.
In fact, the predictions I wrote before an experiment in this series missed more often than they hit. That I knew they had missed was, every time, because I measured.
The denominator behind A through D is only 4 to 12 workers per experiment. What can be said strongly is only where the result split completely between the conditions. Without that 1 line written, none of the 4 called it; with it written, all 4 called it (article 2). Asked to write a number, 5 of 6 counted; not asked, it was 0 of 6 (article 8). A difference such as 1 of 4 against 3 of 4, on the other hand, cannot be told apart from chance. In the "Caveat" of each article I have drawn that line, 1 at a time.
And what was measured is 1 model at 1 point in time. Change the generation and the numbers move. A through D are differences that came out with only the condition changed, on the same day and within the same model, so they ought to last longer than the numbers, but that is no guarantee either.
There is 1 more thing I could not separate across the 9 articles. The workers in the experiments run with this project's instruction file handed to them every time. It holds lines amounting to "do not implement on a guess" and "do not take an old record at face value," and part of the results of articles 6 and 9 can be explained by those lines alone. There is no experiment that compared with those lines taken out. I am writing it down rather than hiding it.
The conclusion
The art of not experimenting is the art of using the results of experiments in combination.
What the 9 articles measured was 9 assumptions. If what you take home stays at 9, the 10th question needs a 10th experiment. That never ends. Because the 9 fold into 4 and the 4 combine, the 10th no longer needs an experiment.
I made one promise in the introduction. My numbers can be checked just by running them. Deciding whether to believe them can wait until after that. Having read this far, you can go 1 step further.
When you read a technical claim, you should be able to read it like this. How large is the denominator. What has not been separated from what. Does this number move with the generation of the model, or is it a matter of a property of the sentence. This way of reading does not work on this series alone.
The first series was titled "The Art of Not Reading." It started from building a shape you do not have to read, and over 5 series it has come as far as giving the power to read back. You will be able to read other technical books too, because you can read the claims written in them from the side of the conditions they stand on.
And once you are there, you no longer have to learn from scratch every time a new tool comes out. Look at which of A through D the new thing falls under, and most of the time a combination explains it.
Please measure only what they do not explain.
In the introduction I had you recall 1 thing you believe, something you hold as "do it this way and it works." When did you last check it, and how many runs' worth of results did you count?
Answering the same question now, mine comes out like this. I counted 9 of them, each worth 4 to 12 workers. Everything else I believe, I have not counted. That is why this article says "a prediction" where it does. Do not put what you counted and what you predicted on the same shelf. That is all this series has to hand over.
For how each article measured, start from the list in the introduction. The production that was the stage for the previous 4 series is still running at typingtube. This series is the record of stepping away from it once and measuring.