I no longer run the product in order to check something. Not running the tests and fixing what falls out, not opening a screen to look, not learning from what happened in production. What I build instead are problems whose answer is known in advance. Whether something was read cannot be measured: what can be measured is only whether a thing that cannot be made without reading it exists. This series measures again, one experiment at a time, everything the previous 4 series wrote down as "do it this way and it works."
First, a word on how to read it. This series can be followed by reading alone. Each article starts from what happened in front of me and goes on through what I put down and what went away. Near the end of each article there is a section called "How to run the experiment." You do not have to read it first. Come back to it when you want to check whether my numbers come out in your own environment.
Learning the remedy from the symptom
Each article starts with a symptom, something that is not working. The remedy comes with the relevant behavior and specifications of the AI, and an explanation of why that remedy fits.
You do not need to learn why it fits before you use it. Put one mechanism down and the symptom is handled.
Even when you know the behavior and the specifications, the AI can still get it wrong depending on how they combine. Ordinary design and development is unlikely to catch these pitfalls, where several behaviors and specifications are tangled together. You learn them from the symptom, one at a time.
Start by checking what you recognize
Start by recalling just one thing.
One thing you believe about using AI, something you hold as "do it this way and it works." How you write your instruction file, or how you ask: either is fine.
When did you last check it, and how many runs' worth of results did you count at the time?
If you are not developing alongside an AI yet, it is natural to have nothing you believe. This series is a record of measuring again what I wrote as "it works" across 4 series, so you can read it as it is.
The grounds for 4 series and 44 articles are mostly my own impression
This is the 5th series.
- Series 1, "The Art of Not Reading": it put the AI's output into a shape you do not have to read
- Series 2, "The Art of Not Listening to the AI's Opinions": it put the AI's proposals and judgments into a shape you do not have to weigh up twice
- Series 3, "The Art of Not Telling": it put the human's input into a shape you do not have to say
- Series 4, "The Art of Not Checking": it put checking itself into a shape you do not have to decide on the spot
- Series 5 (this one): it measures whether those 4 series were really working
All 44 were written from what happened in 1 production product. It is typingtube, a web service for practicing typing along with music videos on YouTube. I run it on my own, and it is still live.
That way of writing had a pitfall common to all 44. I put a mechanism down, and things went well afterward. That is the grounds. I did not compare the days it was down with the days it was not. There was no way to compare. There is only 1 production, and time does not rewind. In every one of the 44, what could be counted was a one-off event.
So I do not know which of the things I have written are true. It may have been working, or it may have been the same without it. Which of the two it is, I can never check as long as I am running production — and that is where this series starts.
What I stopped running is the product
The title is "The Art of Not Running," but this is not a paper exercise where nothing gets run. This series runs more than the previous 4 did (every article runs several AI workers). What I stopped running is the product, as a way of finding something out.
Let me line up what the previous 4 series stopped running. I stopped opening screens and clicking around them. I stopped taking screenshots and running through the whole thing to confirm. I stopped re-running tests. I stopped running a 2nd test run alongside the first. I stopped running all of the heavy E2E. All 5 are runs meant to check that the work in front of me was being done properly. Means, count, scope, number: I cut them one at a time.
What was left is running in order to find things out. Run the product, and learn from what happened. It is how the grounds themselves get made, so this is the one thing I never let go of across the 4 series.
The finale of series 3 was titled the art of doing nothing. That one was about the human not stepping into the work, and what is being stopped is different here. Here I do step in. If anything, I reach in myself every time there is an experiment. The only thing I do not run is the product as material for finding something out.
What I use instead are problems whose answer is known in advance.
For example, I put down 12 files that imitate logs and let a script decide which values ought to be picked out of them. The answer key goes outside the directory that is handed to the workers (if a human makes it afterward, it gets pulled toward the answers that came back). The same task goes to several AI workers that do not carry the conversation over, with exactly 1 condition changed. What comes back is matched against the answer key. There is no gap there for my impression to get in.
What I hand over is 1 script that tells you by running
In the previous 4 series, each article handed over "1 file that works once you put it down." A hook, a wrapper, a line of instruction. That was what the reader took home.
In this series there is nothing to put down. What is handed over instead is how to measure. At the end of each article I put down how I made the material, how the condition was changed, and how it was counted.
You can check for yourself whether the same numbers come out in your environment.
My numbers can be checked just by running them. Deciding whether to believe them can wait until after that.
It is the reverse of what the introduction to series 1 wrote: "A mechanism works simply by being placed. Understanding is soon enough once you have started using it." That one was a promise that you could put it down without reading what was inside. This one is a promise that you can run it without believing what is inside.
Each article is made of 1 assumption, 1 way of measuring, and 1 experiment procedure. The assumption is usually something the official documentation says, or a practice that is widely recommended. I set it up as a hypothesis, throw an experiment at it, and write down the numbers that came out. There is no article that ends by introducing a hypothesis.
The articles that came out wrong are in here as they are
There is one thing to say up front. Some articles came out with the result that it was not working after all. I did not drop them. Dropping them would make it staging, not an experiment.
There was more than 1 way of coming out wrong. An article that came out the opposite of what I expected, an article where no difference showed up across the conditions at all, and an article where what I was trying to measure was wrong along with the design of the experiment, so the 1st run was done over. Each one says so.
And what cannot be said is lined up at the end of each article too. How large the denominator is, what has not been separated from what, which numbers move with the generation of the model. To put down 1 caveat that covers all of it first: the numbers in this series are measurements of 1 model at 1 point in time. They are not written for you to take the numbers themselves home. What you take home is the one thing each article puts at its end, the understanding that fits in one sentence.
One more thing. You can start here without having read the previous 4 series. Each article puts 1 line at the top about what the previous series wrote, so reading that alone is enough.
That said, this series is about someone who has finished putting the mechanisms down. If you have not put them down yet, reading from series 1 is faster.
How the series is laid out
The first 4 are about the lines written in the instruction file not working the way the person who wrote them expected. They measure how far it works, and where it stops working.
- 1. The art of not enforcing a procedure
- 2. The art of not enforcing a wrapper
- 3. The art of not aggregating
- 4. The art of not being useful
The middle 3 are about the answer being decided not by what the instruction says but by how it is asked on the spot and by how long the conversation runs. The same thing is being asked, yet the result changes with one sentence of difference and with the number of exchanges.
- 5. The art of not making the AI pick again
- 6. The art of not closing the escape hatch
- 7. The art of not keeping a session going
The last 2 are about what you cannot see when you look only at what is left once it is over. They deal with the finished files, and with the handoff passed to the next person.
The finale combines the results of the 9 articles. Combined, they answer things that have not been measured yet.
You do not have to read them in order. Article 9 does follow on from article 8, though, so read those 2 in order.
Next time, the art of not enforcing a procedure. I stopped rewriting the instructions the AI did not follow. And even so, those instructions came to be followed.