This is article 2 in the series "The Art of Not Telling." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the introduction.
This time it is procedures. It is about giving up writing the how for the AI, as in "fix it this way," and about what you hand over in its place. A procedure closes off the range of the search in advance, because the AI takes what you wrote literally.
Start by looking for just one thing.
A place in your most recent session where you wrote the how out for the AI. "First do this, then do that." "Use this function." "In this pattern."
Was that procedure the order you thought was best?
Or was it that you would have felt uneasy if you had not written it?
CC BY 4.0
I counted, and not one procedure was left
For a long time I thought that the more specifically I wrote, the better the result that came back. And in fact, the great principle people state for instructions aimed at an AI is to be clear and specific.
So it surprised me when I counted. The instruction file on the production project on my own machine (typingtube, a web service for practicing typing along with music videos on YouTube) is 102 lines. Numbered procedures came to 0 lines.
Even counting what lies outside the instruction file, the skills (registering custom commands and tools) and the catalog of missions for subagents (an AI that starts up separately and works), only 2 places still hold a numbered list.
| Where | What is in it | Is this a procedure? |
|---|---|---|
| The skill for checking the screen | 4 stages (grep → lint → screenshot → smoke test) | The order matters. You try the cheap means first, and go no further once an earlier one has solved it |
| The checklist before saving to memory | 5 items (value, permanence, duplication, consistency, approval) | Conditions to judge by. There is no need to go down them in order |
In those 3, the instruction file, the skills and the catalog, there is nothing that writes down how to go about a task. What is left is only the 2 above: the case where the order itself is the answer, and conditions to judge by that have nothing to do with order. While I was at it I counted the plan documents as well. There are 200 of them, and 0 hold a numbered procedure.
Hand the judgment over in advance and the output shrinks to match
Anthropic's prompting guidance has the passage that puts this most plainly. It is about asking for a review.
Write "only report critical issues" or "be conservative," and the model follows it literally, and reports less. Have it report everything, and do the narrowing down in a separate step.
Writing the how does the same thing. "Use this function" is "you do not have to look at the other options," and "first A, then B" is "even if there is something that ought to be looked at before A, skip it." A procedure arrives as an instruction to close off the range of the search in advance.
The AI keeps to instructions. So when your procedure was not the best one, that gap comes through as a gap in the output.
What I understood fits in one sentence.
A procedure closes off the range of the search in advance. So what you hand over is only the result, the constraints and the way to check.
The mechanism: bind what gets handed over, not the how
What you put down is a single sheet. For each mission you line up the files that must really be there before launch. The main AI that starts the subagents, hands out the work and takes in the results is what I will call "the center" here. It is the side that directs. The general term for it is the orchestrator.
# subagent_missions.yml — the catalog of missions (a shortened version of the code on my machine)
missions:
i18n_translation:
description: Backfill locale translations (1 agent = 1 language)
requires: # ⚠️ if one is not really there, the launch stops
- scripts/i18n_ledger.py # the list of work (the only input handed over)
- scripts/i18n_apply.py # the center does the inserting
- scripts/i18n_verify.py # the acceptance test is the center's too
- docs/reference/i18n_translation_notes.md # the per-language notes
model: [sonnet, haiku]
These 4 lines carry straight on from the one sentence in article 1. What machinery (something that works without being read, like a hook or an automated test) can hold is the running, and what it cannot hold is the reason. What is bound here is not the running either. It is the input (the list of work), the exit (the acceptance test), and the boundary not to cross (the center does the inserting). What is not written is how to go about it.
The hook that cuts in at launch looks only at whether these 4 are really there. It does not measure whether what is in them is any good.
It is the same single sheet I put down in article 4 of the previous series, "The Art of Not Listening to the AI's Opinions." That one was about tool descriptions: I stopped lining up the situations to use them in, and wrote only the contract for what gets handed over, thickly. This one is the side of the instructions you type every time.
There are places where a procedure belongs
This is not "do not be specific." The 2 places above taught me how to draw the line. How specific an instruction is should match how easily that operation breaks. Open ground where judgment is needed gets a rough guide; a narrow bridge where only one procedure is safe gets the exact command.
The 4-rung ladder in the table above is exactly that narrow bridge. The order itself, trying the cheap means first, is the answer, so the order gets written. "How to implement it," on the other hand, is open field: there are many roads, and more of them are roads I do not know.
So the rewrite is not "delete" but "replace."
| What I used to write | What replaces it |
|---|---|
| First do A, then do B | The result: what has to be there for the job to be done |
| Use this function / this pattern | The constraints: the places not to touch, the boundary not to cross |
| Check it at the end (article 1) | The way to check: which automated test has to pass |
What I stopped writing
The how. The route through the implementation, the function to use, the order to start in. The amount I write has not gone down: the result, the constraints and the way to check are written more carefully than before. What went away is only the lines that closed off, in advance, roads I do not know.
Caveat: if what gets handed over is not ready, writing a procedure does not get through either
The 4 in the catalog take time to write. You build the list of work, you write the acceptance test, you gather the notes to watch out for. Saying "just handle it" without doing that is not what this article is about; it is just dumping the work. As the previous series wrote in the same place, what you did not hand over, the subagent makes for itself.
How to verify: throw the same task, handed over in two different ways
Prerequisites
- You can set up 2 places that do not carry the conversation over (start 2 new sessions, or start 2 subagents)
- You have a repository on your machine that you are free to change
Time required: 15 to 30 minutes (it varies with the size of the task)
Steps
- You pick 1 task. There are 2 conditions: it is small enough to finish within 30 minutes, and there are several roads through the implementation (for example, adding a sort order to an existing list). On a task with only 1 road, changing a setting or adding a file in a fixed form, this check shows no difference
- You hand it to the 1st agent with the how written out as well. Write the procedure you think is best, the way you always do ("fix A first, then B, and use this function")
- You hand the same task to the 2nd agent with only the result, the constraints and the way to check. You do not change a single character of the wording of the task from the 1st, and you replace only the part where you wrote the procedure with these 3: what has to be there for the job to be done (the result), the places not to touch (the constraints), and which automated test has to pass (the way to check)
- You line up the list of files the 2 agents touched
What to look at (this is not a pass or a fail. You are looking at whether a difference shows up)
- A file that is not in the 1st and is in the 2nd — that is the road your procedure had closed off in advance
- There is not one such file — your procedure was the best one for that task. That is a result too
What not to look at
- Which was faster, which has fewer lines, whose work you like better
- All this check can measure is the range of the search, not how good the work is
Cleanup
- Delete what the 2 agents produced, if you are not taking it in
Do not run the 1st and the 2nd one after the other in the same session.
The procedure you handed to the 1st stays in the 2nd one's context.
And do not ask the 1st or the 2nd, the ones that did the work, "which way of handing it over was better?"
The one answering is the very AI that wrote under that way of handing it over. Having it mark its own work is not measuring. The list of files you lined up in step 4 is all the material the judgment needs.
Next time, the art of not writing up bug details. I do not explain to the AI what the bug is. Even so, it gets fixed.
Series: The Art of Not Telling
- ← Previous: 1. The art of not asking for verification
- → Next: 3. The art of not writing up bug details
- All articles: Introduction
End of CC BY 4.0