---
title: "The Art of Not Running #1: The Art of Not Enforcing a Procedure"
author: garplab
publisher: TypingTube
license: CC BY 4.0
license_url: https://creativecommons.org/licenses/by/4.0/
license_scope: 「CC BY 4.0」の印から始まる節（仕組み・検証手順・コード）。印の無い本文は著作権を留保
canonical: https://typing-tube.net/articles/en/ugokasanai-01-no-procedure-enforcement
series: "動かさない技術"
language: en
---


> This is article 1 in the series "The Art of Not Running." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the [introduction](https://typing-tube.net/articles/en/ugokasanai-intro).

This time it is about making the AI follow a procedure. A procedure is an instruction that specifies the way you go about the work rather than the thing that comes out of it, and "open each file one at a time and check it" is that shape. Instructions of this kind have something in common. Followed or broken, what comes out is the same, so you cannot tell afterward whether they were followed. If you want to tell, the process has to be left in the output: there is no way other than putting it into a shape you can count afterward.

In [article 6](https://typing-tube.net/articles/en/569d598941c00f) of the first series, "The Art of Not Reading," I wrote that a rule works only when it is read, and that an automated test or a hook fails even when it is not read. In [article 5](https://typing-tube.net/articles/en/iwanai-05-no-emphasis) of the third series, "The Art of Not Telling," I wrote that with emphasis you cannot measure whether it worked, and settled for looking only at how the count grows rather than at the effect. This time I go and measure that "cannot be measured" side.

Start by opening just one thing.

A line in your instruction file (CLAUDE.md, AGENTS.md) that specifies the way you go about the work. "One at a time." "In order." "Always check first."

Could you tell the runs that followed that line from the runs that did not, going by what came out?

## If you cannot tell them apart, it is the same as not measuring

For a long time I thought an important instruction would be followed if I put it somewhere prominent and wrote it short. The official documentation, too, has a passage to the effect that the more instructions you add, the more each single instruction is diluted. I read that as a matter of position and length. Move it up. Make it bold. Cut the lines that are not needed.

On the production product I cannot check whether that reading is right. That is because telling whether an instruction was followed needs a task whose answer is known in advance.

So I stepped away from the product and made the material.

In `logs/` I put down 12 files that imitate logs. The work is nothing more than picking 1 value out of each file and collecting them into `answers.tsv`. Only the lines that start with `RESULT: ` count; the same string inside a comment line or inside a quotation is a decoy, and some files have no valid line at all. The answer key is made by the script at the same time as the material, and goes outside the directory that is handed to the workers (if a human makes it afterward, it gets pulled toward the workers' answers).

The instruction for the work is 1 sheet, `NOTES.md`. Inside it I put 1 sentence about the way to go about it.

> Open the files **one at a time** and check what is in them.

I hand this to workers that do not carry the conversation over. This 1 sheet is all they get, and the prompt they start from says nothing about this 1 sentence. The moment it does, the thing I am measuring disappears.

The 1st experiment took "how many lines you pile into 1 prompt" as its condition. No difference came out. A worker reads 377 lines in 1 command. The length of a single read was not the cause of the dilution.

The 2nd took the amount of material as its condition (12 files / 60 files. `NOTES.md` is the same down to the byte across both conditions, and I check the first 12 digits of its sha256). All 8 workers got full marks on the task, and compliance with the "write a number" instruction placed in 3 spots was 24/24 as well. Not even the number of exchanges was governed by the condition. The median was 13.5 on the 12-file side and 12.0 on the 60-file side, and the 60-file condition produced both the shortest run at 7 exchanges and the longest at 67. Workers fold the material up with `grep` or a script, so its amount does not make the session longer.

Instead, 1 thing turned up. The only instruction that was broken was "open each file one at a time and check it."

| Instruction | Position | Moves needed to follow it | Compliance |
|---|---|---|---|
| Write 1 number (3 spots) | lines 31 / 127 / 219 | **write it once** | **24/24** |
| Open each file one at a time and check it | line 27 | **12 / 60** tool calls | **3/8** |

Even in the session that ran to 67 exchanges, the instruction on line 31 was followed. What got diluted was not position or length but the side that took more moves. Here I almost concluded that instructions fall away in order of how many moves they take. The 3rd experiment went to check that, and missed.

What I understood fits in one sentence.

> **An instruction that specifies the process is not followed when the artifact is the same. If you want it followed, put it into a shape where that process is left in the artifact.**

## How to measure: put 2 instructions of the same effort side by side

In the "how to go about it" section of `NOTES.md` I put 2 instructions side by side, each worth the same 12 items of effort.

| Name | Position | Scope | Trace of compliance | Content |
|---|---|---|---|---|
| **P** | line 27 | 12 items | **none left** | Open each file one at a time and check it |
| **A** | line 28 | 12 items | **left** | For each file, write that file's line count into `checked/log_NN.txt` |

Their positions are adjacent, their scope is the same 12 items, and the moves they take are the same 12. The only difference is whether a trace of compliance is left. The `logs/` I handed over is the same as in the 2nd experiment down to the byte (I check it with sha256). `NOTES.md` is 2 lines longer, by exactly the 2 lines A adds and the turning of P into a bullet.

Here is the result of handing it to 4 workers.

| Worker | Answers to the task | Decoys | **A (trace left)** | **P (no trace left)** | Exchanges | Seconds |
|---|---|---|---|---|---|---|
| 1 | 12/12 | 2/2 | **12/12 (line counts match too)** | **12/12 followed** | 34 | 80.0 |
| 2 | 12/12 | 2/2 | **12/12 (line counts match too)** | **12/12 followed** | 20 | 78.1 |
| 3 | 12/12 | 2/2 | **12/12 (line counts match too)** | **12/12 followed** | 17 | 67.6 |
| 4 | 12/12 | 2/2 | **12/12 (line counts match too)** | **12/12 followed** | 21 | 69.3 |

Across the 3 experiments, compliance with that same 1 sentence comes out like this.

| Experiment | How it was placed | Compliance with P |
|---|---|---|
| 2nd | P alone (12 files of material) | 2/4 |
| 2nd | P alone (60 files of material) | 1/4 |
| **3rd** | **A placed next to P** | **4/4** |

I changed neither the wording nor the position. Adding 1 line next to it that asks for "an artifact covering the same scope" turned 3/8 into 4/4.

P is invisible to the scorer (its being invisible is the point itself). What gets counted is the workers' records: I count the cases where 1 tool call named exactly 1 `log_NN.txt` and no more. A call that swept through them with `grep` or `for` names 0 files or several, so it does not go into the count.

```bash
# Count "did it open them one at a time" from a worker's records (a shortened version of the implementation on my machine)
python3 - "$1" <<'PY'
import json, re, sys
single = 0
for line in open(sys.argv[1]):
    hit = re.findall(r"log_\d+\.txt", line)
    if len(set(hit)) == 1:
        single += 1
print(f"one-at-a-time reads: {single}")
PY
```

## What I stopped rewriting

Instructions that do not get followed. I used to rewrite any line I found that had not been complied with. Move it up. Make it bold. Add "always." Cut the lines around it to reduce the dilution. In all 3 experiments, compliance moved with neither position nor length nor moves.

Now, instead of rewriting, I look at whether I can add 1 line next to it that asks for an artifact leaving a trace of that instruction having been followed. If I can, no cleverness in the wording is needed. If I cannot, I build the rest on the assumption that the instruction will not be followed.

The dichotomy in series 1, article 6 had a boundary. All a hook can do is a judgment whose truth a program can decide. "Did it open them one at a time" cannot be decided, so this instruction had nowhere to sit but the rule side: that is what I thought. Whether it can be decided is something you can create, by changing the output you ask for. The boundary was not a given; it was a line drawn from this side.

## Caveat: this is not "making it comply"

2 causes are still mixed together. A asks for 1 artifact per file, so following A satisfies P automatically. Whether "A raised compliance with P" or "P was satisfied for free thanks to A" cannot be separated by this design.

As a practical answer, I think that is exactly the answer. Writing it so that they cannot be separated is the right way to write it.

As a claim for an article, though, it gets weaker. What can be said goes as far as "all 4 workers that had A followed it," and the difference from 3/8 to 4/4 itself cannot be told apart from chance with a denominator of 4.

There is one more thing this experiment cannot say. Everyone got full marks on the task in all 3 experiments (16 workers across the 3, made up of 4, 8 and 4; counting only the 2nd and 3rd, whose numbers appear in this article, it is 12). Even with 5 kinds of decoys in there, the ceiling has not been broken. An experiment whose difficulty is pinned leaves a narrow reading whether or not a difference shows up. Next I raise the difficulty on the "does each item need a judgment of its own" side.

The direction in which it works is the one thing I can state from the machinery rather than from measurement. In the 2nd experiment, the 5 workers that swapped the instruction for a cheaper route all got full marks on the task. They did not break it; they made the same output by a cheaper route. When the output is the same, that is what happens.

## How to run the experiment: compare the same 1 sentence as an instruction that leaves a trace and as one that does not

**Prerequisites**

- You can start 4 or more workers that do not carry the conversation over at the same time, in the same wave (do it one after another and how busy things are that day turns into a difference between conditions)
- You can split the directory handed to the workers by condition
- You can read the workers' records afterward. What you need, for each call, is "the name of the tool that was called" and "what it was called on": for a shell run, the command line that was typed (`cat logs/log_01.txt`), and for a file-reading tool, the name of the file that was read (`logs/log_01.txt`). Records that keep only the name of the tool cannot measure this

**Time required**: 20 minutes (if you write the material generator and the scorer yourself. 5 minutes if you already have them)

**Steps**

1. Make material whose answer is known, with a script. Have it output 12 files and their answer key at the same time. The answer key goes outside the directory that is handed to the workers
2. Make 1 instruction file. Put in it the description of the task, and in 3 spots far apart an instruction to "write 1 number" (this is the positive control. If it fails, it is the work itself that is broken, not the condition)
3. From that same instruction file, make 2 conditions. One has only P in the section on how to go about it (the instruction that leaves no trace). The other has A next to P (same scope, same moves, and it leaves a trace). Change nothing else, down to the byte (print the hashes and check them)
4. Start both conditions in the same wave. Do not write anything about P or A in the prompt they start from
5. Score the answers to the task. P does not show up here. What shows up goes as far as A
6. Count P from the workers' records. Count the cases where 1 tool call named exactly 1 target file and no more

**Pass conditions** (all of them have to hold)

- The 3 spots from step 2 are correct for everyone in both conditions (= this rules out a failure of the work rather than a difference between conditions)
- The answers to the task are full marks in both conditions (= a difference in difficulty has not turned into a difference in compliance)
- In the condition handed only P, 1 or more workers do not follow P
- In the condition with P and A side by side, A is satisfied on every item

**If it does not pass**

- Everyone followed P even in the condition handed only P: it does not take enough moves. Increase the number of files in the material and raise the cost of opening them one at a time
- A is not satisfied on every item: the suspicion is that A's artifact overlaps with the answers to the task. Do not choose for A a value that comes out naturally in the course of producing the answers
- Nothing changes in either condition: the scope of P and A is out of line. Line the scope (how many items) and the position back up first
- The answers to the task are not full marks: the difficulty is too high. In that state, a difference in compliance and a difference in difficulty turn into the same number

**Cleanup**

- Delete the material you generated, the workers' records, and the answer key. Delete the answer key along with the rest: leave it, and the next time, thinking you have rebuilt the material, you end up scoring against the old answers

An instruction that leaves no trace of compliance is neither followed nor broken. What this experiment could not measure to the end was not the compliance rate but whether you can tell them apart.

---

Next time, the art of not enforcing a wrapper. **I stopped writing "always use it" about the commands we built ourselves.** And even so, the AI uses those commands of its own accord.
