---
title: "The Art of Not Running #2: The Art of Not Enforcing a Wrapper"
author: garplab
publisher: TypingTube
license: CC BY 4.0
license_url: https://creativecommons.org/licenses/by/4.0/
license_scope: 「CC BY 4.0」の印から始まる節（仕組み・検証手順・コード）。印の無い本文は著作権を留保
canonical: https://typing-tube.net/articles/en/ugokasanai-02-no-wrapper-enforcement
series: "動かさない技術"
language: en
---


> This is article 2 in the series "The Art of Not Running." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the [introduction](https://typing-tube.net/articles/en/ugokasanai-intro).

This time it is about making the AI use a wrapper of your own. A wrapper is a project-specific entrance put down in place of a standard command: `scripts/test.sh` put down in place of `rails test` is the shape of it. There are any number of reasons to put one down. You want to stop things running alongside each other, you want the environment variables lined up, you want a record left behind. And once it is down, most people write 1 line in the instruction file. "Run the tests with `scripts/test.sh`." That 1 line loses the moment the AI has an errand that takes it back to the raw tools.

In [article 1](https://typing-tube.net/articles/en/cd93e0c5cac832) of the first series, "The Art of Not Reading," I wrote this about that 1 line.

> Even if you write "use the wrapper" in your instruction file, what surfaces the moment it goes to run the tests is the standard command it has seen hundreds of millions of times. One line in an instruction file has to compete with a habit and a priority ingrained in the AI, and **it usually loses**.

I had not counted that "usually." The official documentation stands on the opposite side. What it gives as an example of something to write in the instruction file is "Bash commands the AI cannot guess," and the line is drawn on the assumption that writing it fills the gap.

The hypothesis is in the official documentation, and the counter-evidence is in my own earlier series. This time I go and measure it.

Start by looking for just one thing.

A line in your instruction file that says to use an entrance of your own in place of the standard command.

"Tests with `scripts/test.sh`." "Deploys with `bin/deploy`."

Once you find it, try to recall how many times that line was actually followed.

You should not be able to recall. That is because whichever one it runs with, what comes out is the same.

## "It is there" and "it gets called" turned out to be different things

I made material whose answer is known. I put 12 files that imitate logs under `logs/`, and the work is nothing more than picking 1 value out of each file and collecting them into `answers.tsv`. Only the ones with `RESULT: ` at the start of the line count; the same string inside a comment line or inside a quotation is a decoy, and some files have no valid line at all.

Into that I put 1 wrapper belonging to this project.

```sh
#!/bin/sh
# The entrance for looking inside logs/ in this project.
# Output is in the form <line number>:<line>.
exec grep -n 'RESULT:' "$@"
```

The contents are `grep -n` itself, and the output does not differ by 1 byte. That is deliberate. If the wrapper were the smarter one, "did it follow the instruction" and "did it get the task done" would turn into the same number. Keep the wrapper to one that gains you nothing by using it, and all that is left is "which of the two words it picked."

I put a 1-line explanation in `bin/README.txt` as well. The wrapper sits in the same place in all 3 conditions, with the same explanation attached. The only thing I changed is the "how to proceed" section of the instruction file.

| Condition | What went into "how to proceed" |
|---|---|
| **Write nothing** | ⚠️ The section itself is not there (the wrapper is in place, and the explanation can be read) |
| **1 line only** | "When you look at the contents of a log, use `./bin/scan <file>`" |
| **Write it strongly** | ⭐ An instruction using every way of writing that the previous 4 series called "effective" (described below) |

I handed it to 12 workers. The answers to the task came out 12/12 for all 12 workers, the decoys 2/2, and the positive control placed in 3 separate spots 12/12 as well. That is not where the difference showed up.

| Condition | Called the wrapper **even once** | Never went back to the bare command **all the way through** |
|---|---|---|
| Write nothing | **0/4** | 0/4 |
| 1 line only | **4/4** | **1/4** |
| Write it strongly | **4/4** | **3/4** |

The 4 workers that had nothing written never opened `bin/` once. It showed up in the output of `ls`, the explanation file sat somewhere they could read it, and still it was 0 times. Between a wrapper "being there" and a wrapper "getting called" there was no connection at all.

And write 1 line, and all 4 workers call it. What I wrote in the earlier series, that "1 line loses to the standard command," comes out wrong here. The 8 workers across the 2 conditions with something written read the contents of the wrapper with `cat bin/scan` first, without exception, and only then used it. A word it did not know was not ignored; it was checked.

The losing comes after that. In the 1 line only condition, 3 of the 4 workers went back to the bare command partway through.

What I understood fits in one sentence.

> **A wrapper gets called if you write it in the instruction file. What decides whether it keeps getting called is not the strength of the instruction but whether an errand that sends it back to the raw tools is still left.**

## How to measure: put down the same wrapper 3 ways, changing only the strength of the instruction

There is 1 condition, the strength of the instruction and nothing else. Into the "write it strongly" instruction I put every way of writing that the previous 4 series called "effective."

| Technique used | Source | What went in |
|---|---|---|
| Write it in the shape you want it done | [article 8](https://typing-tube.net/articles/en/iwanai-08-no-criticism) of the third series | Not "do not type a bare search command" but "please use `./bin/scan`" |
| Write the reason | series 2, [article 2](https://typing-tube.net/articles/en/kikanai-02-reason) | "How it gets typed changes with who types it, so afterward you cannot match up whether the same range was looked at" |
| Split the scope | series 2, [article 3](https://typing-tube.net/articles/en/kikanai-03-scope) | "This is for looking under `logs/`. It is not needed for `NOTES.md` or `answers.tsv`" |
| Give a positive example | the official guidance on format | 2 lines of `./bin/scan logs/log_01.txt` |

For the reason I wrote only things that are true inside this material. Write an advantage like "the wrapper is more accurate," and the reason for complying turns from "it followed the instruction" into "it judged that this was the more correct thing to do."

What was kept identical is confirmed with a script. The specification of the task, the wording of the 3 measured items, the total number of lines in the instruction file (216 lines in all 3 conditions), the contents of `bin/`, the contents of `logs/`. For 2 of the measured items I made even the line numbers the same (what gets cut is only the filler in the first half, so the only one that shifts is the first item). Length and position were already ruled out as conditions in the previous experiment, so that shift in 1 item does not explain the difference.

Whether the wrapper was called does not remain in the artifact. `answers.tsv` comes out the same whichever word it was read with. What gets counted is the workers' records.

```bash
# Count "went through the wrapper / looked with raw tools" from the worker's record (a shortened version of the implementation on my machine)
python3 - "$1" <<'PY'
import json, re, sys
SCAN = re.compile(r"(^|[;|&\n]|\bdo\s)(\s*(?:\./)?bin/scan\b[^;|&\n]*)")
RAW  = re.compile(r"(?<![\w./-])(grep|cat|head|tail|sed|awk)(?![\w-])")
scan = raw = 0
for line in open(sys.argv[1]):
    if "logs" not in line and "log_" not in line:
        continue
    if SCAN.search(line):
        scan += 1
    # ⚠️ Erase the wrapper call before looking for raw tools (both land on the same 1 line)
    rest = SCAN.sub(r"\1 ", line)
    if any(RAW.search(seg) and ("logs" in seg or "$" in seg) for seg in re.split(r"[;|\n]+|&&", rest)):
        raw += 1
print(f"wrapper {scan} / raw tools {raw}")
PY
```

Without that 1 line of "erase, then look," every number is off. At first I did not write it, and I was counting `ls logs && cat bin/scan` (which only read the contents of the wrapper) as "looked at logs with a bare command." A shape where several commands sit inside 1 call always turns into something else when you look at it all together.

## What I stopped rewriting

It is that 1 line, the one written for a wrapper that was not followed. I used to fix the line whenever I found a bare command being typed. Move its position further up. Add a reason. Put the prohibition in bold. In this round of measurement, a way of writing that did all of that still came to 3/4. It is up from the 1/4 of the "1 line only" way of writing, but with a denominator of 4 it cannot be told apart from chance, and above all it is not 4/4.

Now, instead of fixing the 1 line, I look at whether an errand that sends it back to the raw tools is still left. What the 3 workers that went back did was there in the records.

| Condition | Up until just before going back | What it typed after going back |
|---|---|---|
| 1 line only | Ran all 12 files through `./bin/scan` | `cat -A` / `sed -n 'l'` (to see the whitespace at the start of the line) |
| Write it strongly | Ran all 12 files through `./bin/scan` | `sed -n '2p' \| cat -A` (to check the 1 decoy line) |

The wrapper prints the lines containing `RESULT:`, but it does not show whether `RESULT: ` is at the start of the line. Sorting the decoys out needs that information. The moment they noticed it was missing they went back to the raw tools, and once back, they never came home.

The answer I gave in series 1, article 1 was to block off the standard command. I still think that is right.

There is 1 more place to look before you block it off, though. Block it off, and work that has an errand to go back for loses anywhere to go. Block the entrance while the wrapper can still do only part of the job, and the AI does the same thing as in the previous article. It just produces the same thing by another route that is open to it.

## Caveat: this is not proof that "the 1 line worked"

2 readings remain. The reason they went back to the raw tools was that the wrapper could do only part of the job.

So this design does not separate "the prior distribution beat the instruction" from "the wrapper fell short, so it reasonably used another tool." Separating them would need a condition where the wrapper does the decoy judgment as well.

That design makes the task easier, though, so compliance and correct answers turn into the same number. One of the two has to be given up, and this time I gave priority to "you gain nothing by using it."

The answers to the task are full marks for all 12 workers. The ceiling that had held through the 3rd run has not been broken at the 4th either. An experiment where the difficulty is pinned leaves you a narrower reading whether or not a difference shows up.

The denominator is 4 per condition. The gap between 0/4 and 4/4 looks large, but the gap between 1/4 and 3/4 cannot be told apart from chance. What this article can say strongly is only the former: write nothing and it is never called once.

And this experiment does not measure "which is better." The workers that read with a bare command and the workers that read with the wrapper all came out with full marks. What you want the wrapper used for is presumably not the answer but something other than the answer (a record, keeping things from running alongside each other, a uniform environment). If that "something" does not remain in what was produced, we are back to the previous article.

## How to run the experiment: hand over the same wrapper, changing only the strength of the instruction

**Prerequisites**

- You can launch 12 workers that do not carry the conversation over, all at once in a wave with the conditions mixed together (do it one after another and how busy things are that day turns into a difference between conditions)
- You can prepare 1 wrapper whose output is the same as the standard command's
- You can read the workers' records afterward. The command lines that were typed have to be left as they were, so that `./bin/scan logs/log_01.txt` and `grep -n 'RESULT:' logs/log_01.txt` can be told apart. For the turns that used a file-reading tool, the name of the file that was read has to be left as well

**Time required**: 25 minutes (if you write the generator and the classifier for the material yourself. 5 minutes if you already have them)

**Steps**

1. Make material whose answer is known with a script. Have it output the 12 files and their answer key at the same time. The answer key goes outside the directory that is handed to the workers
2. Put down 1 wrapper. Make it one whose output does not differ from the standard command's by 1 byte. Make the wrapper a smart one and compliance and correct answers become the same number
3. Put the explanation file next to the wrapper. Put the same one in all 3 conditions (without it, you end up measuring "it did not know the thing existed" rather than "the effect of the instruction")
4. Make 3 conditions out of the same instruction file. Write nothing / 1 line only / use every technique from the earlier series. Line up the total number of lines (adjust with filler, and print the line count rather than a hash to check them against each other)
5. Launch the 3 conditions in a mixed wave. Do not write anything about the wrapper in the prompt you launch them with
6. Score the answers to the task. The wrapper does not come up here. That it does not come up is the point
7. Count 2 things from the workers' records. Calls that went through the wrapper, and calls that looked at the target with raw tools. Erase the wrapper calls before looking for raw tools (both land inside the same 1 call)

**Pass conditions** (all of them have to hold)

- The answers to the task are full marks in all 3 conditions (= a difference in difficulty has not turned into a difference in compliance)
- The positive control (a 1-step instruction placed in 3 separate spots in the instruction file) is correct for everyone in all 3 conditions
- In the "write nothing" condition, the wrapper is called 0 times (= you have caught that a wrapper merely existing does not get it called)
- In the "1 line only" condition, 1 or more workers go back to the bare command

**If it does not pass**

- The wrapper was called even in the "write nothing" condition: the explanation file stands out too much. Move the wrapper's name away from the words of the task
- The wrapper is not called in any condition: the wrapper is less convenient than the standard command. Line up the shape of its arguments with the standard command
- Everyone got through on the wrapper alone in all 3 conditions: there is no errand that sends them back to the raw tools. Put 1 piece of information the wrapper does not show into what the task has to judge
- The answers to the task are not full marks: the difficulty is too high. In that state, the difference in compliance and the difference in difficulty turn into the same number

**Cleanup**

- Delete the material you generated, the workers' records, and the answer key. Delete the answer key along with the rest: leave it, and the next time, thinking you have rebuilt the material, you end up scoring against the old answers

Putting the wrapper down alone does not get it called; write 1 line and it is called; write strongly and it still does not get carried all the way through. Of these 3, the only one the instruction file could move was the first.

---

Next time, the art of not aggregating. **I stopped making the AI give the same answer every time.** And even so, the conclusion always settles in the same place.
