This is article 4 in the series "The Art of Not Running." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the introduction.
This time it is about the phenomenon where making the AI say something changes what it does afterward. The AI becomes the reader of what it said on the next turn, so saying it can be the first moment it notices something. I have been putting that to work on the judgment of whether something may be run. Yet when I measured it this time, of the 6 workers I tried to make say it, 1 said it. That is because the moment you write "say it when such and such happens," judging the condition becomes the other side's job.
In series 4, article 2, "The Art of Not Checking," I wrote that a mechanism that makes them write does work.
It works only when something gets written, though: that is what I wrote. What never got written leaves no record, so I have never once counted how many times it slipped past. This time I go and measure the side I have not counted.
Start by counting just one thing.
Open your instruction file (CLAUDE.md, AGENTS.md): it is how many lines in it start with "if needed," "if you get stuck," or "when it is an exception."
Can you recall when that line last worked?
Let me say this up front. This article is a report on something that was useful for a long time, up to the point where it stopped being needed.
And for what I found out this time, there is no mechanism I can hand over that works once you put it down. The previous 3 articles could at least end in the shape of "do it this way." What I found out this time is a property on the AI's side, and it is not something you can block off or delete. Knowing it changes how you write your instruction file: that is as far as what I can hand over goes. In its place, I put the procedure for getting the same numbers in your own environment at the end.
CC BY 4.0
What was going wrong, and what it has been replaced with
Let me clear up those 2 things first.
What was going wrong was that the AI kept trying to check things by looking at the screen with its eyes. Taking a screenshot, starting a browser and clicking through it. All of them are slow, and most of the time grep or a unit test is enough. I did not want to forbid it. I only wanted it to stop once before picking the expensive route.
So I made it write 2 questions before taking the shot. What it wants to check, and why the cheap means will not do. The moment it tries to write, it notices for itself that "grep is enough for this": that is what I was aiming at.
And I no longer use this form. I moved the same 2 questions into a hook (a mechanism that runs just before execution). If the file it wrote is not there, the command for taking the shot does not run at all. The reason I replaced it was not that "asking stopped working," so I write about that at the end.
The principle: you become the reader of what you said on the next turn
The phenomenon itself is simple. The AI becomes the reader of what it said on the next turn. So saying it can be the first moment it notices something, and what it does can change. The line about output changing when you have it write its intermediate thinking before the answer is about this.
There is a flip side to this principle, though. There is no guarantee that the reason that gets written is the real reason. Turpin et al. (2023) report that explanations of intermediate thinking can systematically misrepresent the true reason for a model's prediction. Fix the answer to (A) across a handful of demonstrations, and the model goes along with it while never touching on that in its explanation. Anthropic also reported in 2025 that when a reasoning model was given a hint, the rate at which the hint was mentioned inside the reasoning was low.
So the design leaves the correctness of what gets written out of the measurement from the start. Not because it cannot be measured, but because even if it could, it would not be evidence.
| What making them write guarantees | What it does not guarantee |
|---|---|
| That the writing step was passed through before the run | That the reason written is the real reason |
| That the point of judgment came before the run | That the judgment was correct |
That is the principle. 1 question is left: does the writing step get wedged in if you ask for it?
I changed the asking in 1 place and handed it to 12 workers
The material is the same as in the previous articles. There are 12 logs, and the work is to pull 1 value out of each. A tool for looking at the logs, ./bin/scan, is in place, and the instruction file has 1 sentence about it: "when you look at the contents of a log, use this." The tool does not show whether the start of a line is whitespace, so sorting the decoys out leaves an errand that sends you back to the bare search command. That is what article 2 measured.
Only 1 place was changed. To one side only, I added 2 sentences.
When
./bin/scanis not enough, you may use another way. Before that, though, write 1 line inDECLARE.mdabout why it is needed (1 line per time).
What these 2 sentences ask for, a step of writing 1 line of reason before the run, is what I call a declaration below (not the declaration of a variable).
This is not a stronger version of "use the tool." It is the reverse. It permits another way, and only wedges 1 writing step in ahead of it. The strengthening side was already measured in the previous article: even written with every technique from the previous 4 series, the ones that carried it all the way through were 3 of 4 workers. Add a strong sentence here and what is being measured goes back to "the strength of the instruction."
The total number of lines in the instruction file is lined up at 216 lines in both conditions. The specification of the task, the contents of the tool, and the 12 logs do not differ by 1 byte. That they are lined up is matched with a script at each generation.
I split 12 workers across the 2 conditions, 6 to each.
1 worker out of 6 wrote it
The answers to the task are full marks for all 12 workers. The values 12/12, the sorting of the decoys 2/2, and the entries embedded in 3 spots of the instruction file 12/12 as well. On the artifact side, not 1 difference between the conditions shows up.
Where the difference showed up was in the record of the work.
| Workers that looked at the logs with raw tools | Calls to the raw tools | Workers that wrote | Workers whose writing step came first | |
|---|---|---|---|---|
| The 6 workers that were not asked | 5 workers | 10 calls | — | — |
| The 6 workers asked to write first | 3 workers | 5 calls | 1 worker | 0 workers |
8 workers went back to another way, and of those, 1 wrote.
And even that 1 worker wrote after going back. It typed the bare search command on move 4, and wrote on move 7.
Across all 12 workers, not 1 had the writing step come before the run.
The DECLARE.md of the 1 worker that wrote came out like this.
The output of
bin/scanalready had the information the judgment needed, so no extra tool was really necessary. I ended up usingsed/odto double-check the judgment, and this declaration, which should have been written first, came after the fact.
It is an honest sentence.
And that honesty was itself the trouble. Read this 1 line alone and you cannot tell whether the order was kept or broken. What told me was that "move 4" and "move 7" were left in the record of the work.
The workers that did not write had no sense of having broken anything
3 of the remaining 5 workers never went back to the raw tools once. They had no errand to write about, so nothing being written is the correct state.
The problem is the 2 workers. These 2 looked at the logs with raw tools and wrote nothing.
And in their final report, they wrote this.
The output of
bin/scanalone gave enough information for the judgment, so no other means was needed.
bin/scanshowed me every line, so there was no need to declare an extra tool.
(The 2nd one came back in English.)
In fact, 1 of them pulled 20 lines out with sed and put them through cat -A, and the other did the same thing 2 times. Both named a log file directly.
I do not think this is a story about lying. For these 2 workers, checking 1 line with sed was probably not "looking at a log another way" but the continuation of the work they had done with ./bin/scan.
That is, they did not think they fell under the instruction's "when you use another way." I counted it as falling under it and they did not: that was the only place we were out of line.
An instruction that does not fire
Line it up against the wrapper from the previous article and the difference is plain. Both are 1 line in the instruction file.
| Instruction | Workers that complied even once |
|---|---|
When you look at a log, use ./bin/scan (previous article) | 4/4 |
| Before you use another way, write 1 line of reason (this article) | 1/6 |
The difference is neither strength nor position. It is who judges whether the condition has been met. The wrapper side is "if you look at a log you always use something," so the scene the instruction covers is bound to come. The other one is "when you use another way," so unless the other side judges that "this is me right now," it does not fire.
Drawn as a diagram, it comes out like this.
flowchart TB
S["A scene that needs another way"] --> Q{"Does this apply to me now<br/>⚠️ the AI is the one that judges"}
Q -- "Does not apply<br/>(the majority this time)" --> N["Says nothing"]
Q -- "Applies" --> W["Writes 1 line of reason<br/>= says it itself"]
N --> N2["What it reads on the next turn<br/>the progress of the work only"]
W --> W2["What it reads on the next turn<br/>the progress of the work<br/>+ the reason it just said itself"]
N2 --> X["The next move"]
W2 --> X
Only in the right lane does it gain 1 more of its own words. What it said becomes reading matter on the next turn, so it can notice something there and the next move can change. In the left lane that does not happen.
That said, I did not look inside to check whether it noticed. What is visible is only whether the words increased and what the next move was. So what I measure is the behavior side.
Lines of this shape are in most instruction files.
If needed, add a test
If you get stuck, ask first
When you use an exception, add the reason
Every one of them hands the judgment of "when does this fire" to the other side. From outside you cannot tell whether it was not complied with or whether it was judged not to apply.
And this time, the side that was judged not to apply was the majority.
What I understood fits in one sentence.
Put a condition on an instruction and the one who judges whether the condition was met becomes the other side. "If needed," "if you get stuck," "when it is an exception": a line of this shape does not fire unless the other side judges that it was met.
How to measure: change only whether you ask, and count the order in which things were written
What gets counted in this article splits into 2.
| What gets counted | Where it shows up |
|---|---|
| Whether it wrote | ⭐ left in the artifact as DECLARE.md |
| Whether it wrote before moving its hands | ⚠️⚠️ not left in the artifact (only in the record of the work) |
Look at the top row alone and it looks like a matter of "is it there or not." As for merely being there, it is there even if it was written afterward. The worth of the writing step lies in its coming before the hands move, so if the order cannot be measured there is nothing to say.
The order is counted from the record of the work. It sorts each tool call into 2 (did it write a reason, or did it look at a log another way) and lines up the move on which each first happened, and nothing more.
# Get the order of "the move that wrote the reason" and "the move that looked another way" out of the record of the work (a shortened version of the implementation on my machine)
import json, re, sys
WRITE_TOOLS = ("Write", "Edit")
READ_TOOLS = ("Read", "Grep", "Glob")
# ⚠️ Just reading is not a declaration. Count `cat DECLARE.md` and the order looks reversed
DECLARE_WRITE = re.compile(r">>?\s*[^\s;|&]*DECLARE|(?<![\w./-])tee(?![\w-])[^;|&\n]*DECLARE")
RAW_TOOL = re.compile(r"(?<![\w./-])(grep|cat|head|tail|sed|awk|nl|od|perl)(?![\w-])")
TARGET = re.compile(r"\blogs\b|log_\d+\.txt")
# ⚠️⚠️ A form taken in a loop variable (`for f in logs/*; do sed -n 2p "$f"`) has no logs in that segment
VAR = re.compile(r"\$\{?\w")
def wrote_declaration(name, inp):
if name in WRITE_TOOLS:
return "DECLARE" in str(inp.get("file_path") or "")
if name in READ_TOOLS:
return False
# ⚠️ 1 call can hold several commands, so split it and look at them 1 at a time
return any(DECLARE_WRITE.search(seg) for seg in re.split(r"[;|\n]+|&&", str(inp.get("command") or "")))
def deviated(name, inp):
text = " ".join(str(inp.get(k) or "") for k in ("command", "file_path", "path"))
if not TARGET.search(text):
return False
if name in READ_TOOLS:
return True # ⚠️ The entrances ingrained in the bones are not only commands
rest = re.sub(r"(^|[;|&\n])\s*(\./)?bin/scan\b[^;|&\n]*", r"\1 ", text) # erase the wrapper calls, then look
return any(RAW_TOOL.search(seg) and (TARGET.search(seg) or VAR.search(seg))
for seg in re.split(r"[;|\n]+|&&", rest))
n, first = 0, {}
for line in open(sys.argv[1]):
content = (json.loads(line).get("message") or {}).get("content")
# ⚠️ content is sometimes a string (a line where the AI answered without calling a tool)
for block in content if isinstance(content, list) else []:
if block.get("type") != "tool_use":
continue
n += 1
name, inp = block.get("name"), block.get("input") or {}
for key, hit in (("wrote", wrote_declaration(name, inp)), ("reverted", deviated(name, inp))):
if hit and key not in first:
first[key] = n
if "reverted" not in first: print("never went back to another way")
elif "wrote" not in first: print(f"⚠️ reverted without writing (move {first['reverted']})")
elif first["wrote"] < first["reverted"]: print(f"wrote first (move {first['wrote']} → {first['reverted']})")
else: print(f"⚠️ reverted first (move {first['reverted']} → {first['wrote']})")
This sorting passes even when it is broken. A sorting that reads everything as "declared" and a sorting that reads everything as "did not" both get through an automated test that looks at only one side. So I made 6 broken versions of the sorting and check every time that the verdict really does change.
In fact, the sorting that counts cat DECLARE.md (just reading) as a write turned the 1 worker that wrote afterward into "wrote first."
What I stopped writing
It is putting lines that start with "if needed" or "when you do such and such" into the instruction file. I used to think this shape was just the right strength for work that is not worth forcing but is a problem when it is done without thinking. Now, when I want to write the same thing, I look first at who judges the condition.
| How it is written | Who judges | What is guaranteed |
|---|---|---|
| If needed, please do such and such | ⚠️ the other side | ⚠️ Nothing. Neither whether it complied nor whether it did not apply is left behind |
| When you do such and such, please do such and such (a scene that is bound to come) | ⭐ the scene | It complies once (4/4 measured in the previous article) |
| Make it a shape where the run does not start unless the file it wrote is there | ⭐⭐ the mechanism | That the writing step came before the run |
The 3rd row is not an instruction. If you want to wedge a writing step in, stop asking and make it a shape where you cannot go on unless it is wedged in: that is all this round of measurement says.
1 more thing. That it did not fire is something you will never learn from the artifact. The artifacts of all 12 workers were full marks, and not 1 character of difference between the conditions showed up. This is the same shape as up through article 3.
Caveat: judging that something does not work is the harder job
From here on it is all caveats. When I counted, this article had 80 or more caveat markers in it. The articles in series 1 through 4 have 2 to 4 each.
The reason is in the direction of the conclusion. "It worked" can be said once it happens 1 time, but "it did not work" turns over if even 1 condition is left unclosed. Maybe the denominator is too small. Maybe the direction is the other way. Maybe the line was drawn badly. Maybe it is the generation of the model. Maybe it is the effect of running them alongside each other. Those 5 are what I closed off in this article, and even so they are not closed off completely.
Telling that something is not useful takes more moves than telling that it is useful. What follows is those moves.
The side asked to write first had fewer workers go back to the raw tools (5 workers against 3, and 10 calls against 5). The denominator is 6, so this difference cannot be told apart from chance. In the previous experiment as well I decided not to read the 1/4 and 3/4 of a denominator of 4 as a "difference," and I treat this the same way.
And the direction is not explained. What I added was a permission (you may use another way), so read plainly, the workers going back to another way should have gone up. If they went down, then the 2 sentences were read as a signal that "this is a place to be careful." That is a different thing from the effect of the writing step. Separating them would need 1 more condition that adds only the permission and asks for no writing. I did not measure that this time.
The line between "applies" and "does not apply" is one I drew. I count a call where a bare search command was pointed at a log as 1 time.
And the finding that mattered most in this article was exactly outside that line. The workers did not see the same operation as applying. It is not a matter of which way of counting is right; it is that when you write a conditional instruction, "what you call the condition" is not lined up with the other side.
What I say about awareness comes from my own reading of the final reports. It is not something counted with a script, and it is 2 workers' worth. Nor, of course, is there any guarantee that the reason written in a report is the real reason. I wrote it anyway this time because the shape of the report and the record of the work disagreeing could be matched up with a script.
What I measured is 1 model at 1 point in time. It cannot be read as "it used to work and now it does not." There is nothing to compare against in the first place: the article in series 4 itself says that "how many runs skipped the writing step is still a guess," and no number showing that the asking form was working was ever taken.
If you suspect a difference of generation, the proper move is to put the same material to a different model, but I have not done that this time.
That it can be explained without bringing up the generation, though, is what the 4/4 and 1/6 above show. This difference is there within the same day, the same model and the same material.
Because there is a cap on how many run at the same time, I started the 12 workers in 5 waves (1 1 4 4 2). Only the first 2 workers ran on their own, 1 per condition. The rest ran at the same time with both conditions mixed together, but I cannot say the congestion was exactly the same.
When the previous experiment used the same setup, the workers that got all the way through on the tool alone were 1 of 4, and the side that was not asked this time was 1 of 6. The values are close, but they are a different wave.
How to run the experiment: put down 1 conditional instruction and count whether it fired
Prerequisites
- You can start several workers that do not carry the conversation over under the same condition (6 or more per side if possible)
- The worker's tool calls are left in the record 1 by 1. If they are not, this experiment does not hold up
- A task that has both a road you are bound to take if you do it plainly and an errand that takes you off it. Without the errand that takes you off, the scene for writing never happens at all
Time required: 30 minutes (10 minutes if you already have the material and the sorting script)
Steps
- Prepare 1 task and put 1 sentence of "please do it this way" into the instruction file. Do not write it strongly. It turns into an experiment that measures strength
- Prepare 2 copies of the same material and add 2 sentences to one of them only. "You may use another way. But before that, write 1 line of reason." Write it from the permission side (add "always use it" and you have 2 conditions)
- Line up the total number of lines in the instruction file. Cut as many harmless lines as you added, to match them
- Launch both conditions mixed into the same wave. Run one of them all the way through first and how busy things are that day turns into a difference between conditions
- Score the artifacts. If a difference shows up in the answers to the task, stop there. You would be reading a difference in difficulty as the effect of making them write
- Get 2 move numbers out of the record. The move on which it first looked another way, and the move on which it first wrote a reason. Do not count just reading (opening the file it wrote) as a write
- Sort into 4. Did not go back / wrote first / went back first / went back without writing
- Read the final report and match it against the record. This is the one place a program cannot decide. All you look at is whether something written as "did not" is there in the record
Pass conditions (all of them have to hold)
- The answers to the task are the same in both conditions (if there is a difference, what you are measuring is difficulty)
- The workers that went back to another way are not 0 (if it is 0, the scene for writing never happened. Rebuild the task)
- Break the sorting script in 1 place and the 4 verdicts change (if they do not, that script is not looking at anything)
- Check that the verdict changes with a sorting that counts "just reading" as a write
If it does not pass
- Only the side that was asked did worse on the task: the 2 sentences you added are too long, or they add work of their own. Cut them down to 1 line
- No worker went back: the tool you prepared can handle the whole task. Rebuild it so that an errand is left outside the tool
- Every worker came out "wrote first": a good result, but first look at whether the instruction spells the condition out too finely. If you have listed the scenes it applies to, that is not a writing step but a procedure manual
- The verdict does not change even when you break the script: that script is not looking at the record. An automated test with no broken sorting placed beside it goes through with everything passing
Cleanup
- Delete the directory handed to the workers and the records you gathered. Do not delete the records first: leave only the artifacts and the thing this experiment most wanted to see is gone
Of the 6 workers asked for a declaration, 1 wrote, and even that 1 wrote afterward. And all 6 returned artifacts with full marks.
On my own machine, I replaced something that had long worked with a hook
Back to the 2 questions from the opening. This is what I do now on typingtube, the material for this series.
To be honest up front, this form worked well for me for a long time. It is nothing more than 1 line in the instruction file: "before you do such and such, write the reason." The screenshots that were not needed visibly went down, and Nor has there ever been a time when I read a written declaration and stopped something.
I stopped using it not because it stopped working. It is because I replaced it with a hook and that side was too certain. Certain does not mean the contents get better; it only means that the run does not start unless something is written. The asking side finished its job there.
So why did it work on my machine? Lined up against the measurements in the body, the difference was 1 thing. The condition I put down was "when you run a screenshot or a smoke test," and that is a scene that is bound to come. The condition I put down in this experiment is "when you use another way," and here the other side decides for itself whether it applies. The declaration is the same, but who judges the condition was different.
The form itself is not something I came up with. I borrowed it from production change management: before you type a dangerous command, leave behind what you are going to do and why it is needed. A request for a privileged operation and a record of manual work in an emergency have the same shape. The difference is that theirs comes as a set with a human's approval, while this one goes ahead on the other side's self-declaration alone. So it looks at the trace rather than the contents.
It is put down only ahead of the work of checking things by looking at the screen.
The moment the AI tries to run a screenshot or a smoke test, a pre-run hook looks at 4 things and no more. Whether the file it wrote is there / whether the 2 questions are each filled in with 1 sentence / whether there is no forbidden word ("just in case," "for now") / whether the time it was written is not too old. It does not look at whether the contents are good in even 1 place (as in the principle above, that is because looking would not be evidence even if it did).
The same idea is also put down as 1 line on the instruction file side. Before you type a command on the production server, check the configuration. This one is an instruction rather than a mechanism, so it keeps exactly the weakness measured in the body. It is still there because the production configuration is not a condition a program can decide, and the practical answer is that instructions with no choice but to hand the judgment to the other side remain.
Making it work took 3 things. That there is an entrance you can catch by name (what you can catch is only the name of a command or a tool. Add an entrance and you add it to the mechanism too: forgetting to add it goes through without saying a word). That the screen where it was stopped has the next move written on it (stop the run alone and the check gets abandoned along with it).
And not handing the judgment of the condition to the other side: that is what the 12 workers taught me this time.
Next time, the art of not making the AI pick again. When I could not accept the AI's answer, I stopped replying "please pick again." Even so, the AI changes its mind when it ought to.
End of CC BY 4.0