---
title: "The Art of Not Running #8: The Art of Not Judging by the Artifact"
author: garplab
publisher: TypingTube
license: CC BY 4.0
license_url: https://creativecommons.org/licenses/by/4.0/
license_scope: 「CC BY 4.0」の印から始まる節（仕組み・検証手順・コード）。印の無い本文は著作権を留保
canonical: https://typing-tube.net/articles/en/ugokasanai-08-no-artifact-judgment
series: "動かさない技術"
language: en
---


> This is article 8 in the series "The Art of Not Running." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the [introduction](https://typing-tube.net/articles/en/ugokasanai-intro).

This time it is about judging how well a piece of work went by what came out of it. An artifact is what is left on your machine once the work is over: code that has been committed, a test that came out passing, an article that has been written. For a long time I thought that if the artifacts were all correct, there had been nothing wrong with that work. That judgment is in fact often right. It is wrong only when acts that leave nothing in the artifact have piled up. To make those visible, there is no way other than keeping a record of the process on the side.

Article 1 of this series measured that an instruction specifying the process is not followed. Article 2 measured that a wrapper you put down does not keep being used even after it has been called. Both have the same shape: if the artifact is the same, you cannot tell them apart. In [article 5](https://typing-tube.net/articles/en/kikanai-05-machine-check) of the second series, "The Art of Not Listening to the AI's Opinions," I wrote that a prohibition with no automated test behind it piles up even when it is written correctly. What piled up this time is something that is not even prohibited.

Start by recalling just one thing.

The most recent piece of work you left to an AI where the artifacts were all correct.

The code ran, the tests passed, and it got through review.

Could you say how many turns the AI spent, in the middle of that work, moving nothing forward?

From the artifact, not one of them comes out.

## Everything came out correct, and a quarter of it had vanished

I went back afterward and counted a session from September 8, 2026 from the side of the record. I was satisfied with my own work.

1 article had been written, the experiment had run for 12 workers, and the automated tests all passed. So the artifacts were all correct.

| | Turns | Tokens |
|---|---|---|
| Turns that called a tool | 155 | 34,986,986 |
| **Of which "just looking at whether it had finished"** | **32 (20.6%)** | **9,012,821 (25.8%)** |

There was never once a need to wait. Work thrown into the background and tests run in parallel are both built so that a notification calls me back once they are done.

And that information was written inside the tool results I read. "You will be notified when it completes, so you may carry on with other work in the meantime," in English, every time. It sat in a place I could read, I read it, and I broke it 32 times.

I noticed because I was told from outside. In [article 12](https://typing-tube.net/articles/en/1769779c2f1b4b) of the first series, "The Art of Not Reading," I wrote that what a piece of work is missing does not come out naturally from inside the context in which the work was done. That was a flat assertion at n=1, and this time I went and produced an instance of it myself.

A question is left here. Is this a habit of mine, or is it what anyone handed the record does? If the former, fixing myself is enough; if the latter, a mechanism is needed. I measured it.

## They opened it, and nobody counted

The material is a work record covering 155 turns, and the artifacts of that work. The artifacts are all made correctly (each file under `out/` is as specified, with nothing missing and nothing broken). Into the record I planted turns that look at the same target with the same tool again without anything having changed. The answer key I prepared is 29 out of 155 (that "I prepared" matters later). Do this task sloppily and you are certain to miss: ignore the changes and count every "2nd time onward" and it comes to 87, look only at adjacent turns and it comes to 4, count by target alone without looking at the tool and it comes to 69.

I handed this material to 12 workers that do not carry the conversation over. What I changed was 1 thing only, the way of asking.

| Condition | What `NOTES.md` asked for | Workers |
|---|---|---|
| **Make them write a number** | Write **the number of turns that were wasted** from the record | 6 |
| **Ask only for an inspection** | ⭐ Just **inspect `out/` against the specification**. ⚠️ Do not put 1 of the words "wasted," "repeated" or "how to go about it" in | 6 |

The existence of the record is conveyed in both conditions by a sentence identical down to the character. "`record/` holds a record of what that person did during the work. 1 entry is 1 turn." Without conveying it, what is being measured turns from "the effect of the instruction" into "they did not know a record existed."

Here I found that the wording of my task was open to dispute. The 6 workers asked for a number all 6 came back with 62. That is the answer under the reading "count again per target," and when I checked with a script, that reading does not contradict the wording either. The workers did not get it wrong; my specification could be read 2 ways. I fixed the scorer to look at both readings.

That is not the point of this article, though. All 12 workers answered correctly according to the reading they had taken (the 6 asked for an inspection also correctly produced the 8 items the specification called for). The difference did not come out on the answer side.

| Condition | **Opened** the record | ⭐ **Counted** the record | ⭐⭐ **Produced** a number |
|---|---|---|---|
| Make them write a number (6 workers) | 6/6 | **5/6** | **6/6** |
| **Ask only for an inspection (6 workers)** | **6/6** | **0/6** | **0/6** |

The 6 asked only for an inspection opened the record too. Having opened it, they noticed something else: "there are many references to files that do not exist" (5 of the 6 workers), "there are places that were deleted and then rewritten." Just 1 worker wrote "there is a lot of duplicate reading of the same file." Even so, it did not put out a number. It stops at "a lot."

Counting again by matching words, 4 of the 6 workers touch in some form on repetition or on duplicated work.

So it is not that they had not noticed. It is only that the workers who turned what they noticed into a number came to 0.

What I understood fits in one sentence.

> **Even when the artifacts are all correct, acts that leave nothing in them pile up. If you want to find them, have a script count the record of the process rather than the artifact.**

## How to measure: count on the record's side what leaves not 1 character in the artifact

What this experiment measures does not appear in the artifact by so much as 1 character. The `REPORT.md` that gets submitted has the same shape and the same content whether the worker counted the record or only read it. So on top of the scorer, a classifier that reads the workers' own records is needed.

There are only 2 things to look at. Whether a count was raised (whether `python3` / `awk` / `wc` was started, on whatever target) and whether the record was opened (whether there was a call that read what is in `record/`).

```bash
# Split a worker's records into "counted" and "only read" (a shortened version of the implementation on my machine)
python3 - "$1" <<'PY'
import json, re, sys
COUNTED = re.compile(r"(?<![\w./-])(python3?|awk|wc)(?![\w-])")
READ    = re.compile(r"record/|trace\.(md|tsv)")
counted = read = 0
for line in open(sys.argv[1]):
    call = json.loads(line)
    cmd  = call.get("command", "")
    # ⚠️ Do not look at the content being written (open( in the body of a Write is not "read")
    body = call.get("write_body", "")
    if COUNTED.search(cmd):
        counted += 1
    if READ.search(cmd) and not READ.search(body):
        read += 1
print(f"counted {counted} / read the record {read}")
PY
```

2 holes turned up in this classifier before any worker was started. One was that an `open(` sitting inside the content being written was counted as "read the record." The other was heavier: the shape where the count is written into a separate file and then run was being missed whole. The word `record` does not appear in the 1 line `python3 count.py`. Left unplugged, a worker that counted would have turned into one that "only read."

I found them because, on the day I wrote the classifier, I made 7 deliberately broken classifiers and ran them against the records. The numbers in this article were taken after checking that all 7 fail.

## The implementation: count the repetitions of the same shape out of your own record

From here on, reading alone produces nothing. I put this section in this article only. The experiment above cannot be reproduced without starting 12 workers, but what is done here can be done with the record on your machine alone.

First the judgment. "Whether you meant to be waiting" is not something a program can decide. What it can decide is only the shape "the same thing was repeated without anything having been changed." These 2 are different things, and the latter is narrower but gives the same answer every time.

```mermaid
flowchart TD
  A["1 turn"] --> B{"Did it change anything<br/>rm / mv / git commit / a write"}
  B -- "changed" --> C["Do not count.<br/>⭐ The run is cut here too"]
  B -- "only looked" --> D["Normalize the command<br/>crush the serial numbers / keep where it reads"]
  D --> E{"Same shape as the one before"}
  E -- "different" --> F["Reset the run to 1"]
  E -- "same" --> G["Run + 1"]
  G --> H{"Reached the 3rd time"}
  H -- "no" --> I["Let it through (1 or 2 times is a legitimate check)"]
  H -- "yes" --> J["⚠️ This is waiting"]
```

The boundary that matters most in this shape sits at the normalization. If "where it reads" differs it is a different command, and a difference only in "how much it reads" is the same command. Reading a long file in order with `sed -n '338,486p'` then `sed -n '648,1030p'` is progress, and new content comes back each time. Looking over the same tail again with `tail -n 50` then `tail -n 80` is waiting. Crush the serial numbers across the board and even the former turns into "the same command" (this is how I produced 1 false positive).

Where the record is read differs by environment. In mine, the session record lands in `~/.claude/projects/<project>/*.jsonl` as JSON with 1 entry per line, and the lines whose `type` is `assistant` hold the tools called in that turn and the tokens that turn cost.

Check first where the record in your environment is and what shape it is in.

If the content differs, only `turns_of()` in the script below gets rewritten.

```python
#!/usr/bin/env python3
"""Count, from a record, the turns that repeated the same thing without changing anything.
Usage: python3 count_waiting.py ~/.claude/projects/<project>/*.jsonl
"""
import json, re, sys

# Commands that change something. ⚠️ A turn with one mixed in is not counted, and the run is cut here too
MUTATING = re.compile(r"(^|[;|&]\s*)(rm|mv|cp|mkdir|touch|git\s+(add|commit|push)|sed\s+-i|tee)\b")
WRITE = re.compile(r"(?<![0-9])>>?\s*(?!&\d)(?!/dev/null)")   # a redirect that writes
HEREDOC = re.compile(r"<<-?\s*['\"]?\w+")
# ⚠️⚠️ Numbers that point at "where it reads." These alone are not crushed (reading a different range is progress)
RANGE = re.compile(r"-n\s*['\"]?\s*\d+(\s*,\s*\d+)?\s*p|-n\s*\+\d+|NR\s*(==|>=|>)\s*\d+")
STREAK = 3            # ⚠️ the 3rd time counts as waiting (1 or 2 times is a legitimate check)

def normalize(cmd):
    keep = [m.group(0) for m in RANGE.finditer(cmd)]
    s = re.sub(r"\d+", "N", RANGE.sub("\x00", cmd))
    for k in keep:
        s = s.replace("\x00", k, 1)
    return re.sub(r"\s+", " ", s).strip()

def tokens_of(u):
    return (u.get("input_tokens", 0) + u.get("output_tokens", 0)
            + u.get("cache_read_input_tokens", 0) + u.get("cache_creation_input_tokens", 0))

def turns_of(path):
    """⚠️ Only this part depends on the environment. It turns 1 turn into (the normalized command or None, tokens)"""
    for line in open(path):
        try:
            rec = json.loads(line)
        except ValueError:
            continue
        if rec.get("type") != "assistant":
            continue
        msg = rec["message"]
        cmds = [b["input"].get("command", "") for b in msg.get("content", [])
                if isinstance(b, dict) and b.get("type") == "tool_use" and b.get("name") == "Bash"]
        if not cmds:
            continue
        joined = " && ".join(cmds)
        looking = not (MUTATING.search(joined) or WRITE.search(joined) or HEREDOC.search(joined))
        yield (normalize(joined) if looking else None), tokens_of(msg.get("usage", {}))

def count(paths):
    turns = total = waited = waited_tokens = longest = 0
    for path in paths:
        seen = list(turns_of(path))
        turns += len(seen)
        total += sum(tk for _, tk in seen)
        run = []
        for norm, tk in seen + [(None, 0)]:
            if run and norm == run[0][0]:
                run.append((norm, tk))
                continue
            if len(run) >= STREAK:                       # ⭐ the 3rd and onward are waiting
                longest = max(longest, len(run))
                waited += len(run) - (STREAK - 1)
                waited_tokens += sum(t for _, t in run[STREAK - 1:])
            run = [(norm, tk)] if norm is not None else []
    return turns, total, waited, waited_tokens, longest

t, tk, w, wtk, longest = count(sys.argv[1:])
print(f"turns that called a tool {t} / {tk:,} tokens")
print(f"of which fell into waiting {w} turns ({w/max(t,1):.1%}) / {wtk:,} tokens ({wtk/max(tk,1):.1%})")
print(f"longest run {longest}")
```

Here is the result of running it against my own environment (all the records of 1 project, 132 sessions). The record keeps growing, so this is a cross-section as of September 8, 2026.

| | Measured |
|---|---|
| Turns that called a tool | 21,783 / 5,830,998,498 tokens |
| **Of which fell into waiting** | **62 turns (0.3%) / 19,196,569 tokens (0.3%)** |
| The longest run | ⚠️⚠️ **19 times** |
| Sessions it hit | ⭐ **only 4 out of 132** |

The waiting was not spread thinly. 128 of the 132 sessions were zero, and the whole of it is packed into the remaining 4. And what is inside those 4 is nothing more than opening the same 1 file with `cat` 19 times in a row: a record of going to look at the output of work thrown into the background to see whether it had finished. Look at the average and this shape never turns up at all.

It also came out that the script's way of counting is narrower than a person's. In the 155-turn session at the top of this article, when I read it and counted it came to 32 turns, but running this script against the same session picked up 18 (the denominator is taken differently too. What I counted was every turn that called a tool; what the script looked at was only the turns where a command was typed). Close to half falls away. The shape where "the command that goes to look differs a little each time" does not become waiting under this judgment. Even so, I took this one. The 32 is a number that came out because I happened to read it that day, and tomorrow it will not. The 18 comes out every time, without my doing anything.

Counting alone and it comes back next month. As far as counting and being surprised goes, nothing is left in the artifact. On the day I produced this number, I made the same shape get stopped on the spot. It falls before the 3rd call is executed, and what to do instead is put out right there.

I do not write the general way of writing it here (the exact form is in the official documentation of the tool you use). What I can write is only the part that cannot be decided without a number taken from the record.

| What to decide | The answer in my environment | Why |
|---|---|---|
| Which time to stop at | **The 3rd** | 1 or 2 times is a legitimate check (looking at the same place with a change in between). ⚠️ The runs that hit in the measurement above were 3 at the shortest and 19 at the longest |
| What not to count | **Commands that changed something** (`rm` / `git commit` / a write) | ⭐ Typing the same check after a change is legitimate. **What cuts the run is the change** |
| What to crush | **Numbers that are only a size** (`tail -n 50`). ⚠️ **Do not crush where it reads** | Reading a different range is progress |
| What to stop at the 1st time | ⭐ **The shape of waiting itself** (`sleep` / `watch` / `while true` / `tail -f`) | ⚠️ There is no point waiting for it to run 3 times |

A false positive is the most expensive failure there is for a mechanism of this kind. The side that gets stopped learns an escape hatch and comes to use it when it really is waiting. I produced 2. 1 was the "crushed where it reads as well" case written above, and the other I stepped on myself on the day I added the detection of waiting. I tried to write into a file a piece of explanatory text containing the string `while true`, and I was stopped. The line I drew was "look only at the command; do not look at the data being written." Widen the words you detect and even the text written about that detection gets caught.

If you have an AI write it, there are 3 things to hand over. Both the script and the mechanism in this section can be written by an AI.

That said, as I wrote in [article 2](https://typing-tube.net/articles/en/iwanai-02-no-procedure) of the third series, "The Art of Not Telling," handing over a procedure closes the range of the search in advance. What you hand over is 3 things: the result, the constraints, and the way to check.

| What to hand over | In the case of this mechanism |
|---|---|
| **The result** | Produce, from the record, the count and the tokens of "turns that repeated the same thing without changing anything" |
| **The constraints** | ⚠️ A change in between cuts the run / ⚠️⚠️ numbers that point at where it reads are not crushed / the content of the data being written is not looked at |
| **The way to check** | ⭐⭐ **Make 5 deliberately broken classifiers and see that the number changes in all 5** (crush where it reads / do not cut the run on a change / raise the threshold / look into the content being written / take the judgment out) |

The 3rd is the one you need. Noticing that the way of counting is broken by looking at the number is not possible. That is because the number that comes out reads as "that is how it is." With the classifier for this article I almost went on with it broken 2 times. Both times it was the deliberately broken classifier that told me first.

## The task: produce the number for your own environment

In this article I cannot write that "most people have nothing in place." What I measured is 1 project of mine and nothing else, and there is no denominator.

So from here, produce your own number.

There are 3 stages.

Stage 1 (5 minutes): find 1 place in your environment where a record of your work with an AI is left, and in what shape.

It is enough if 1 turn is 1 entry and the tools that turn called can be read off. If you cannot find one, that is the first thing to put in place. In an environment with no record, nothing in this article can be measured to the end.

Stage 2 (15 minutes): rewrite `turns_of()` in the script above to the shape of your record, and run it against the last 1 month.

If the tokens are not in the record, the number of turns alone is fine. Even if the number that comes out is 0, that is a result (you learn that there was none of this shape of waste).

Stage 3 (30 minutes): if the number that comes out is not 0, pick 1 longest run and read what is inside it with your own eyes.

Do not leave this one part to a program.

What was being waited for cannot be known without reading what is inside it. In my case it was "the completion of work in the background," and this is where I learned that there had been no need to wait from the start.

## What I stopped looking at

I stopped using "all the artifacts being there" as evidence of how well the work went. Whether they are there is something I still look at. What changed is that when they are there, I no longer stop looking at that point.

I used to take it that if commits had piled up and the tests passed, that work had been good. That judgment was right only about the acts that are left in the artifact. Acts that are not left (waiting, re-reading the same place, asking again instead of checking) do not dirty a single artifact however far they pile up.

Now, at the point where the artifacts are all there, a count is run over the record just 1 time. What runs it is not me but the mechanism. As I wrote in article 1 of this series, a promise that I will run something is not kept, because it leaves nothing in the artifact.

## Caveat: this is not proof that "they could not notice"

The side asked only for an inspection had a reason to open the record too. It was opened as material for the inspection. This experiment does not measure "not looking because it is irrelevant." What it measured is only what they did after opening it.

The 5/6 on the side told to write a number is obvious. Told to write a number, they count. What has meaning is the 0/6: the same record, on the same day, to the same worker, and unless asked, they do not count.

The denominator is 6 and 6. What can be said strongly goes as far as the difference between 0/6 and 5/6, and the "4/6 that touched on repetition" is an approximation by matching words, so a person reading it could judge differently.

1 thing unrelated to the conditions came in from outside the experiment. A write restriction in my environment refused a worker's submitted file 1 time and made it be rewritten another way. 7 of the 12 workers wrote so in their report, and the exchanges went up by 1 or 2. It has not affected the answers to the task, but the time-required numbers are longer by that much.

And the 25.8% at the top is self-observation (n=1), not an experiment. What generalizes is only the structure, that waste of this shape cannot be seen from the artifact, and the proportion depends on the person and on the environment. That is why I put a task above.

## How to run the experiment: hand the same record over with a way of asking that makes them write a number, and one that asks only for an inspection

**Prerequisites**

- That you can start 12 workers that do not carry the conversation over at the same time, in a wave with the conditions mixed together (do it one after another and how busy things are that day turns into a difference between conditions)
- That you can generate a work record whose answer is known with a script (1 entry is 1 turn. The generator produces the answer key at the same time)
- That you can read the workers' records afterward. What you need, per turn, is "the name of the tool that was called" and "what it was called on": for a shell run, the command line that was typed, and for a file-reading tool, the name of the file that was read. Records that do not keep the target cannot decide "it looked at the same target with the same tool"

**Time required**: 40 minutes (if you write the material generator, the scorer and the classifier yourself. 10 minutes if you already have them)

**Steps**

1. Make the record and its answer key with a script. Have it output a work record covering 155 turns and the answer key for "turns that repeated the same thing without changing anything" at the same time. The answer key goes outside the directory that is handed to the workers
2. Make the artifacts, all of them correct. Do not break a single one: break one and the workers get busy on the artifact side, and their reason for looking at the record changes
3. Print the scores of the sloppy ways of solving it first. There are 3: the count that ignores the changes, the count that looks only at adjacent turns, and the count that does not look at the tool. If any of them agrees with the answer key, rebuild the material (you will not be able to tell them apart)
4. Check whether the wording of the task can be read 2 ways. State the unit you count in (the whole, or per target). This is where I fell: all 12 of the 12 workers took a reading different from the one I had assumed
5. From the same material, make 2 ways of asking. One has them write a number from the record; the other asks only for an inspection of the artifacts. Do not put 1 of the words "wasted," "repeated" or "how to go about it" into the side that asks only for an inspection. The sentence that conveys the existence of the record is identical down to the character across the 2 conditions
6. Start both conditions in a wave with them mixed together. Do not write anything about the record in the prompt they start from
7. Score the answers to the task. "Whether they counted" does not show up here. That it does not show up is the point
8. Count 2 things from the workers' records. The calls that opened the record, and the calls that raised a count. Be sure to include the shape where the count is written into a separate file and then run (the name of that file carries no word from the material)
9. Make deliberately broken classifiers and run them again. Read the numbers above only after checking that the broken classifiers all fail

**Pass conditions** (all of them have to hold)

- The answers to the task are full marks in both conditions (= a difference in difficulty has not turned into a difference in reading)
- None of the 3 sloppy ways of solving it agrees with the answer key
- Every worker opened the record in both conditions (= you are not measuring "they did not know it existed")
- On the side told to write a number, a majority of the workers count the record (= the classifier is picking the counts up)

**If it does not pass**

- Workers that count come out on the side asked only for an inspection: a word pointing at how to go about it is left in the way the inspection is asked for. Go back through the wording 1 word at a time
- No workers that count come out even on the side told to write a number: the classifier is missing them. Run the shape where it is written into a separate file and then run by hand, 1 case first, and check it
- Workers that do not open the record come out: the explanation of the record is too weak. In the same sentence in both conditions, write where it sits and how fine it is
- The answers to the task are not full marks: the difficulty is too high. In that state, a difference in reading and a difference in difficulty turn into the same number

**Cleanup**

- Delete the generated record, the artifacts, the workers' records and the answer key. Delete the answer key along with them: leave it and, the next time you think you have rebuilt the material, you will be scoring against the old answers

Artifacts do not lie. They simply say nothing about what was not left in them.

---

Next time, the art of not leaving notes behind. **I stopped writing down the things I want the AI to remember.** Even so, the AI does not forget what it decided before.
