This is article 1 in the series "The Art of Not Chasing." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the introduction.

This time it is about going on digging even after the answer is in. In the previous 2 series I wrote the opposite over and over. Series 4, "The Art of Not Checking," rebuilt the way of checking itself; series 5, "The Art of Not Running," measured the flat assertions of the previous 4 series again against problems whose answer is known. Measure, is what I have been writing. This time I write what looks like the exact reverse of that. Digging deeper does not get detected: because it reads a different place every time, not 1 of it is stopped by a mechanism that watches for repetition.

That said, what I forbid is not measuring. It is starting to measure without having decided where it ends. The problems I measured in series 5 were all "problems whose answer is known." That the answer is known means the end is decided before you start. This time is the record of a day when that had not been decided.

Start by recalling one thing.

The last time you tracked down a cause, at which round did that investigation end?

Was it the answer that came out that decided the end?

Or was it a condition you had decided before you started?

If you cannot recall the latter, this article is about you.

CC BY 4.0

Of 108 calls, 89 changed nothing

I counted a session from September 9, 2026 from the side of the record. That day, to carry the series on, I was reading the scripts for the experiments and the records of the past 8 rounds.

Calls
Tool calls133
Of which, calls that changed something28
The longest run of reads that went on changing nothing35

The top 2 moved while this article was being written too. When I started writing it was 108 and 19; when I finished writing it was 133 and 28. The only one that does not move is the 35, and that is because it is a stretch that had already ended before I started writing the article. Self-observation includes the very act that is observing.

The 35 run from the 6th call to the 40th. Over that stretch, not 1 byte of any file changed.

The first question was a single line: "Go ahead with the next task. Do you have a clear outlook on being able to finish it?" To answer on the outlook alone, a few calls are enough.

I read a different place every time. The records of the experiments, the README of a script, the code that writes the judgment, the implementation of the scorer, the actual past artifacts kept in storage. Every one of them returns new content each time. That is why not 1 reason to stop comes out from the inside.

The mechanism that stops waiting did not stop 1 of them

This environment already has a mechanism that stops waiting. Going to look at whether work thrown into the background has finished gets stopped once it has gone on 3 times in a row. In article 8 of the fifth series I measured that 25.8% of the tokens I used had vanished into "just looking at whether it had finished," so I made it a mechanism at that point.

[observed] Stopped by the waiting detector (whether the same shape went on 3 times): 0

Go on for 35 calls and still not 1 of them is stopped. That is because what it looks at is "whether the same look-only command went on 3 times." Digging deeper does not look like a repetition of the same thing. It reads a different place in a different file every time.

What is more, that escape hatch was one I opened myself. On the day I built the mechanism I put in a condition like this: "if where it reads differs, it is a different command." That is because reading a long file in order from the front is not waiting but progress, and new content comes back each time. This is the right condition, and without it false positives come out. The right condition had become, just as it was, the passage for going on digging.

⭐⭐⭐ That it did not stop is not a failure of the mechanism. It is only that what the mechanism looks at is different.

What the digging turned up was a finding that the article does not need 1 line of

Let me write where the 35 calls went. I found that the "gate" that decides whether the experiment may go on to the next round was failing, and I dug into why it fails. I read the code of the judgment, read how that judgment reads the record of the previous round, read the storage format of the record, and actually ran the scoring again.

What I found is that the gate stays failing whatever you do. That is because what the judgment reads is the "saved scoring result" of the previous round, and that number does not change even when I swap the yardstick out.

As a place to fix in the script, this is genuine.

And the article does not need 1 line of it. The material the article I was about to write needed was already all there before I started digging. The 35 calls were not used in order to check that.

The options were strung across the bottom of the hole I had dug

After the 35 calls, I put out 4 options to ask for a decision. Fix the gate / take it as failing / do only a different small fix / swap the task itself.

All 4 sit inside "what to do about the gate." Not 1 of them is the option "whether to carry this experiment forward at all."

The author did not pick an option; he turned the question itself down. "Do not answer whether the next experiment finishes; answer the work up to finishing the article." With that one line, I got back to where I had started for the first time.

In the finale of series 2, "The Art of Not Listening to the AI's Opinions," I wrote that you do not pick an option. That was about "do not pick from among the options the AI has laid out." This time the cost side of it came into view. To make the 4 strung across the bottom of the hole, 35 calls' worth of compute was used.

The 35 calls I dug became the grounds for rejecting the next proposal

This is the part of this article that bit hardest.

Right after I was stopped, a proposal came from the author. "How about writing an article called the art of not digging deeper." It is this article you are reading now.

I rejected it, laying out 3 reasons. The definition of the series is not set up that way. It goes against the rule for how titles are made. That material has already been settled as something another article takes on. All 3 come out of the records I had read that day.

And the central 1 of them was off. I had judged that, because the words "dig deeper" are negative from the start, it does not hold as advice. The advice that sits on the other side of taking the negative off is not there. It is "chase the cause to the end. Do not stop halfway." It was on the side that gets taught as an engineer's virtue.

Look closely and the remaining 2 were decisions the author can change as well. The side that had read the records was treating them as things that cannot be moved.

⭐⭐⭐ The more you dig, the stronger the frame you dug in becomes. 35 calls' worth of reading turns, just as it is, into 35 calls' worth of "premises that cannot be moved."

I think this is a shape you often see in meetings. The person who investigated it most closely rejects a new proposal on the grounds of what they investigated. There is no ill will anywhere. What was investigated is true, the citations are accurate, and the conclusion is off all the same.

In my case the reason it was off is clear. That is because the frame I built over the 35 calls did not have the option "throw out that frame itself" inside it.

What I understood fits in one sentence.

⭐⭐⭐ Digging deeper is not shaped like waiting. Because it reads a different place every time, not 1 of it is stopped by a detector that watches for repetition. What they have in common is only this: not 1 byte of the artifact has been changed.

The principle: the condition for stopping can only be decided before you start digging

While you are digging, the next 1 move looks like the last one. This is not a matter of mood but a matter of structure. Each time you read, new content comes back, and each time, "1 more read and I will know" is updated. What keeps being updated cannot be stopped by the very person it is updating.

The record of the same day holds the side that did stop as well. The 8th experiment, run the day before, was built around the question "does the sign that came out on the 7th survive even when the denominator is doubled." It is over at the point the answer comes out.

In fact I ran 12 workers, checked that the sign survives, and stopped there.

The difference is not the tool and not the person. It is only whether, before starting, it had been decided what state counts as the end.

This is where it does not contradict the previous 2 series. What I was measuring in series 5 was all "problems whose answer is known." Preparing the answer first is a discipline of experiment and at the same time a device for stopping. What I did not do this time is not the experiment but that device.

How to measure: count by whether it changed something, not by whether it is the same shape

The detector that watches for waiting looks at "whether the same command went on 3 times." To watch for digging deeper, you shift where you look by 1: I do not look at whether it is the same shape, only at whether that call changed something.

Waiting⭐ Digging deeper
How it looksGoes on looking at the same placeReads a different place every time
DetectorThe count of same shapes in a row after normalizing⭐ The count of reads in a row with no change in between
In common⚠️⚠️ Neither changes 1 byte of the artifact

"Whether you have dug too far" is not something a program can decide. What it can decide is only the shape "how many times it read while changing nothing." These 2 are different things, and the latter is narrower but gives the same answer every time.

flowchart TD
  A["1 call"] --> B{"Did it change anything<br/>a write / rm / mv / git commit"}
  B -- "changed" --> C["Reset the run to 0"]
  B -- "only read" --> D["Run + 1<br/>⭐ Whether it is the same place is not looked at"]
  D --> E{"Did it pass the threshold"}
  E -- "no" --> F["Let it through"]
  E -- "yes" --> G["⚠️ This is digging deeper"]

The threshold changes with the environment. On my machine 35 is the measured value, and because there was not 1 reason to stop inside it, the threshold ends up placed far short of that.

The implementation: count, out of your own record, the runs of reads with no change in between

Where the record sits differs by environment. In mine, the session record lands in ~/.claude/projects/<project>/*.jsonl as JSON with 1 entry per line, and the lines whose type is assistant hold the tools called in that turn.

Check first where the record in your environment is and what shape it is in.

#!/usr/bin/env python3
"""Count, from a record, the runs of reads that went on without changing anything.
Usage: python3 count_digging.py ~/.claude/projects/<project>/*.jsonl
"""
import json, re, sys

# ⚠️ Calls that change something. The run is cut here
MUT = re.compile(
    r"(^|[;|&]\s*)(rm|mv|cp|mkdir|touch|git\s+(add|commit|push)|sed\s+-i|tee)\b"
    r"|(?<![0-9])>>?\s*(?!&\d)(?!/dev/null)"          # a redirect that writes
)

calls = []
for path in sys.argv[1:]:
    for line in open(path, encoding="utf-8"):
        if not line.strip():
            continue
        row = json.loads(line)
        if row.get("type") != "assistant":
            continue
        for c in (row.get("message") or {}).get("content") or []:
            if isinstance(c, dict) and c.get("type") == "tool_use":
                cmd = (c.get("input") or {}).get("command") or ""
                calls.append(bool(MUT.search(cmd)))

streak = best = best_end = 0
for i, mutating in enumerate(calls, 1):
    streak = 0 if mutating else streak + 1
    if streak > best:
        best, best_end = streak, i

print(f"calls {len(calls)} / changed something {sum(calls)}")
print(f"longest run of reads that changed nothing: {best}"
      f" (call {best_end - best + 1} to {best_end})")

There is 1 condition I have deliberately left out of this way of counting. It does not look at "whether that read was useful." That is because whether it was useful is not known even to the person reading, until they have finished reading. What it looks at is kept to the shape alone.

1 more thing. This detector does not ask where the write went. A tallying script written out to a temporary file is counted as "changed" too. Narrowing it to the artifact alone makes it accurate, but then a definition of "artifact" has to be written for each environment, and writing that definition is itself the entrance to digging deeper.

The task: produce the number for your own environment

In this article I cannot write that "anyone digs for 35 calls." What I measured is 1 session of mine and nothing else, so there is no denominator.

So from here, produce your own number.

Run the script above against the record of your most recent session.

There are 2 numbers that come out.

  1. The longest run of reads that went on changing nothing
  2. What question that run started against (open the relevant place in the record and you will see)

Be sure to look at 2.

Look at the number alone and it ends at "mine is 12, so it is light." What you should look at is whether the question those 12 were answering is the same as the question you set at the start. My 35 calls were answering a question that had sprouted partway through, while I was digging.

What I stopped digging deeper into

  • I stopped tracking down on the spot the reason a judgment fails. I leave only the fact that it failed in the record, and hand it on to the next decision
  • I stopped reading on "1 more read and I will know." Before reading that 1, I write first what reading it will decide. If I cannot write it, I do not read it
  • Before laying out options, I look at whether there is an option outside the hole I am digging now. If all 4 are in the same hole, that is not a set of options but a report of how deep I have dug

Caveat: this is not an experiment

The numbers above are self-observation (n=1). They are not a controlled experiment. What generalizes is only the structure, that consumption of this shape cannot be seen from the artifact by so much as 1 character, and the number 35 depends on the person, on the environment and on the work. That is why I put a task above.

There is 1 more thing this article cannot say. I cannot say that the 35 calls of digging deeper were completely wasted.

In fact, 1 place to fix in the script was found. What can be said goes only as far as the allocation of time: "those 35 calls were not used on the question that should have been answered then."

Judge the value of what was found, and the decision to put resources into it, separately.

How to verify: line up the round that stopped and the round that did not, with the same script

Here is the procedure for when you want to get past n=1. What you prepare is not workers but 2 requests that set the question up differently.

Prerequisites

  • That you can start 12 workers that do not carry the conversation over at the same time, in a wave with the conditions mixed together (do it one after another and how busy things are that day turns into a difference between conditions)
  • That you can prepare material whose cause settles on 1 thing. Make it something whose correct answer you can write yourself in advance (with material you cannot write it for, you cannot tell stopping apart from giving up)
  • That you can read the workers' records afterward. What you need, per call, is "the tool that was called" and "the command that was typed": records that do not keep the command cannot decide "it changed nothing"

Time required: 30 minutes (if you write the material and the scoring yourself. 10 minutes if you already have them)

Steps

  1. Prepare the same material (a record or code whose cause settles on 1 thing). What is handed to the 2 conditions is the same down to the byte
  2. Make 2 versions of the request. What changes is only the 1 sentence that gives the condition for ending
  3. The request for condition A is only "track down the cause"
  4. The request for condition B adds just 1 sentence to that: "write, in 1 line first, the condition that says this is where it ends, and then start investigating"
  5. Hand the same number of workers that do not carry the conversation over to each condition (6 each if you can)
  6. From the workers' records, count the longest run of reads with no change in between, using the script above
  7. Be sure to score the correctness of the answers too (= whether the answer has dropped by as much as it stopped)

Pass conditions (all of them have to hold)

  • The rate of correct answers is the same in both conditions (= having them write the condition for ending has not made the answer itself drop)
  • In condition B, a majority of the workers actually wrote the condition for ending (= the instruction is getting through)
  • The count agrees with 1 case counted by hand (= the detector is not letting them slip)

If it does not pass

  • The run does not shrink even in condition B: the condition for ending has become a tautology like "until the cause is known." Give examples of things you can write before reading (the number of files to read, the number of hypotheses to check)
  • The rate of correct answers dropped in condition B: either the material is too hard or the condition for ending comes too early. In that state, stopping and giving up turn into the same number
  • Workers that do not write the condition for ending come out: in the request, specify where to write it (the first 1 line, for instance)
  • The count does not agree with the hand tally: the judgment of writes is missing. Decide first whether writes to a temporary file are counted as "changed" too, or conversely whether it is narrowed to the artifact alone

Cleanup

  • Delete the generated record, the workers' records and the scoring table. Delete the scoring table along with them: leave it and, while you think you have rebuilt the material, you will be scoring against the old answers

Chase the cause to the end is right advice. Only, that advice does not have "until when" written in it. What is not written does not stop where the person who wrote it thought it would.


Next time, the art of not using the cheapest model. I stopped making the model with the cheapest unit price the default. Even so, the total for finishing 1 piece of work has come down.

End of CC BY 4.0