---
title: "The Art of Not Reading #5: Read Your Subagents' Output, and Only That"
author: garplab
publisher: TypingTube
license: CC BY 4.0
license_url: https://creativecommons.org/licenses/by/4.0/
license_scope: 「CC BY 4.0」の印から始まる節（仕組み・検証手順・コード）。印の無い本文は著作権を留保
canonical: https://typing-tube.net/articles/en/62f714d7c2b599
series: "バイブコーディングにおける読まない技術"
language: en
---


> This is article 5 in the series "The Art of Not Reading." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the [introduction](https://typing-tube.net/articles/en/34b02627c718fa).

This time it is subagents. A subagent is an AI that starts up separately and works without carrying the main conversation over. It is a different worker, so nothing reaches it apart from the prompt you hand over: not the policy you settled on a moment ago, not the memory you have accumulated. Throw a job over there, and what is happening stays out of sight until it is finished.

This series has been about not reading what the AI outputs.

But in this one article I am going to say the opposite.

The subagent logs are the one thing you should read, once.

The main AI that starts the subagents, hands out the work and takes in the results is what I will call "the center" from here on. The general term for it is the orchestrator. Not looking before the guardrails are up is not the art of not reading; it is just neglect.

Nobody is watching a worker who started today, one for whom every "they will know that without my saying it" is gone.

What I understood fits in one sentence.

> **A subagent is a different worker, not a continuation of the main conversation. So you make it read how it is used before you put it to work.**

## Read, just once

Pick one task you recently left to a subagent, and start from what the center instructed and what it received.

Open the call on the main screen and you can read the prompt that was handed over in full, along with the report that came back (on Claude Code, in the detailed view of the conversation, or in the session log file).

Is the text that went over what you meant?

What does the report say it did, and how far did it get? While you are there, look at the numbers as well: how many calls, which model, how many tokens spent.

The point is not to keep watch. It is to learn which accidents are happening and decide where the guardrails go. That is why once is enough. When I read mine, I saw four things.

- Repeated failure. What was displayed: "Done." What had actually happened: the report came at the end of more than ten retries against the same error, and the result was unusable
- Accumulation ignored. I assumed it would work under the usual rules. In fact not one line of the memory or the startup rules had gone over, and the subagent walked back into a loophole I had spent weeks closing
- The wrong model. I meant it as light, routine work, but leaving the model unspecified meant it took over the center's expensive model
- **Heavy token consumption**. One day I handed the job of filling in the missing translations for [typingtube](https://typingtube.net) (a web service for practicing typing along with music videos on YouTube, offered in more than ten languages) out to one agent per language: **13 of them started in parallel, and the reported figures added up to about 1.73 million tokens**. Each agent **rewrote its own insertion script and its own verification script**, translated without being handed the per-language notes, and kept repairing the mistranslations that verification found and sending them back around. Work that should have been over in one pass swelled up in the back and forth between finding and fixing. I noticed only when I hit the usage limit

Every one of those lines looked, on the main screen, like something that was "making progress," right up until I read the log.

## Why it is not a continuation of the main conversation

The [official documentation](https://code.claude.com/docs/en/sub-agents) lists what does not load. The center's conversation history, the skills the center invoked, the files the center read, and the memory that accumulates on its own. What goes over is the prompt the center writes, and the instruction file. The "accumulation ignored" I walked into was not something I forgot to hand over. It works the way it is specified to. If the files the center read do not go over, that also means the subagent reads the same things again for itself. The reason 13 of them rewrote their insertion and verification scripts separately is that not one line of what the center had written had reached them.

The model is on the same page. If you do not name one at launch, the subagent runs on the same model as the center. My "light, routine work" ran on the center's model because I did not name one.

What comes back is the last message and the numbers for what it consumed, and nothing else. Counting the 102 records left on my machine, one record was 220,000 bytes at the median and the report that came back was 2,570 characters at the median. What the center sees is 1-2% of what the subagent read and wrote. Reading once means looking at the remaining 98% once.

The 1.73 million tokens above is those consumption numbers added up. It is the sum of the figures reported by the 10 that finished, and a reported figure is the size of that one agent's last round trip. It is not the total from start to finish. Adding up the whole records of the 19 that ran that day comes to more than 200 million tokens. Most of that is the part that gets read again on every call.

More than ten retries do not show up in the report. The way the report leans to "Done." has an official phrasing, too. The Claude Code documentation writes that [the AI stops where the work looks done](https://code.claude.com/docs/en/best-practices). The same page also says to make it show evidence rather than claim success, and to keep the side that did the work from becoming the side that grades it. Those two are why the acceptance test does not sit on the subagent's side.

There is an explanation for the more than ten repeats against the same error, too. Unless new information comes in from outside, [an AI cannot correct its own answer, and can come out worse after trying](https://arxiv.org/abs/2310.01798). The Claude Code documentation, for its part, says that if two corrections on the same problem do not fix it, the conversation is contaminated by the approach that failed, so throw it away.

## One mechanism: just apply the one from article 4

You take article 4's "hook that checks for the existence of a file that is not created unless you look at the documentation" and apply it at the entrance where subagents are started. You write one sheet, a catalog of missions: the jobs it is all right to throw at a subagent, and the list of what has to be ready before one starts.

```yaml
# subagent_missions.yml — the catalog of missions a subagent may be started for (a shortened version of the code on my machine)
missions:
  i18n_translation:
    description: fill in the translations (1 agent = 1 language)
    requires:                      # if these are not there at launch, the hook stops it
      - scripts/i18n_ledger.py     # the list of work (the only input handed over)
      - scripts/i18n_apply.py      # the center inserts the results
      - scripts/i18n_verify.py     # the center runs the acceptance test as well
      - docs/i18n_notes.md         # the per-language notes (the context handed over)
    model: [sonnet, haiku]         # unspecified or unpermitted models do not get through
    max_parallel: 5
    max_total: 20
```

What the launch hook looks at is whether the declared mission is in the catalog, whether the `requires` are really there, whether the `model` is one of the permitted ones, whether the count is within the limit: that and nothing else. If one of them is missing, the launch itself stops and the catalog is pointed at. An AI that is stopped reads the guidance and picks its way forward again.

The hook is not judging whether anything was read. What sits in `requires` are files that only someone who has already broken the job down into simple work can produce, and their existence is the mechanical test for whether that breaking down has happened. Each accident on the list above gets an answer of its own.

- Repeated failure is caught by the center's acceptance test, which drops the false "done"
- Accumulation ignored disappears with the context that the documentation in `requires` carries over
- The wrong model is stopped because naming one is mandatory
- Heavy token consumption is taken out by the three above together. The tools are the center's, what has been learned travels in the documentation under `requires`, and mistakes are picked up in one pass by the center's acceptance test, so no place is left for the back and forth to happen

The 5 and the 20 written in the catalog are separate from Claude Code's own limits. What Claude Code decides is that at most 20 can run at once and that nesting can go three deep, and nothing else; how many may be started in one session is not something it decides. The numbers in the catalog are ones I decided myself, inside those. Every extra worker is one more set of preliminaries to hand over.

One more thing. When it stops, it goes back to the center and from there back to a human. A subagent cannot ask a person. The tool for asking questions is taken away from every one of them. So there is no way to write a form where it stops and asks by itself; it can only go through the center. Letting it try again only rebuilds the token problem, that pile of retries.

## And after this, I do not read them

Once the guardrails are up, I do not read the logs. What goes over, which model, how many: all of it is settled at launch, and the only results I take in are the ones that have passed the center's acceptance test. What that acceptance test counts is whether the variable markers line up, whether the tags are unbroken, whether empty values are left, whether the item counts match: structure only, not whether the translation is right. I read once, and that was that.

## How to verify: take one away, and confirm that it stops

**Prerequisites**

- You have written the one catalog of missions and finished registering the hook that cuts in at launch
- The catalog holds at least one mission, and that mission's `requires` (the list of files to have ready before launch) are really there

**Time required**: 10 minutes

**Steps**

1. You move one of the files listed in that mission's `requires` out of the way. You move it aside, you do not delete it. You put it back in the cleanup
2. You have the center (the main AI) start a subagent under that mission

**Pass conditions** (all of them have to hold)

- It stops before the subagent starts. Failing after it has started is too late. The tokens are already spent
- The reason it stopped names the file that is missing
- The location of the catalog is shown

**If it does not pass**

- It started anyway — either the hook is not looking at whether the `requires` are there, or the mission name is not tied to a line in the catalog
- It stopped, but the name of the missing file does not appear — the center's AI stops without knowing what to prepare. Put the file name in the message the check prints

**Cleanup**

- Put the file you moved aside back, and have it started under the same mission once more. You are done when it runs without stopping

---

Next time, "The art of not reading rules." **I stopped writing the things I want honored as rules.** And those rules are still not broken.

---

**Series: The Art of Not Reading**

- ← Previous: [4. The art of not using skills](https://typing-tube.net/articles/en/8f9ae8791b448d)
- → Next: [6. The art of not reading rules](https://typing-tube.net/articles/en/569d598941c00f)
- All articles: [Introduction: I Barely Read What the AI Outputs Anymore](https://typing-tube.net/articles/en/34b02627c718fa)
