This is article 4 in the series "The Art of Not Listening to the AI's Opinions." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the introduction.

This time it is subagents. A subagent is an AI that starts up separately and works without carrying the main conversation over. One day I handed the translations for more than ten languages to 13 subagents, one agent per language. The job I thought I had handed over was carried out by the 13 of them in 13 different ways. All that went across was the wording of the call, and the contract (what gets handed over, where not to use it, the boundary not to cross) was never written down.

This series has been about not weighing up the AI's proposals.

But in this one article I am going to say the opposite.

Check just one thing, once: the input going across to your subagents.

Article 5 of the previous series, "The Art of Not Reading," wrote this accident up from the side that reads the logs that came out. This time we look at it from the side that hands the work over.

Start by opening just one thing.

The full text of the prompt you handed over for the job you most recently threw at a subagent.

Is the name of the script it is meant to use written in there?

Is it written in there who decides whether the result passes?

Probably not. I had not written it either.

CC BY 4.0

What the previous series only half said

In article 4 of the previous series I wrote this. Tool descriptions eat into each other the more similar condition lines stand side by side, and the one you most want used is the one that gets buried. So I stopped lining up conditions.

As a description of what happens, that is right. In my environment there are 84 scripts, and skill registrations (registering custom commands and tools) come to one.

But I wrote it in a way that can be read as "the shorter the description, the better." That is wrong. Anthropic's guidance on defining tools says the opposite outright. Write an extremely detailed description. It is the single most important factor in how well a tool performs. What it asks for is concrete too, close to a manual page. What the tool does, when to use it and when not to, what each parameter means and how it changes behavior, caveats and limitations, and the information it does not return. At least 3 to 4 sentences, it says, and more than that for a complex tool.

It looks like a contradiction, and it is not. The line can be drawn here. The text that decides whether a tool gets called (the heading that says when to use it) and the text that decides how it behaves (the body of the contract) are two different things. What I said not to grow in the previous series was the first one. The second should have been grown.

13 agents each writing its own scripts

On the production project on my own machine (typingtube, a web service for practicing typing along with music videos on YouTube), I handed out the translations for more than ten languages at one agent per language, and the 13 of them used about 1.73 million tokens.

Broken down, each subagent had rewritten its own insertion script and its own verification script. 12 of them in the scratchpad, and 68 to 97 tool calls per agent.

This did not happen because the "conditions for use" were vague. It happened because nothing at all had been written of the contract: what gets handed over, which scripts to use, and where the agent's own job ends.

What I understood fits in one sentence.

What you must not grow is the list of "when to use it," and what you should write thickly is the contract: what gets handed over, where not to use it, the boundary not to cross.

Why what went unwritten took a different shape in each of the 13

All that I handed to the 13 was the wording of the call. On my side sits everything we had exchanged up to that point: how much of each language was left, what I had tripped over before, where in which file the text goes. Not one byte of it went across to the 13.

Anthropic's documentation on subagents lists what does go across. A subagent starts in a fresh, isolated context. Your conversation history, the skills you invoked, the files you read: none of that is visible there. Six things go across: the subagent's own system prompt, the delegation prompt you wrote, the project instruction file, the state of git, the skills you named in advance to be preloaded, and the roster of the other subagents running in the same session.

On the day of the accident my project instruction file had no section on subagents. There were no skills I had named in advance to be preloaded either. What reached the 13 was the wording of the call, one thing.

What went unwritten does not stay blank. The subagent's side fills it in. With last article's prohibition, a single line with no scope written on it was all that arrived, and the scope of the thing counting locally became the scope of the prohibition. This time there was nothing on hand at all. How to insert a translation comes from what each of the 13 carries over from its training. The 12 scripts, alike and each a little different, were born that way.

The subagent documentation also tells you where to write what. A subagent's description eats context, so keep it short, and move the detail into each subagent's system prompt, which loads only when that subagent runs. The text that decides whether it gets called and the text that decides how it behaves are given separate places to live.

The mechanism: write the contract into the catalog of missions

What you put down is a single sheet. On the list of jobs it is all right to throw at the AI, you line up what has to be ready before launch. The main AI that starts the subagents, hands out the work and takes in the results is what I will call "the center" here. It is the side that directs. The general term for it is the orchestrator.

# The catalog of missions a subagent may be started for (a shortened version of the code on my machine)
#
# ⚠️ Do not start one for a mission that is not here. What belongs here is bulk work that breaks down into simple steps.
#   Throw a job that has not been broken down and each agent rewrites its own tools and melts the budget.
# ⚠️ requires is what has to be ready before launch. If even one is missing, the hook stops the launch.
#   That is the mechanical test for whether the job was broken down into simple work.
# ⚠️⚠️ Do not let subagents run the tests. The concurrency guard lets only one through, so
#   it fights the main side and one of the two always fails. Tests are run in the main session.

session_total_checkpoint: 25   # past this line, stop whatever the mission is and report to a human

missions:
  i18n_translation:
    description: Backfill locale translations (1 agent = 1 language)
    requires:
      - scripts/i18n_ledger.py                     # the list of work (the only input handed over)
      - scripts/i18n_apply.py                      # the center does the inserting
      - scripts/i18n_verify.py                     # the acceptance test is the center's too
      - docs/reference/i18n_translation_notes.md   # the per-language notes
    model: [sonnet, haiku]     # ⚠️ do not make opus the default (the translated text is under a tenth of the whole)
    max_parallel: 5
    max_total: 20

The description stays one line, and the comments ended up longer than the body. That is the "when not to use it," and it is exactly what the manual page above asks for. What article 2 attached to a prohibition was a reason. A tool description needs the same thing.

The script that cuts in just before launch (a hook, on Claude Code) is looking at this catalog. There are only 5 checks, and they split into the ones you can get around and the ones you cannot.

#CheckCan it be got around?
1A mission name is declaredWrite a name and it passes
2Is the mission in the catalogSame as above
3The requires files are really there⚠️ No
4The model is on that mission's permitted list⚠️ No
5The count of launch records (parallel, total, restarts under the same name)Start it under another name and it slips past

1 and 2 are the text that decides whether it gets called, 3 and 4 are the contract, and what this table shows is that the contract is the part doing the work. What sits in requires are files that only someone who has already broken the job down into simple work can produce, so their existence is the test for whether that breaking down has happened. Whether it was read cannot be measured. Whether it was written can.

One thing I got wrong. At the time I thought most of those 1.73 million tokens had gone on rewriting scripts, and that writing the contract and pulling the inserting and the acceptance testing into scripts at the center would cut that much out.

After pulling them in I measured 4 languages. The earlier way, each agent writing its own scripts, came to about 172,900 tokens per agent averaged over 10 languages. The later way, handing over the center's scripts, came to 190,325 tokens for the language with the most keys. The tokens did not go down. What was setting them was not duplicated scripts but the translation itself. Either way, reading the existing translations to line the terms up and writing the sentences that were not there yet cost the same.

What the contract bought was not speed. It was that the operations that break the YAML left the subagents' hands. If that YAML breaks the production server will not start, so that was the only route by which it could be broken. And it was that the acceptance test became the same one for every language.

And after this, I do not check

Check once, write it into the catalog, and that is the end of it. From then on the hook checks every time, in my place. At each launch it looks at whether the requires are all there, and stops if any are missing. I do not open the input.

There is one thing I stopped saying as well: the explanation that "this job is not suited to a subagent." I used to stop the AI every time it tried to start one and talk through why it did not fit. Now it cannot start what is not in the catalog, and the reason a job is not there is written at the top of the catalog. I never give the same explanation twice.

I checked exactly once.

Caveat: you do not write a contract for every tool

A thick contract is only needed at the entrances where a mistake is expensive: in my environment that is running the tests, committing, and starting a subagent, a handful of places.

The rest, close to 80 scripts, stay plain, and a mistake there shows up at once and is cheap.

Get this wrong and you end up writing 84 manual pages, doing to yourself the very "it dilutes as you add" the previous series warned about. Choose what to give a thick contract by the damage done when the contract is broken.

How to verify: take one requires away, and confirm that it stops

Prerequisites

  • You have written the one catalog of missions and finished registering the hook that cuts in just before launch
  • The catalog holds at least one mission, and that mission's requires files are really there

Time required: 10 minutes

Steps

  1. You move one of the files listed in that mission's requires out of the way. You move it aside, you do not delete it. You put it back in the cleanup
  2. You have the center (the AI that directs the subagents) start one under that mission

Pass conditions (all of them have to hold)

  • It stops before the subagent starts. Failing after it has started is too late. The tokens for the job you handed over are already spent
  • The reason it stopped names the file that is missing
  • The location of the catalog is shown

If it does not pass

  • It started anyway — the hook is looking only at the mission name, not at whether the requires are really there
  • It stopped, but the catalog is not shown — the center's AI stops without knowing what to prepare next. Add the one line that points at it

Cleanup

  • Put the file you moved aside back, and have it started under the same mission once more. You are done when it runs without stopping

A mechanism that stops things breaks in the direction of not stopping. A check that has never once stopped anything looks exactly like a check that is working. Both of them let the launch through in silence.


Next time, the art of not writing automated tests. For more than half of my prohibitions, I have not written an automated test. Even so, not one of them is a prohibition that is merely written down.


Series: The Art of Not Listening to the AI's Opinions

End of CC BY 4.0