This is article 6 in the series "The Art of Not Telling." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the introduction.

This time it is subagents. A subagent is an AI that starts up separately and works without carrying the main conversation over. The instructions that go across to it are the one thing in this series that I do write out in detail. I write them out not to teach it the content, but to fix the shape of the output.

This series has been about cutting down what I say to the AI.

But in this one article I am going to say the opposite.

Spell out your subagents' output in detail, and only that.

In article 4 of the previous series, "The Art of Not Listening to the AI's Opinions," I wrote Check your subagents' input, and only that. This time it is the exit side. In all 3 series the exception falls on the same party. The first series, which gets you out of reading the AI's output, made its exception "read your subagents' output, and only that" (article 5), and the previous series, which does not weigh up proposals, made it "check the input, and only that." And this time it is spell out the output.

Start by recalling just one thing.

The last time you ran a subagent, which model did it run on?

Was the way you wrote what you handed over something you had tried on that model?

CC BY 4.0

Just 1 procedure was left

In article 2 I wrote that the instruction file, the skills and the catalog of missions hold not one procedure that writes down how to go about a task. At the time I thought there were no procedures left on my machine at all. I counted again, and there was 1.

What is sitting there is the guide handed to a subagent. What I counted in article 2 was the instruction file, the skills and the catalog; I did not look as far as inside the files the catalog points at as "have these ready before launch." The 1 that was left was in there. This is where the exception is.

I threw the same task at 2 models, 2 ways each

In article 1 I compared an AI I had told to verify against one I had not. That was measured on opus. Subagents, though, run on a different model.

My catalog of missions says model: [sonnet, haiku] for the translation mission, with the comment "do not make opus the default" attached to it. So I cannot carry the results from article 1 straight over.

So I ran the same task 2 ways on sonnet as well.

ModelTold to verifyTool calls
Aopusyes8
Bopusno4
Csonnetyes13
Dsonnetno3

The answers from all 4 were an exact match with the answer key I had built beforehand with a script. All 4 avoided the trap as well, the one where only the 10th file does not have the sentence in question.

My hypothesis was wrong. I had thought that a lower model needs detailed instructions. sonnet produced 10/10 with no instruction, and made fewer moves than opus while doing it (3 calls against 4).

The way it was wrong splits in 2.

One: the growth differs. Told to verify, the moves doubled on opus (4 to 8), and on sonnet they went to 4.3 times (3 to 13). How much "asking makes the moves grow," the thing I wrote about in article 1, differs from model to model.

Two: traces of having checked against the source on its own were there only on opus. The opus I had not instructed put numbers at the top of its report: "grepped for the marker word. Exactly 1 hit in each of 9 files, 0 in the 10th." sonnet only stated that it "confirmed the phrase in all 9 of the others," with no numbers from any check.

All that can be observed is whether a trace was left in the report. Matching answers are not evidence that anything was checked against the source. As article 9 of "The Art of Not Reading" says, a self-report leans toward "done."

Anthropic's prompting guidance carries a proviso of its own. Treat a technique that comes with a particular model's name as something measured on that model, and check it for yourself before you apply it to another model.

What I understood fits in one sentence.

The techniques up to here hold only on the model they were measured on. For a party that runs on a different model, you write the shape of the output: what counts as done.

The mechanism: write the shape of done out in full on a single sheet

What you put down is a single sheet. For each mission you line up the conditions under which what came out is taken in. The one that takes it in is the main AI that started the subagents and handed out the work, and here I will call it "the center." It is the side that directs. The general term for it is the orchestrator.

# The guide to translating the series articles (a shortened version of the code on my machine)

| | |
|---|---|
| Input | 1 original article, and the translation of that language's "introduction to the series" (the authority on terms) |
| Output | `<series>/<lang>/<file>.md`, **1 file and no more** |

⚠️ Do not touch the list of registered articles (adding a language is the center's job. If everyone writes the same 1 file, they collide)
⚠️ Do not touch the original either. If you find the original needs work, stop translating and report it
⚠️ Do not run the tests (the concurrency guard lets only 1 through). The acceptance test is run by the center

## There are 3 kinds of fenced block, and whether you may translate them differs
  code      = what the reader runs on their own machine. Must match exactly (translate only the line comments)
  markdown  = prose the reader writes into their own instruction file. Translate it
  mermaid   = a diagram. The labels are prose, so translate them

## What the acceptance test drops (if it is not lined up, it fails)
  link URLs (order included) / the number and level of headings / the number of bullet items / the number of paragraphs /
  code fences (word for word) / the numbers in the body (thousands separators as they are)

## What the acceptance test does not drop (= a human looks at these)
  the meaning, the tone and the consistency of terms / whether the title reads at a glance in a listing / the real screen

What article 4 of the previous series wrote was the contract in the catalog: for which mission, and with what ready, you may launch. That is the entrance. This one is the exit, and it fixes the shape of what came out.

And this single sheet does not judge anything itself. What looks at "what gets dropped" is the acceptance test, and all it compares is whether the original and the translation have the same structure. The guide and the acceptance test both sit on the catalog's "have these ready before launch" side (they are a different mission's row in the field we looked at in article 2).

What grows here is the shape of the output, not the how. Even the 1 procedure found at the top of this article is, on the inside, the boundary between what the center handles and what the subagent handles.

And after this, I do not spell anything out

Once I had written the 1 guide, there was nothing left to write in the request each time. What I hand over is only "this article, into this language." The same shape as the exception article in the previous series: write it out in detail exactly once. That once is what made the requests that followed short.

Caveat: none of the 4 got it wrong

One, I cannot write that "a lower model gets it wrong." All 4 were 10/10. What I measured was 1 task, 1 run per condition, and that is not statistics. All I could observe was that "on an easy task, no difference showed up."

Two, you do not write a guide for every subagent. The ones I have written are only for missions where the same shape of output comes out many times over, and I have not written one for read-only investigation. The previous series draws the line in the same place: write thickly only for what gets called again and again.

Three, a guide too works only when it is read. This article is no exception to that. The only one that still drops it when the guide is not read is the acceptance test.

How to verify: leave out 1 thing the guide names, and confirm that it fails

Prerequisites

  • You have written the 1 guide to hand to a subagent, and it holds a list of "what the acceptance test drops"
  • You have finished writing the acceptance test (the thing that judges whether what came out is in a shape you can take in)
  • You do not start a subagent. You are checking only the receiving side

Time required: 10 minutes

Steps

  1. You get 1 correct piece of work ready and see it pass the acceptance test first (something a subagent produced in the past is fine). Skip this and, when step 3 fails, you will not know whether it failed because of what you left out or because it was failing all along
  2. You pick 1 item from the guide's "what the acceptance test drops" and make a copy with that element deliberately left out (delete 1 heading, take 1 line off a list). Leave the original file alone and break the copy
  3. You put the copy through the acceptance test

Pass conditions (all of them have to hold)

  • The acceptance test fails
  • What is missing comes out (something of the form "1 heading short")

If it does not pass

  • It passed anyway — the acceptance test is not looking at the "what the acceptance test drops" you wrote into the guide. The guide and the judgment disagree
  • All it says is "failed" — whoever fixes it (the subagent, or you) ends up hunting for the difference themselves

Cleanup

  • Delete the copy you broke

What it looks at is not meaning but sameness of structure. Whether the contents are right is for a human to look at. As far as a program can decide, let the program decide.


Next time, the art of not interrogating the AI. I do not interrogate the AI even when it gets something wrong. Even so, the mistakes get fixed.


Series: The Art of Not Telling

CC BY 4.0 はここまで