This is article 4 in the series "The Art of Not Chasing." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the introduction.

This time it is about leaving work that finishes in the main session to a separate worker. A worker here means a counterpart that starts up apart from the conversation you are in now, does only the range it was told to, and comes back with a report. Last time I handed the same task to both the expensive side and the cheap side and made a mechanism only out of the differences that failed. That was about "who to have do it," and this time it becomes about "whether to do it yourself, or hand it over." There is 1 line, and it is that you do not hand over what you already have — that is because the place you hand it to re-reads, once more, what you have already read.

I use this feature a lot. In the records on my machine, I had handed work to 342 workers. So this time is not "do not use it." When I counted the records of those 342 I had handed it to again, the line between work that was worth handing over and work that was cheaper not to hand over came out clearly. This is about that line.

Start by recalling one thing.

The last time you left a piece of work to someone, did the one you left it to have what you had read up to then?

If they did not, that premise is being read again from somewhere. The question is who is paying for it.

CC BY 4.0

I handed it to 342 workers, and was paying 16 moves of the main session for each one

I counted from the side of the record. I did not run any new experiment. That is because the records of 46 past sessions, and of the 342 workers started from them, were still sitting there as they were.

Let me settle the unit for counting first. What I count here is only "the amount newly read in." When you go on handing the same context over, most of it becomes reuse of what was handed over before. Add the part that could be reused as well and you end up counting the same text dozens of times and the digits change, so what I add up is only "the amount read in for the first time this round."

Amount newly read in (median)
1 move of the main session1,593
The 1st move of 1 worker I handed it to6,340 (4.0 times the main session)
The total for 1 worker I handed it to25,390 (16 moves of the main session)

Hand it to 1 worker and you have as much newly read in as going 16 moves forward in the main session.

And 1 worker I handed it to came back after 8 exchanges at the median. So work worth 8 exchanges has the price of 16 moves on it.

Line the number of exchanges up and compare, and the ratio comes out more plainly.

Amount newly read in, per 1 exchange
The main session1,593
1 worker I handed it to3,083

1.9 times. Even if it hits the tools the same number of times and thinks the same amount, handing it over makes it about 2 times.

There is only 1 place where the difference comes out

Why does it become 2 times? There is only 1 reason. The counterpart you hand it to does not have what the main session has piled up until then.

Take "the amount that could be reused" out of the same records and you see that place directly.

Amount that could be reused in 1 move (median)
The main session267,532
The 1st move of 1 worker I handed it to21,849

That is 1 in 12. The main session goes forward reusing 260,000 tokens' worth on every move, and what it was newly having read in was only 1,593. As a proportion, 99.1% is reuse. ⭐ The side handed the work, too, piles up its own context as the exchanges build, so it goes up to 95.3%.

Only on the 1st move, though, is there nothing piled up.

This explains both "a good deal" and "fast." They are not separate mechanisms; they are the same 1.

flowchart TD
  A["The premise the main session piled up"] --> B["Go on in the main session<br/>⭐ the premise is already loaded<br/>what you add is only the difference"]
  A -.->|"not carried over"| C["1 worker I handed it to<br/>⚠️ collects the premise again<br/>4 times on the 1st move"]
  B --> D["1 exchange 1,593"]
  C --> E["1 exchange 3,083"]

"There is no reuse of the premise at all" was going too far

This is a place where the explanation I had set up before I counted came out wrong, so I write it as it is.

I had thought that "the counterpart you hand it to is a new conversation, so none of the earlier reuse works at all." If that were right, the amount reused on the 1st move ought to be zero.

Workers whose 1st move had zero reuse: 74 of the 342 (22%)

Close to 80% were not zero. Split them up and the reason comes out. When several were started together, the 1 that moved first and the ones that followed after differed.

Amount that could be reused on the 1st move (median)Share with zero reuse
When only 1 was put out (90 workers)16,42129%
Put out together, the 1 that moved first (74 workers)3,18350%
Those that followed after it (178 workers)24,5426%

The 74 in the middle are a different thing from the "74 of the 342 with zero reuse" above (the numbers matching is a coincidence, and the breakdown is 26 + 37 + 11 = 74). Between the middle and the bottom there is a difference of 8 times. The order is taken not as "the order they were ordered in" but as "the order they actually started moving in." The ones put out together have no order of ordering, so line them up by the filename order of the records and this difference disappears (they become 24,538 and 24,540).

Workers share a "head" with each other. The same descriptive text goes to every worker first, so the ones that followed after can ride on the part the 1st one had read in. So the accurate way to put it goes like this:

⭐⭐ What you throw away by handing over is "the context the main session had piled up until then," not the shared head.

This correction does not change the direction of the price (paying 4 times on the 1st move has not changed). The ones that follow after do come out cheaper, though, so the continuation of this works on next time's story about running in parallel.

The speed side, too, had got slower for the same reason

The time comes out of the same records.

Seconds (median)
1 move of the main session (from the tool result coming back to it starting to speak next)6.7
Until 1 worker I handed it to first starts to speak11.4
Until 1 worker I handed it to finishes109.6

Up to the first word it is 1.7 times, and up to finishing it is 16 times. The number 16 times came out separately from the "16 moves" that came out on the price side, but it is only the same thing seen in a different unit. Increase the amount read in and the time to read it increases too.

Here I have made a mistake in measuring, 1 time. At first I took "finishing" as the time the result came back to the main session's side. Then a number came out saying "1.6 seconds until it returns." That is shorter than the first word (3.6 seconds) (that 3.6 seconds is the number that came out first at the time, and the 11.4 seconds in the table above is after I took it again). A worker run in the background returns only "received" right after it is ordered. I was counting that as the finish. I have taken the finish again, as the last utterance on the side I handed it to.

The condition in the other direction: not being piled up in the main session is itself a gain

Up to here it is only the "handing over is expensive" side. The condition in the other direction comes out of the same records too.

Median
The total amount 1 worker I handed it to dealt with in that time288,318
The amount added to the main session as a result3,244

That is 1 in 89. The counterpart you hand it to reads 280,000 tokens' worth and leaves only 3 thousand behind in the main session.

This is the decisive place. Do the same reading in the main session and what was read piles up in the main session.

And the main session's context is reused on every move after that. Reuse is cheap, but it is not zero. Read 30 files in the main session and, until that session ends, you go forward carrying those 30 files on every move.

So gain or loss is not decided by the weight of that work. It is decided by how many moves the main session goes on for afterward.

flowchart TD
  A{"Does the main session<br/>already have that premise"} -- "it has it" --> B["⚠️ do not hand it over<br/>hand it over and you pay 4 times on the 1st move<br/>to collect the same thing again"]
  A -- "it does not" --> C{"Will what was read<br/>be used afterward too"}
  C -- "it will" --> D["⚠️ read it in the main session<br/>hand it over and only the report is left"]
  C -- "it will not" --> E["⭐ hand it over<br/>reads 280,000 and leaves only 3 thousand behind"]

The principle: do not hand over what you have, hand over only what you do not have

Let me put the numbers so far into 1.

⭐⭐⭐ Work that the premise the main session already has is enough for: do not hand it over. Work that collects what the main session does not have: hand only that over.

This is not a story of "do not use it." It is a story of the place where you use it being settled.

In fact, I have not thrown this feature away. What I threw away is only "handing over, just in case, an investigation that finishes in 8 moves in the main session."

The line for the judgment can be drawn from the 2 tables above.

  • Work that has what the main session has already read be read again: hand it over and you pay 4 times on the 1st move and have it collect again what the main session has
  • A large volume of reading that the main session has not read yet: hand it over and it reads 280,000 worth and only 3 thousand worth of report comes back. The main session does not get fatter
  • Work whose contents the main session then refers to over and over: this you must not hand over. Only the report comes back, so the main session re-reads it every time it refers to it

How to verify: before handing it over, count whether the main session has that premise

Prerequisites

  • That you have an environment where you can start workers, and that its records (the amount read in, the amount reused, the times) are kept. If they are not kept, this procedure cannot be used. On a guess the sign changes (1 time I mistook the time of the finish and put the direction of the speed out backward)
  • That what you count is only "the amount newly read in." Add the amount that could be reused and you count the same context over and over and the digits change
  • That 1 record does not settle it. Look at the same kind of work over several runs together

Time required: 30 minutes (if the records are already on your machine. If you start from the setting that keeps the records, start by piling up 1 day's worth)

Steps

  1. From your most recent records, get the total "amount newly read in" for 1 worker you handed it to
  2. From the same records, get the "amount newly read in" per 1 move of the main session
  3. Divide 1 by 2. This is "how many moves of the main session 1 worker handed it to comes to"
  4. Count how many exchanges the counterpart you handed it to actually made. Compare 3 and 4: if the number of moves is larger than the number of exchanges, that much is what handing it over added on
  5. Get the total amount the counterpart you handed it to dealt with, and the amount added to the main session as a result. The larger the ratio, the more the work was worth handing over
  6. Line up the work where the ratio in 5 was small, and put "did the main session already have that premise" to each one

Pass conditions (all of them have to hold)

  • The number in step 3 has come out (if there were no records and it became a guess, this procedure has not held together)
  • In step 4, the difference between the number of exchanges and the number of moves can be explained (if it cannot, you have the units mixed up)
  • Among the work lined up in step 6, there is 1 or more that comes out as "the main session already had it." If there are 0, the way you hand things over is already right, so this article is not needed

If it does not pass

  • The amount for 1 worker is about the same as 1 move of the main session: you are adding the amount that could be reused as well. What you add is only the amount newly read in
  • The time of the finish comes before the start, or is shorter than the first word: you are counting the reply that accepts the order as the finish. Take it again from the last utterance on the side you handed it to
  • Counterparts that should have been started at the same time all come out as "the 1st worker": you are bundling them by the time they were ordered. Bundle them again by whether they were put out together in the same 1 utterance
  • The amount added to the main session is about the same as the amount the counterpart you handed it to dealt with: that report is too long. The point of handing it over lies in not bringing back what was read

Cleanup

  • Delete the intermediate files you used for the tallying. Do not delete the records themselves: delete them and the denominator for the next time you count again is lost
  • Keep the results you counted with a date on them. The records keep growing, so run the same script tomorrow and the numbers move

Caveat: what these numbers cannot say

This is not a controlled experiment. I did not measure the same work side by side as "done in the main session" and "handed over"; I only counted again, after the fact, the 342 cases I actually handed over.

So a bias in how they were selected is left in. The work I handed over is all work I "handed over because it looked heavy," and work that finished lightly in the main session is not in these 342 cases in the first place. This bias works in the direction of making "1 worker = 16 moves" look larger than it really is.

There are 2 numbers the bias does not reach.

  1. 1.9 times per 1 exchange: the number of exchanges is lined up, so it holds as a comparison even when the weight of the work is not lined up
  2. The amount that could be reused on the 1st move being 1 in 12: this has nothing to do with the contents of the work. That the counterpart you hand it to does not have the main session's context does not depend on what you asked for

The other way round, "16 moves" and "16 times" cannot be carried out as they stand. What you can use is your own number, from counting the same way on your own machine.

1 more thing. These numbers move with how you choose the model as well. Use the cheap side for the counterpart you hand it to and the price comes down even for the same amount read in.

The amount read in itself does not change, though: the work of collecting the premise again comes up whoever you have do it.

What I stopped handing over

  • I stopped having files the main session has already read be read again. Hand them over and you pay 4 times on the 1st move and have it collect again what you have on your side
  • I stopped "having it investigated in parallel just in case." A counterpart handed work just in case still reads in 1 worker's worth. It is a just-in-case worth 16 moves of the main session
  • I stopped handing over work whose contents get used afterward. What comes back is only the report, so the main session re-reads it every time it refers to it
  • There is work I have not stopped handing over: investigations the main session has not read yet, in large volume, whose contents do not have to be brought back. Reading 280,000 and leaving only 3 thousand behind is something the main session can never do

The judgment that leaving it to someone is faster looks right most of the time. That is because it saves you from stopping your own hands. Only, what the side it was left to does first was to re-read what I had already read.


Next time, the art of not using an agent team. I stopped letting workers talk to each other. That is because an interruption does not end in 1 go; it stays in both of their histories from then on.

End of CC BY 4.0