This is article 5 in the series "The Art of Not Chasing." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the introduction.

This time it is about splitting 1 piece of work across any number of workers and running them at the same time. Last time it was "do it yourself, or hand it to 1 worker." This time the place you hand it to becomes n workers. The "hand it over and it is about 2 times per exchange" that I gave last time I will not repeat. What I add this time is only where, and why, the price on the side you split it across grows. The price of splitting is decided by the number of workers × how long 1 worker's conversation runs. The number of workers is fixed the moment you decide to split, and the length of the conversation is decided by how you hand it over. And keep them waiting and it is cut off, so while they are stopped the time limit on the reuse runs out, and every restart ends up rebuilding the whole conversation.

First, let me deny the title. Where speed is needed, you should split. The only time you do not split is when you have no need to buy speed.

And the price is the kind of thing you may pay, if you are prepared to pay it.

On top of that, a confession. The shape where workers go on talking to each other as they work, I have never used in earnest. So not 1 measured number about that shape comes out here. What comes out is 2 things: what differs between handing it to 1 worker and splitting it across n, and where that difference turns into a price.

Recall one thing.

When you split it 4 ways, did the work finish 4 times faster?

And the price: had it gone to 4 times?

On my machine the speed went to 2.5 times at most, and the amount to nearly 4 times.

Let me say up front what there is to take home. The price is decided by "the number of workers × how long 1 worker's conversation runs." The number of workers is over at the point you decide to split, so the only side you can put a hand to is the length of the conversation. From here I count those 2 separately.

CC BY 4.0

Handing it to 1 / splitting it across n / having them talk — there are only 3 differences

First, let me sort out the words. These 3 are continuous with each other, in a relation where what grows gets added 1 at a time.

What growsWho waits
Hand it to 1 workerThe conversations grow by 1We wait on 1 worker
Split it across n workersThe conversations grow by n. ⭐ A join point appearsWe wait on the slowest 1 worker
Have them talk to each otherOn top of that, what each says enters the other's conversation⚠️ They wait on each other

What matters here is that every one of them is "a conversation grows." It is not that a price tag is attached to the feature. The price comes out as the amount by which the contents of the added conversation get read through again. From here that is what I count.

Split it across 4 workers and the amount you have it read goes to nearly 4 times

I counted the same records as last time again, this time from the batch side. I have not run any new experiment. A "batch" here means the cluster of workers started all at once in 1 and the same message. Across 46 sessions there were 164 batches / 342 workers.

Here it becomes a lie unless I give 2 ways of counting. Even in the same input there is a part read for the first time and a part that can be reused from the previous exchange, and the latter is billed cheaply. Only cheaply, not zero.

Workers put out at onceBatchesAmount read for the first timeAgainst 1 worker aloneAmount with the reuse added inAgainst 1 worker alone
1 worker9032,4661.00 times282,4861.00 times
2 workers1660,9981.88 times675,0632.39 times
3 workers1378,8702.43 times1,075,7373.81 times
4 workers4493,8602.89 times1,072,7713.80 times

Looking at the left column alone and reading it as "split it across 4 and 2.89 times covers it" is a mistake. The right column is 3.80 times — almost exactly the number of workers.

What has gone down is not the amount but the price attached to that amount. The input for as many workers as you split across does occur, properly, for that many workers. Where the part that comes out cheap is coming from I give in the next section.

There should be 1 place in the table that catches. 3 workers (3.81 times) and 4 workers (3.80 times) are almost the same. Even though the number of workers goes up by 1, the amount does not grow. This becomes clear in the section after next. That is because it is not the number of workers alone that decides the price.

There is 1 more thing this table alone does not settle. It may be only that the bigger the batch, the lighter the work I handed to each worker. That is because the one deciding how many to split across was me. The numbers that have nothing to do with the weight of the work are in the next section.

The only one warming things up was always the first worker

From the same records, this time I look at the amount that could be reused. This one has nothing to do with what was asked for, so it fills in the weak place in the table above.

Amount that could be reused on the 1st move (median)Share where the reuse was zero
When only 1 worker was put out (90 workers)16,42129%
Put out at once, the 1 worker that moved first (74 workers)3,18350%
The workers that followed after it (178 workers)24,5426%

A difference of 8 times. Every worker gets the same explanatory text handed to it first. The only one that has that text read in is the 1 worker that actually moved first, and the ones that followed were riding on it.

This is the reason why, in the table just now, only "the amount read for the first time" came in under the number of workers. It is not that the amount went down. The same amount simply moved over, from the 2nd worker on, to the cheaper way of counting.

flowchart TD
  A["split it across 4 workers at once"] --> B["the 1 worker that moved first<br/>⚠️ has the explanatory text read in<br/>amount reused 3,183<br/>half of them from zero"]
  A --> C["the 3 workers that followed<br/>⭐ can ride on that part<br/>amount reused 24,542<br/>zero for 6%"]
  B --> D["the amount read for the first time<br/>comes to 2.89 times"]
  C --> D
  D --> E["⚠️⚠️ but add the reuse in and it is 3.80 times<br/>what went down is not the amount but the unit price"]

There is a pitfall in how you take the order. Workers put out all at once have no order of dispatch (that is because it is 1 and the same message). Line them up by the filename order of the records and this difference disappears.

Do that and it comes to 24,538 and 24,540, and the conclusion that there is no difference comes out. What you line them up by is "the order they actually moved in."

The other half of the multiplication: how long 1 worker's conversation runs

This is the most important section this time. I looked at the side of the number of workers just now. This time the right-hand side: I line the same 342 workers up again by how many exchanges 1 worker's conversation had. What the table lines up is 339 workers (the 3 workers that ended in 1 exchange are left out, because a per-exchange figure cannot be produced for them).

Exchanges for 1 workerWorkersTotal history read through again (median)Per exchange (median)
2 to 3 exchanges1764,23821,577
4 to 6 exchanges118173,63832,740
7 to 12 exchanges110295,11035,550
⚠️⚠️ 13 exchanges or more941,538,16554,243

Between 2 to 3 exchanges and 13 or more, it is 24 times. The number of exchanges differs by only about 5 times.

The reason is in the right column. The weight per exchange grows along with it (21,577 → 54,243). When the conversation stretches, everything up to that point gets read through again every time, so the more exchanges there are, the heavier 1 exchange becomes. That is why the total works not as addition but as multiplication.

This is the reason 3 workers and 4 workers were almost the same in the table just now. Even if you add 1 to the number of workers, the amount does not grow if that 1 worker finishes short. The other way around, even without adding to the number of workers, it grows if 1 worker talks long.

The only side you can put a hand to is this one. The number of workers is over at the point you decide to split, but the shape you hand it over in, the one that settles how many exchanges it ends in, can be decided before you hand it over.

A median of 8 exchanges. That is the place doing the paying

The 342 workers on my machine came to a median of 8 exchanges per worker. ⚠️ The longest 1 worker is 107 exchanges.

Line them up by use case and it becomes clear.

Use caseWorkersExchanges (median)Total read through again (median)
Translation (1 worker = a whole language)5554,840,746
Investigation handed over without writing a use case64201,142,747
Experiments (the material made first, then handed over)1396206,526

Let me write the surprising side first. I got this one wrong. I had thought "work where I have prepared even the procedure and the acceptance test should be the cheapest." Translation is that work. The list to hand over, the script that slots the result in, the acceptance test that produces pass or fail: all of it is prepared on my side.

That came out the most expensive. There is 1 reason. That is because I left a whole language to 1 worker in one piece, and so it took 55 exchanges.

The cheapest, the other way around, were the workers put out for experiments. The material I make with a script on my side, and the side that receives it only does it once and returns. 6 exchanges.

It was not decided by how carefully I prepared. What decided it was the shape of the work handed over: how many exchanges it ends in.

Splitting it really does make it faster. Only, the hold-up is the slowest worker

Let me give the speed side as well. Leave it out and this becomes a story with only one side.

Workers put out at onceUntil the batch finishesThe same number done 1 at a time in orderHow much it shrank
2 workers129.9 seconds197.9 seconds1.52 times
3 workers99.1 seconds258.4 seconds2.61 times
4 workers146.6 seconds372.4 seconds2.54 times

Split it across 4 and it shrinks by only 2.54 times. It has got worse than when I split it across 3 (2.61 times).

The reason shows up in the spread inside the batch.

Workers put out at onceThe fastest workerThe middle workerThe slowest workerSlowest ÷ middle
2 workers82.9 seconds98.9 seconds129.9 seconds1.31 times
3 workers78.5 seconds85.4 seconds99.1 seconds1.16 times
4 workers78.5 seconds88.4 seconds146.6 seconds1.66 times

When I split it across 4, the middle worker finishes at 88.4 seconds. Even so, the batch finishes at 146.6 seconds. That is because what we are waiting on is not the middle worker but the slowest worker. The more workers you add, the more times you draw a losing ticket.

Keep them waiting and the part that can be reused disappears

Up to here I have been writing that "it is cheap because there is a part that can be reused." That part has a time limit.

I split the 68 times I sent a worker a "carry on" by how long that worker had been kept waiting before I sent it.

Time kept waitingTimes"Amount read for the first time" on the move that resumedSame, the amount that could be reused
Under 5 minutes6034639,058
⚠️⚠️ 5 minutes to 1 hour815,10424,879

44 times. While it is kept waiting the time limit on the reuse runs out, and the move that resumes rebuilds the whole conversation up to that point.

The denominator is only 8 times. Do not read the number itself; look only at the direction.

This is the place where it works hardest on the shape that has them talk to each other. Have workers converse with each other and, while one side waits for the other's reply, it stops, without fail. While it is stopped the time limit runs out, and every restart ends up rebuilding the whole conversation.

And as the previous section showed, the longer the conversation, the larger the amount to rebuild.

That said, this is an inference from the mechanism. What I have on hand is only the interruptions I sent to a worker from my side, and I have not once measured workers waiting on each other.

If you do not have them talk, it does not have to be a "team"

Turn everything so far over and the judgment becomes clear.

The reason to choose that shape is only when you want them to converse with each other. If you do not have them converse, what has grown is only the number of conversations, so it is the same as simply handing work to n workers at the same time. If it is the same thing, I choose the way of handing over where no waiting on each other happens.

What you want to doThe shape to choose
⭐ Clear the same kind of work away quickly, togetherJust hand it to n workers at the same time. ⚠️ There is no need to have them converse
⚠️ One side cannot be decided without seeing the other's result⭐⭐ Do it in 1, without splitting. Solve it by having them converse and the waiting on each other piles up on both
⚠️⚠️ You really do want them to consult each other⭐ It is a shape you pay for speed. ⚠️ What grows then is the exchanges (1 worker's conversation stretches), and the rebuilding at every wait, the 2 of them

The principle: the price is "the number of workers × how long 1 worker's conversation runs"

The numbers so far fold into 1 multiplication.

⭐⭐⭐ The price of splitting is the number of workers × how long 1 worker's conversation runs. The number of workers is fixed the moment you decide to split, and the length of the conversation is decided by how you hand it over. And the one deciding the time it ends is only the slowest 1 worker.

The 2 numbers this article produced are the left and the right of that formula. The side of the number of workers is, at the point you split, almost exactly the number of workers (3.80 times for 4). The side of the conversation's length moves by 24 times depending on how you hand it over. The only one you can move is the right.

flowchart TD
  A["split it across n workers"] --> B["the conversations grow to n"]
  B --> C{"how many exchanges<br/>1 of them ends in"}
  C -- "short (a few exchanges)" --> D["⭐ only the number grows<br/>64,238 at 2 to 3 exchanges"]
  C -- "long (more than ten exchanges)" --> E["⚠️⚠️ number × length is what works<br/>1,538,165 at 13 exchanges or more"]
  A --> F["a join point appears"]
  F --> G["⚠️ it ends when<br/>the slowest 1 worker ends"]

If speed is needed, it is a price you may pay.

That said, in a situation where speed is needed there is no time to check this carefully. So in practice you end up paying without knowing where the waste is. At the very least, know in advance where it grows: that is what I wrote this section for.

The shape for going parallel without raising the price

The moves I wrote in the earlier articles work as they are. This is not a new story. It is only applying the same things once more.

  1. Before handing it over, make the material on your side (the same shape as article 3's "hand the pattern over in explicit writing"). Leave it to them from the searching on, and all the exchanges spent searching pile up in that worker's conversation
  2. What you hand to 1 worker goes up to the size it can return in 1 go (the same line as last time's "hand over only the work of gathering what the main session does not have"). Not a whole language, but a single file
  3. Do the pass or fail judgment on your side. Leave the judgment to them as well and the redoing of that judgment piles up in that worker's conversation
  4. Even the weights out before splitting. Mix 1 heavy item in and you end up waiting even after the rest have finished

Do this and it leans toward the "workers put out for experiments" side of the table above (6 exchanges / 206,526). With the same number of workers, how you hand it over alone made a difference of 23 times.

How to verify: count the left and the right of the multiplication separately

Prerequisites

  • That you have an environment where several workers can be started at once, and that its records (the amount read in, the amount reused, the exchanges, the timestamps) are kept. If they are not kept, this procedure cannot be used
  • A "batch" is bundled by whether they were put out at once in 1 and the same message. You must not bundle them by the time of dispatch (what was put out in parallel scatters in the record, and all of it gets treated as 1 worker)
  • 1 batch does not settle it. Look at batches with the same number of workers, several of them together

Time required: 30 minutes (if the records are on your machine)

Steps

  1. From the records, split the batches up by number of workers
  2. For each number of workers, produce the median 2 ways — the amount read for the first time alone, and the one with the amount reused added in. Taking 1 worker as 1.00, produce how many times each of them is
  3. Compare the 2 multipliers from step 2 against the number of workers itself. Do not look at one alone and read it as "it came in under the number of workers"
  4. Inside a batch, split it into the 1 worker that moved first and the workers that followed, and produce the amount that could be reused on the 1st move. The order you line them up in is "the order they actually moved in" (with dispatch order or filename order the difference disappears)
  5. Forget the number of workers for a moment and sort the workers into classes by "exchanges for 1 worker" (2 to 3 / 4 to 6 / 7 to 12 / 13 or more). For each class, produce the total history read through again, and that total divided by the exchanges
  6. For each batch, produce the slowest worker's time divided by the middle worker's time
  7. For each number of workers, line up the time until the batch finishes against the total for doing the same number 1 at a time in order

Pass conditions (all of them have to hold)

  • In step 3, the 2 multipliers are apart (if they are the same, you are not counting the amount reused)
  • In step 4, the 1 worker that moved first has the smaller amount that could be reused (if they are the same, you have the order wrong)
  • In step 5, the more exchanges a class has, the larger its "per exchange" is as well (if they are the same, you are not counting the resending of the history)
  • There is 1 or more batch whose ratio in step 6 is larger than 1.0. If all of them are 1.0, the weights on the side you split across are perfectly even, so this article is not needed

If it does not pass

  • The 2 multipliers come out the same: you are not counting the amount reused separately. The amount read for the first time and the amount reused are in different columns of the record
  • No difference comes out between the first worker and the ones after: you are lining them up by dispatch order. Workers put out at once have no order of dispatch. Line them up again by the time the first reply came back
  • "Per exchange" does not change even as the exchanges grow: you are looking only at the amount newly read in, not at the amount that could be reused. Count the history read through again instead
  • The slowest worker does not match the end of the batch: you are taking the end from the time the result came back to the main session. Take it again from the worker's last message

Cleanup

  • Delete the intermediate files used for the tallying. Do not delete the records themselves
  • Keep the results with a date on them. The records keep growing, so run the same script tomorrow and the numbers move

Caveat: what these numbers cannot say

Let me write the biggest weakness first. On the shape that has workers talk to each other, I have measured nothing. That is because I have never used it in earnest. What I wrote in this article goes as far as "split it across n workers," and beyond that it is only generalities.

Nor is this a controlled experiment. I did not line the same work up as "done by 1 worker" and "split across 4 workers" and measure them; I only counted the 164 batches I actually split again, after the fact.

The one deciding how many to split across was me. So "it gets cheaper per worker" may be only that the more widely I split, the lighter the work I handed to each worker. A bias in this direction I could not remove.

There are 2 numbers the bias does not reach.

  1. The amount that could be reused differs by 8 times between the 1 worker that moved first and the workers that followed — this has nothing to do with what was asked for. Every one of them gets the same explanatory text handed to it
  2. The slowest worker takes 1.66 times as long as the middle worker — this is a ratio inside a batch, so it holds as a comparison even when the weights across batches are not even

The table of exchanges has 1 bias that cannot be removed as well. Whether the exchanges grew because the work was hard, or grew because I handed it over in a shape that makes exchanges grow, is not settled by these numbers alone. What I can say goes as far as "when exchanges grow, it works as multiplication," and not "change how you hand it over and it will always go down."

The other way around, "3.80 times" and "2.54 times" cannot be carried out as they stand. What you can use is your own number, from counting the same way on your own machine.

Let me fix 1 place in the previous article. In its first draft I wrote this way of splitting as "the 1st worker put out at the same time (164 workers)." I fixed the previous article's table to 3 rows on the day I wrote this one. Open article 4 now and what is there is the table after the fix. Those 164 workers had 90 workers from batches where only 1 worker was put out mixed into them. Take the mixture out and split it again and the difference widens, to 8 times. The conclusion of the previous article (the workers that follow can ride on the first one) does not change; the direction was toward getting stronger.

1 more thing. Batches split across 5 workers came to only 1 in the records. The numbers do come out, but with 1 case I do not read them. That is why there is no row for 5 workers in this article's tables.

What I stopped splitting

  • I stopped "splitting it 4 ways for now." At the point you split, the amount you have it read is fixed at almost exactly the number of workers (3.80 times). The speed is decided by the slowest 1 worker
  • I stopped handing 1 worker "1 whole thing." Leave a whole language to it and it takes 55 exchanges, and the reading through again went past 4.8 million. I cut the same work into sizes it can return in 1 go, and hand those over
  • I stopped splitting things whose weights are not even. Mix 1 heavy item in and you end up waiting even after the other 3 workers have finished
  • I stopped thinking "I prepared the procedure and the acceptance test, so it should be cheap." It is not decided by how carefully you prepare. What decides it is the number of exchanges
  • There is work I have not stopped splitting — investigation whose weights are even, where they do not need to see each other, and which can be returned in 1 go. There it is 3.80 times the amount and 2.54 times the speed. What balances it out is that the part that can be reused is cheap

The feeling that adding workers makes it faster mostly looks right. That is because 3 of them finish first, and the time spent waiting is not visible.

And the price was settled for that many workers at the point I decided to split, and from there on it was decided by the 1 worker that talked the longest.


Next time, the art of not having it reviewed midway. I stopped having another worker look my work over partway through. That is because more than 80% of the places I hand it to are a different model, and they do not know what I was already keeping.

CC BY 4.0 はここまで