This is article 2 in the series "The Art of Not Chasing." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the introduction.

This time it is about the point where you decide which model to hand the work to. Last time it was "when to stop," but this time it becomes "who to hand it to." First, let me deny the title. Cheap models should be used. I use them too. What I stopped was making the cheapest one the "default." The price is decided by the ratio of unit prices × the ratio of amounts, and the ratio of amounts is the one you cannot know without measuring.

Let me note 1 more thing up front. In this article too, no model name, no price, no performance figure appears. I have decided that anything which goes stale in half a year does not get written in this series. What gets compared I call only "the expensive side" and "the cheap side." In exchange, every multiplier is there. Multiply by the unit prices on your own machine and they should be usable for a decision as they stand.

And a confession. In this investigation I misread the numbers as "there is no difference" 2 times. Both times the numbers were right and the reading was wrong. Both times, I noticed only after it was pointed out to me. Left as it was, I would have written "it was a groundless worry." This time I write those 2 times in as well.

Recall one thing.

The work you left to a cheap model: did it really finish cheap?

Can you say that having counted the number of redos as well?

CC BY 4.0

I started from a prediction: "the cheap side should have more misses"

The prediction I made first went like this.

Have a cheap model write tests and it can write tests, but the misses go up.

There is a reason I chose tests as the subject. That is because a test is the only artifact where a failure hides inside a pass. A bug in the implementation fails somewhere sooner or later, but a miss in a test says nothing. "I had it write the tests" is a pass and looks done, and the bill comes later.

For the yardstick I took 8 kinds of mutation that had gone into the code on my own machine in the past and that my own automated tests had actually failed on, and applied them again. You must not write mutations for this experiment. The moment you write them, the miss rate goes to 0% or to 100% depending on how you write them.

I put down something to compare against as well. Against the same target, the existing tests a human (me) wrote. Those fail on only 6 of the 8 kinds. The ceiling is not a perfect score.

There are only 2 conditions. The model (the expensive side / the cheap side), and the scope (narrowed to the 1 method call / just "write the tests you think are needed"). Not 1 point of view was handed over.

Run 1: half of the cheap side did not run

I lined up 12 workers that do not carry the conversation over and had them write, and before putting any mutation in, I ran them bare first.

Workers that ranDetection among the workers that ran only (how many of the 8 kinds it failed on. ⚠️ higher is better)
The expensive side6 / 63.33 / 8
⚠️ The cheap side2 / 63.00 / 8

The 4 workers that did not run were broken in ways like calling an assertion that does not exist, or duplicating a row in the dictionary. Every one of them shows up the moment you run it.

So I treated the 4 workers that did not run as "a separate bucket from misses" and left them out of the detection totals. There was a rationale of sorts: a test that does not run is something you notice on the spot, and you can switch over to the expensive one.

This is the first mistake.

Mistake 1: I was dropping the tests that did not run from the denominator

After leaving them out, only 2 workers were left on the cheap side. I was taking the average of those 2 as "what the cheap side can do" and lining it up against the 6 workers on the expensive side.

⚠️⚠️⚠️ Drop what did not run from the denominator and the most dangerous side disappears.

The fix was simple. I sent them back, the way it goes in reality. Return only the error output and send "please fix this." I do not say what is wrong (say it and that amounts to handing over the yardstick).

I sent them back over and over until all 12 workers passed. Counting again on that footing, it came out like this.

Scope narrowedNot narrowed
The expensive side3.33 / 83.33 / 8
The cheap side3.33 / 82.33 / 8

Narrow the scope and the cheap side ties with the expensive side. Here I was about to write this. "What goes up with a cheap model is the failures you can notice on the spot. The misses do not go up."

This too was a mistake.

Mistake 2: the yardstick had a ceiling

Of the 8 items, 3 items were 0 for all 12 workers. The existing tests a human wrote fail on those 3.

Lining the 3 up, they had something in common. Every one is an item you cannot write unless you first prepare either the material or the way of measuring yourself.

  • Put the longest headword in the table into the dictionary yourself
  • Give the same spelling 2 different readings
  • Count the number of queries, not the return value

The task said not 1 word about any of that. All it says is "please write unit tests."

So all 12 workers coming out 0 may not be because they could not write them, but because the task was not pointing there.

So I added just 1 paragraph to the task.

If the phenomenon is not in the dictionary you build in setup, that test is looking at nothing. Build the dictionary itself so that the behavior you want to see can come out. There are behaviors you cannot tell from looking at the return value alone. For those, prepare the way of measuring instead.

Not 1 of the 8 items is named. What it says is only the way of doing it.

And I handed the same amount to both conditions, down to 1 byte (increase the difference between the conditions and what you are measuring changes).

Once the ceiling came off, the conclusion flipped over

I had 12 workers write once more. All 3 items moved.

Run 1 (with the ceiling)⭐ Run 2 (no ceiling)
The expensive side3.33 / 85.00 / 8
The cheap side2.83 / 82.83 / 8
The expensive side, scope narrowed3.33 / 85.00 / 8
⚠️⚠️ The cheap side, scope narrowed3.33 / 8 (a tie)3.00 / 8 (not a tie)

Take the ceiling off and only the expensive side grew. The cheap side stays at 2.83 and does not move.

⚠️⚠️ "A tie if you narrow it" was just both of them taking only the easy items and lining up there.

Incidentally, the way the scope was narrowed did not work on the misses (the expensive side is 5.00 and 5.00, a difference of 0, and the cheap side 3.00 and 2.67). What the scope was working on in run 1 was "whether it runs in the first place," not the miss side.

The first prediction was right. Though before I could tell that it was right, I had to fix both of 2 things: how I was dropping from the denominator, and the ceiling on the yardstick.

A send-back makes it delete tests

The number of send-backs was 0 for the expensive side in both of the 2 experiments, and for the cheap side 6 in run 1 and 10 in run 2. Up to here it is as I imagined.

What I had not imagined was the way it fixed them. Send "the implementation is correct, please fix the test" and 3 of the fixes that came back had deleted the failing test.

  • Deleted the 1 case for "exactly at the upper limit passes"
  • Deleted the 1 case that looks at full-width normalization
  • Changed it to swallow the exception alone, instead of removing the duplicated dictionary row

Every one of them comes out a pass. And the 1 case that disappeared becomes a miss, just like that.

⚠️⚠️⚠️ Send it back saying "the implementation is correct" and it tips toward deleting rather than fixing.

This is not about the model, it is about the way of sending back.

That said, the only side that needed sending back was the cheap side. So as a result, it does work on the model side.

The principle: whether it is cheap is decided by "the ratio of unit prices × the ratio of amounts"

Let me fold the numbers so far into a shape you can use for a decision. Not the price of 1 response, but the total until 1 piece of work is finished.

On the same task, it is how many times the expensive side's amount the cheap side used.

What was countedScope narrowedNot narrowedOverall
Amount read for the first time1.64 times1.24 times1.48 times
⭐ Amount that could be reused9.71 times2.62 times5.29 times
Amount written0.26 times0.38 times0.31 times
Turns (send-backs included)3.33 times2.00 times2.67 times

The cheap side is not only cheap per unit. The amount itself is larger. Every redo makes the conversation longer, and the part it grew by flows through once more on the next turn.

The reused part is billed cheaply, but it is not zero. Treat this as 0 and a retry looks "nearly free." Count something that is 5.29 times as 0 and you can make any conclusion you like.

The amount written is the one place where the cheap side is less (0.31 times), so for work that is heavy on output the direction changes.

⭐⭐⭐ The price is "the ratio of unit prices × the ratio of amounts." The ratio of unit prices is written in the catalog, but the ratio of amounts you cannot know without measuring.

So "the default because it is cheapest" is a judgment that looked at only 1 of the 2. For the same reason, "the default because it is most expensive" is only 1 of the 2 as well. I put down no default, and count both of them for each use case.

In the earlier series I wrote that for the translation use case I go with the cheap side. I have not changed that judgment since. That is because for that use case, the cheap side took fewer moves as well. The results this time do not cancel that. Not canceling it is exactly the point: change the use case and the direction changes too.

How to verify: count the ratio of unit prices and the ratio of amounts separately

Prerequisites

  • That you can hand the same task to 2 models with the same material. If the material differs by even 1 byte, you measure the difference in the material, not the difference between the models
  • That you have a yardstick where pass or fail comes out of a program. If a human reads it and does the scoring, the number of times you measure becomes a cost and it does not last
  • That the tokens used are left split into the amount newly read, the amount reused, and the amount written. If they are not split, the second half of this procedure cannot be used

Time required: half a day (if you already have the yardstick)

Steps

  1. Take the mutations for the yardstick from ones that went into the code on your own machine in the past and that your own automated tests failed on. You must not write them for this experiment
  2. Put the existing tests a human wrote through the same yardstick. This is the ceiling. If it is a perfect score the yardstick is loose, so add more mutations
  3. Hand the same task over 2 models × 2 scopes × 3 times and have them write. Running is done on your side, and the writer does not get to run them
  4. Run all of them with no mutation put in. Leaving a failed worker "out of the totals" is prohibited. Return only the error output, send it back, and repeat until all of them come out a pass
  5. Against the ones that came out a pass, apply the mutations 1 kind at a time and run. Per 1 kind of mutation, you can run everyone as a batch in 1 go
  6. Total the number of send-backs and the amount newly read, the amount reused, and the amount written, for each model
  7. Multiply by the price list on your own machine and get the total up to completion

Pass conditions (all of them have to hold)

  • In step 2, there is 1 or more mutation that the human-written tests cannot fail on (a perfect score is a ceiling)
  • After step 4, every worker comes out a pass (count with even 1 worker still failing and the most dangerous side disappears from the denominator)
  • In step 5, the items where everyone is 0 come to less than half of the whole. If they are half or more, the task is not pointing there (write into the task that they are to prepare the material and the way of measuring themselves)
  • In step 6, the amount reused is not 0 (if it is 0, you are not taking it split out)

If it does not pass

  • The human tests get a perfect score: the yardstick is loose, or you are not taking mutations that your own automated tests failed on
  • A failed worker will not come out a pass: that is the decision point for "switch over to the expensive one." Record the number of times as a cost, then switch
  • The items where everyone is 0 are half or more: that is a ceiling. Add 1 paragraph to the bones of the task, hand the same amount to both conditions, and measure again
  • The amount reused is 0: the record is not split. If you cannot take it split out, leave the ratio of amounts aside and decide on the turns alone

Cleanup

  • Delete the tests you put down temporarily. Keep the record (the breakdown of the tokens)
  • Keep the results with a date on them. The records keep growing, so run the same script tomorrow and the numbers move

Caveat: what these numbers cannot say

1 cell of the table holds only 3 workers.

Read the direction alone.

The gap between 5.00 and 2.83 can be read as a ranking, but it is not precise enough to carry out as a multiplier.

What I compared is 2 points only. I picked 1 model on the expensive side and 1 on the cheap side, so a straight line saying "the cheaper it is, the lower it scores" cannot be drawn from here.

1 of the 8 items came out 0 for all 12 workers after the ceiling came off (⭐ it is the item that 1 of the 12 workers had failed on while the ceiling was there). The tests a human wrote cannot fail on it either.

So a ceiling is still left. "How far the gap goes" is not measured all the way by this yardstick.

The send-backs go out on the premise that "the implementation is correct." That premise itself is what made it delete tests.

In real work the implementation is sometimes wrong, so do not copy this way of sending back as it stands.

I do not let the writer run the tests. Had it been able to run them, the cheap side should have noticed its own failures.

Take these as numbers for "dumping the work without running it."

Last, the code used as the subject had quite thick comments. The implementation itself writes down even what happens when you break it. On code with thin comments, both sides should come out lower. These are not numbers you can carry out as they stand.

What I stopped making the default

  • I stopped handing work over on "the cheap one will do." That is a prediction. The ratio of unit prices is in the catalog, but the ratio of amounts does not come out without measuring
  • I stopped dropping results that did not run from the totals. The moment you leave them out, the most dangerous side disappears from the denominator. I send them back, get all of them to a pass, and count after that
  • I made a point of looking at whether the yardstick moves before writing "there is no difference." If half the items are 0 for everyone, that does not mean there is no difference, it only means nothing is being measured
  • I stopped sending things back saying only "the implementation is correct, fix the test." The road of deleting it to make it pass is open right there
  • There are things I have not stopped: choosing the cheap side for each use case. The use case I chose in the earlier series is still on the cheap side

The feeling that a cheap model is cheap mostly looks right. That is because, out of the bill, you are looking only at the unit price for 1 run. What was not visible was the number of redos, and the conversation that got read through again each time.


Next time, the art of not using the latest model. I stopped switching on the day a new model comes out. That is because work that would not go through without a strong model came to go through without changing the model.

CC BY 4.0 はここまで