This is article 3 in the series "The Art of Not Chasing." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the introduction.

This time it is about work that would not pass unless a strong model did it, and how it came to pass. First, let me deny the title. You should use a new model. Use it as soon as it comes out. The only case where you stop needing to is when you can put the pattern of that work into explicit writing, and on top of that you have checked that "it fails if you do not hand it over." On the day I started writing series 1, I was throwing roughly 80% of it at a model said to be strong at prose. Now the same work comes out with the default model alone. What I did in between was not to pick the model again. I only piled up conventions and automated tests. It would not pass unless a strong model did it because I was filling in what the mechanisms were short of with the model's power.

One more thing to say up front. Model names, prices, performance figures: none of them appear in this article. I have decided not to write, in this series, anything that goes stale in half a year. What I write is the shape of it: where the difference comes from, and what you put down to make it go away. From here on, what gets compared I call only "the expensive side" and "the cheap side."

Start by recalling one thing.

The last time you decided "this one is not going to work unless it is the stronger side," what did you compare with what?

Did you hand the same thing to both and line up the places where they failed?

Or did you switch over the moment one of them failed?

CC BY 4.0

The breakdown across 1 week did not amount to evidence

From the records on my machine, I counted the responses by model across the 1 week I was writing articles.

DayWhat I did that dayBreakdown by model
Day 117 drafts for series 1 (0 conventions / 0 automated tests)⭐ the side strong at prose 83% / the default 17%
Day 39 article bodies for series 3 + piled up conventions worth 76 commitsthe default 78% / the specialized side 22%
Day 62 article bodies + 3 automated tests for how it is written⭐ the default 100%

Line up only the same kind of work (writing the body of an article) and it is specialized 83%, default 78%, default 100%. What I piled up in between was 29 clauses of convention and 6 automated tests that look at how it is written.

Even so, this table does not amount to evidence. There are 3 reasons.

  1. There are days I left out. On Day 4 and Day 5 I was making an appendix for a book rather than articles, so the kind of work is different. Line them up in the same column and that day's breakdown looks like "the effect of what I piled up"
  2. A new version of the side strong at prose came out right around Day 1. It overlaps with the period when "more became available to use," so the movement in the breakdown can be read as either reason
  3. This is the heaviest one — what these numbers show is "what I used," not "what was enough." I did the switching by hand, so as it stands this is a record of my own subjective view

None of the 3 is something statistics can remove after the fact. The records have already been taken, and the conditions are ones I chose. So I prepared a controlled side separately.

I handed the same task to 12 workers, on the expensive side and the cheap side

I started 12 workers that do not carry the conversation over, all at once, and handed them the same task. The task is "write 1 unpublished article." There are 2 conditions.

  • The mechanism condition: hand over sentences that put the pattern of how to write into words, or do not hand them over
  • The model condition: the expensive side, or the cheap side

No human does the scoring. For the articles in this series, automated tests already look at the pattern of how they are written (the thesis sentence in the opening, the shape of the next-time preview, the 6 labels in How to verify, the length of the body). I diverted those automated tests as they are into the scoring table, so not once does my own judgment enter into pass or fail.

[observed] run 1 (12 workers, out of 11 points)
pattern not handed over  expensive side 8.0 / cheap side 6.3
pattern handed over      expensive side 11.0 / cheap side 9.0
mechanism difference +2.8 / model difference 1.8

The mechanism difference (+2.8) came out larger than the model difference (1.8). And here is the 1 line that works hardest —

⭐⭐⭐ The cheap side with the pattern handed over (9.0) came out above the expensive side with the pattern not handed over (8.0).

There is something this 1 run still does not let me say, though. The side handed the pattern has an input longer by exactly the sentences handed over. Whether what worked was "the pattern" or "the amount handed over" cannot be separated in this shape.

What worked was the pattern, not the amount handed over

So for run 2 I added 1 more condition. A condition that hands over sentences of the same length with not 1 word about the pattern in them. I raised only the amount, by as much as the pattern, and took the content out.

[observed] run 2 (12 workers, converted to the same 11 items as run 1)
pattern not handed over                 7.0
unrelated sentences of the same length  7.5
pattern handed over                     10.2

The unrelated sentences of the same length (7.5) landed next to handing nothing over (7.0). Only the side handed the pattern sits +2.7 away from there. What worked was not the amount but the pattern.

And the strongest thing in this run is not the averages themselves. It is that even though I ran it on a different subject, with different concrete examples and a different ordering, almost the same spread as run 1 came out (7.2 and 10.0 against 7.0 and 10.2). A difference over 1 run is noise, but if the spread reproduces, that is a shape.

Here too I leave 1 thing I have not settled. That the unrelated sentences did not work can be read as "the amount does not work" and equally as "it saw they were unrelated and skipped them." In this shape they cannot be separated.

Even with the pattern put into words, there were things that were not kept

When I lined up the items that failed, the 1st place was the same throughout. The length of the body. 7 of the 12 workers missed it, and 8 in the next run.

The range for the length is written in the conventions as a number. They miss it even so. The reason is clear: it is because the side doing the writing cannot count the characters it has written.

⭐⭐ There are patterns that are not kept even when put into words — the ones you cannot know without counting for yourself.

This draws 1 line through the choice between "write it in the conventions" and "make it an automated test." A pattern that is kept once read is served by a sentence. A pattern that is not settled without counting does not reach through a sentence. The latter can only go on the automated test side from the start.

This result does not deny "put the pattern into words and the cheap side passes too." It is a redrawing of the line: there are patterns for which "a sentence" is not enough as the form of putting it into words.

With no pattern handed over, the expensive side fell apart more

There is something that disappears when you look only at the averages. The lowest score among the 12 workers came from the expensive side.

That 1 worker scored 2 points out of 13 before conversion (the scoring in this run is 13 items, and the averages above are those converted to the same 11 items as run 1). It got the level of a heading wrong in 1 place, and put down not 1 verification section. Missing 1 place made the items hanging off it fail all together.

⭐⭐ When the pattern is not put into words, the way it fails is sometimes not "drops a little" but "misses the whole thing."

That 1 worker alone flipped the sign of the model difference for that run. That is why I cannot write "the cheap side is better." What I can write is the spread of how they fail, not the average. With no pattern handed over, neither side fails in the same way.

There is no point putting a mechanism on a pattern that is kept even when you do not hand it over

As the runs piled up, the number on the other side came out too. It was when I added a condition that hands over nothing of what has been learned.

[observed] run 5 (12 workers, out of 11 points)
hand nothing over 10.75 / hand over concrete examples 10.75 / hand over the pattern 11.00
of the 11 items, the ones these 12 workers never once failed: 10

Even with nothing handed over, both sides kept 10 of the 11 items. A difference came out on 1 item only.

⭐⭐ There is no point putting a mechanism on a pattern that is kept even when you do not hand it over.

This is not a counterexample to the theme, it is about the order of the steps. What to look at before putting a mechanism down was not "will it be kept once I put it down" but "does it fail if I do not." Add 1 clause of convention for an item that does not fail and all that grows is the conventions.

The cost: put a mechanism down and something gets blocked by exactly that much

This is a side effect that hits the theme head on, so I write it without hiding it. Hand the pattern over in explicit writing and points get lost from the handing over.

  1. Only on the side handed the explicit text did ways of writing that my automated tests cannot see increase. 0 cases in the 10 workers handed nothing, 5 cases in the 10 workers handed the pattern (both are worker counts summed across 2 experiments). The pattern is kept, and at the same time the eye of the automated test watching that pattern gets blocked
  2. Handing the explicit text over made the very item it was handed over for fail. In all 3 workers that failed, the line that failed was a comment and nothing else. It is the shape where they not only kept what had been learned and handed to them, but wrote down the reason they were keeping it, and that 1 line caught on the very automated test in question

Both are the cost side of "the difference gets filled by a mechanism." Mechanisms are not free. For every one you put down, a way of failing that comes from putting it down is added.

The principle: the model difference points to where the holes in the mechanisms are

Let me put the numbers so far into 1.

⭐⭐⭐ A difference you were filling with the model's ability gets filled when you put a mechanism down. So "I need a stronger model" was, most of the time, "the mechanisms are not enough."

What matters is that erasing the difference itself is not the goal. The places where a difference came out point to where the mechanisms are not enough. The expensive side fills the missing mechanism in with ability, so the hole stays invisible. The cheap side fails at that same hole as it is.

So handing it to both is itself a detector. The places where the expensive side passed and the cheap side failed — those alone are the differences a mechanism can fill.

The other way around, a difference whose reason for failing cannot be brought down into an automated test sometimes remains. That difference is not a hole in the mechanisms, it is a difference in ability. The decision is settled by "could it be made into an automated test."

flowchart TD
  A["hand the same task to both"] --> B{"was it only<br/>the cheap side that failed"}
  B -- "both passed" --> C["⚠️ no mechanism needed<br/>putting one down only adds a pattern that is kept anyway"]
  B -- "only the cheap side failed" --> D{"can the reason it failed<br/>be made into an automated test"}
  D -- "it can" --> E["⭐ 1 more mechanism<br/>the difference gets filled here"]
  D -- "it cannot" --> F["⚠️ this one is a difference in ability<br/>leave it to the expensive side"]

You pay 1 time. Handing it to both costs more on the spot. For the items where the difference is gone, you can run everything after that on the cheap side.

How to verify: hand it to both and turn only the differences that failed into mechanisms

Prerequisites

  • That you can start workers that do not carry the conversation over on the same task, all at once (run them 1 at a time in order and the congestion of the moment turns into a difference in conditions)
  • That the task is one where pass or fail is decided automatically. On a task a human scores, this procedure costs far more (you need as many human judgments as there are workers)
  • That the subject matter is unpublished. Have it rewrite something already published and where the right answer lies gets muddy

Time required: 30 minutes (if you already have the automated test that does the scoring. If not, start by writing 1 automated test)

Steps

  1. Prepare the same task. What you hand to both is identical down to 1 byte
  2. Hand the same thing to both the expensive side and the cheap side. A difference over 1 run is noise, so run the same task several times
  3. Line up only the places where the expensive side passed and the cheap side failed. This is where the mechanisms are short
  4. For each one you lined up, put "does it fail if you do not hand it over" to it first. Items both sides keep get thrown out here
  5. Put the remaining differences down as conventions or automated tests. A pattern that is not settled without counting goes on the automated test side, not the conventions
  6. Hand the same thing to both once more

Pass conditions (all of them have to hold)

  • There is 1 or more difference lined up in step 3 (0 of them and that task is not showing you the holes in the mechanisms)
  • What you put down in step 5 does not ask for a human judgment even 1 time (if it does, that is not a mechanism, it is an operating practice)
  • In step 6, the places where the cheap side fails have gone down
  • The expensive side's score has not gone down (= the mechanism you put down has not broken the side that was passing)

If it does not pass

  • The differences came to 0: the task is too easy. Make the subject matter harder (tighten the yardstick instead and next time everyone fails at the same 1 place, and again no difference can be measured)
  • There are differences, but they cannot be made into automated tests: that one is a difference in ability. Leaving that item alone to the expensive side is the right answer, and forcing it into the conventions adds 1 more clause that is not kept
  • The expensive side went down in step 6: the mechanism you put down is banning even the writing that was passing. Narrow the range of the clause down to the 1 actual case that failed
  • The way they fail is scattered from worker to worker: the number of runs is not enough. 1 worker's extreme result changes the sign of the average (it actually happened: the 1 worker with the lowest score flipped the model difference)

Cleanup

  • Delete what the workers wrote and the scoring table. Delete the scoring table along with it — leave it and you will be comparing against old scoring while believing you rebuilt the subject matter
  • Keep the conventions and automated tests you put down in step 5. The only thing you may delete is the traces of the measurement

Caveat: what this experiment cannot say

I cannot yet say "you can make the cheap side the default." The number of runs per condition is small, and the model difference was never more than large enough for 1 worker's extreme result to change its sign. What I can say goes as far as the mechanism difference coming out larger, and that spread reproducing on a different subject.

1 more thing. The difficulty was decided by the material, not by the yardstick. With the same yardstick and the same task, changing only the wording of the material sent the number of workers failing a given item back and forth across 4, 12 and 0. As long as you measure by pass or fail, the difficulty flips on 1 line of the material. That is why a measurement in this shape can be read not as an absolute score but only as a difference between conditions inside the same material.

And the breakdown table at the top does not amount to evidence, right to the end. What the controlled side can say goes as far as "hand the pattern over and the difference shrinks," and "why I did not switch" is not contained in it.

What I stopped switching

  • I stopped moving up to the model above when I saw a failure. First I hand the same thing to both. If there is no place where only one of them failed, switching will not fix it
  • I stopped adding "a pattern that is not kept" to the conventions. Before adding it, I look at whether it is kept even when it is not handed over. If it is kept, I do not add a clause
  • I stopped writing patterns that are not settled without counting as sentences. Lengths, counts and the numbers in an ordering go on the automated test side from the start
  • I stopped swapping out the setup on the day a new version comes out. I do use the new one. The point is that the material for whether to swap is not the announcement but the list of the places failing on my own machine

The judgment that a strong model is needed looks right most of the time. That is because it is a judgment made after looking at where things failed. But whether what failed there is on the side "a mechanism can fill" is not something you know until you hand the same thing to both.


Next time, the art of not using subagents. I stopped throwing work that finishes in the main session at a subagent. That is because I counted how much of the premise I end up re-reading on my side while the 1 worker I handed it to is coming back.

End of CC BY 4.0