---
title: "The Art of Not Chasing #7: The Art of Not Choosing a Model"
author: garplab
publisher: TypingTube
license: CC BY 4.0
license_url: https://creativecommons.org/licenses/by/4.0/
license_scope: 「CC BY 4.0」の印から始まる節（仕組み・検証手順・コード）。印の無い本文は著作権を留保
canonical: https://typing-tube.net/articles/en/owanai-07-no-model-choice
series: "追わない技術"
language: en
---


> This is article 7 in the series "The Art of Not Chasing." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the [introduction](https://typing-tube.net/articles/en/owanai-00-intro).

This time it is about the point where you settle which model to use. Article 2 was "do not make the cheapest one the default," and article 3 was "do not switch on the day a new one comes out." This time it is about stopping comparing in the first place. That is because mechanisms do not increase while you are comparing. What a comparison returns is "which one is above," and what a mechanism returns is "either one gets through." What comes back has a different lifespan.

First, let me deny the title. Comparing itself is not a bad thing. I compared too. 2 articles' worth of this series (article 2 and article 3) are made entirely of comparison. What I stopped is going on comparing in order to choose.

1 more thing. In this article too, no model name, no price, no performance figure appears. What gets compared I call only "the expensive side" and "the cheap side." The cost I do not put out in yen or in seconds either. What I put out is only "how many mechanisms it comes to." Yen and unit prices go stale in half a year, but a ratio does not go stale.

And a confession. In 1 day, I used 483 minutes comparing models. In the same time, I could have put down 20 mechanisms that fail.

What I actually put down was 5.

And the difference that came out of the comparison was 0.17.

Recall one thing.

While you were comparing models, how many mechanisms did you add?

Once you have finished comparing, how many months does that answer last?

## The difference that came out of comparing was 0.17

I handed the same task to both the expensive side and the cheap side. The condition is "what you hand over." For 1 experiment I lined up **12 workers that do not carry the conversation over**, and ran that 2 times, with different random numbers. **Run 1 had 3 variants** (hand nothing over / hand concrete examples over / hand the pattern over in explicit writing), and **run 2 had 2 variants** (hand nothing over / hand the pattern over in explicit writing). The differences below are taken from the 2 variants at the ends, lined up the same way.

| Condition | Difference (out of 8 points) |
|---|---|
| ⭐⭐ **What you hand over** (hand nothing over → the pattern in explicit writing) | **+2.00** and **+2.17** |
| ⚠️ **Which model it is** | **0.17** and **0.17** |

The mechanism difference came out 12 to 13 times the model difference (2.00 ÷ 0.17 and 2.17 ÷ 0.17). And it came out at the same spread both times. With different random numbers, the same number came back.

This continues from what I wrote in article 3 (a difference you were filling with the model's ability gets filled when you put a mechanism down). Article 3 was "you do not have to switch to the new one." This time it goes 1 step on from there, and becomes "you do not have to decide which of the 2 to use in the first place."

## What is more, that 0.17 swaps direction

Let me break the 0.17 of the total score down by condition.

| What was handed over | Which one is above |
|---|---|
| ⚠️ Nothing handed over | **The expensive side** is 1.00 above |
| ⚠️ The pattern handed over in explicit writing | **The cheap side** is 0.67 above |

The sign was reversed. Hand over none of what has been learned and the expensive side wins; hand the explicit text over and the cheap side wins. Look only at the total score and it looks like "there is no difference," but inside, they had swapped.

This is what "no ranking comes out" means. "The expensive side is above" and "the cheap side is above" can both be said from the data on my own machine. What can be said is only this: decide what you hand over, and which one is above is decided too.

Notice that the order is the wrong way round.

Choosing the model first and thinking about what to hand over afterward is the usual order.

What is actually working is what you hand over.

## Of the 11 items, 10 are kept by both even when you do not hand them over

In a different experiment, I split what I want kept into 11 items and counted. I added 1 condition that hands over nothing of what has been learned, and looked, item by item, at "how many workers failed."

Of the 11 items, 10 were kept by both models even with nothing handed over. A difference came out on 1 item only.

So the place to put a mechanism is that 1 item. Add explicit text for the remaining 10 items and all that happens is that what was being kept goes on being kept; nothing changes.

And 10 items' worth of explicit text means having all of it read every time you hand it over.

Read this part together with article 3.

The side effect that handing the explicit text over blocks the eye of the automated test watching that explicit text comes out of the same data too. Mechanisms are not free. That is why narrowing down where you put one to 1 item is worth it.

## The answer a comparison gives flips over just by changing how you read it

There is 1 more thing, a property of comparison itself. In the experiment in article 2, each time I scored the same output again, the conclusion changed, 3 times in all. The numbers stayed right all 3 times. What changed was only how I took the denominator, and 1 paragraph of the task.

The answer a comparison gives depends strongly on how you compare. Where I misread it, I will write next time.

It is a different experiment, so the scores cannot be compared as they stand (the task and the yardstick both differ). There is only 1 thing that can be said here. In both experiments, what you hand over and how you read it worked more strongly than the difference between the models.

## The principle: while you are comparing, not 1 mechanism gets added

What a comparison returns is "which one is above." What a mechanism returns is "either one gets through."

What comes back has a different lifespan.

| | What comes back | How long it can be used |
|---|---|---|
| ⚠️ Comparison | Which one is above | ⚠️⚠️ **Until the day one of the 2 changes** |
| ⭐⭐ Mechanism | Either one gets through | ⭐ **Until the task changes** |

Models go on being updated after you have finished choosing. The grounds on which you chose disappear on the day of the update.

The automated test you put down, on the other hand, fails at the same place whichever model comes.

And a comparison, the whole time you are comparing, adds not 1 mechanism. This is the cost. What you are paying is not the fee but the mechanisms you could have made in that time.

## So how many mechanisms could I have made with the time I spent comparing

I counted. I did not run a new experiment. All I did was split 1 day's worth of records (62 commits) into 5 by hand.

I do not let an automated test decide the split. I tried to cut it by the files touched and it would not split. A move that writes an article also raises the lower bound on the automated tests, so it touches the automated tests' files. So I write the classification by hand, and have the automated test watch only "whether not 1 of them has been left unclassified."

    [observed] classified all 62 commits of 2026-09-09 (0 unclassified)
    [observed] moves: compare 19 / mechanism 5 / write 8 / upkeep 12 / other 18
    [observed] elapsed (minutes. gaps over 60 minutes cut off, 5 times): compare 483 / mechanism 117 / write 166 / upkeep 70 / other 134
    [observed] minutes per move: comparison 25.4 / mechanism 23.4
    [observed] mechanisms that could have been put down for the effort spent comparing: by moves 19 / by elapsed 20.6
    [observed] mechanisms actually put down: 5

The answer is 20. With the 483 minutes I used comparing models, I could have put down 20 mechanisms that fail.

What I actually put down was 5.

This number comes with a cross-check. The time per move was 25.4 minutes for comparison and 23.4 minutes for mechanisms, almost the same. So counting by moves (19) and counting by time (20.6) come out at the same answer. Had these been far apart, I was going to throw the moves side away, since it would mean the size of 1 move differs.

Of the 19 moves, the experiments themselves are 10. The remaining 9 were the script that does the scoring, the gate that judges whether an experiment may be set up, and the design of the next condition. The prep work for comparing came to almost as much as the moves spent comparing.

## Caveat: what these numbers cannot say

The comparison side is a lower bound. The list goes stale with every experiment, so the touch-ups afterward (12 moves) are really a cost of comparing.

The same touch-ups come up after writing an article too, though, so they could not be split apart. Put them in and the 20 grows further.

In my case, the comparison itself became an artifact, an article. This series sells comparison, so the cost has been recovered in full. A comparison probably does not turn out that way.

Read it with that subtracted.

I count with 1 move = 1 finished piece of work. This is not a measurement, it is my way of counting. The measurement is the time (483 minutes / 23.4 minutes), and that is what produces the 20.6. The 2 came out close, so I put out both.

These are numbers for 1 day and 1 person. All 62 commits are mine, and what I compared is only 2.

Read it not as a multiplier but as an order of magnitude.

And let me draw the most important line. This is not a story that says do not measure. All I stopped was measuring "in order to choose a model." Whether a mechanism I have put down is working, I still measure every time. How to measure that, I will save for next time.

## What I stopped comparing

- I stopped deciding "which one suits it" first. It is decided after you decide what to hand over. The order was the wrong way round
- I stopped reading "there is no difference" from the total score. Break it down by condition and the sign was reversed. The total score had only flattened that out
- I stopped handing over 11 items' worth of explicit text. 10 of them were kept even when they were not handed over. What I put down is only the 1 item that failed
- I stopped thinking of a comparison as "free." What you are paying is not the fee but the mechanisms you could have put down in that time
- There are things I have not stopped: fixing on one of the 2 for each use case. I just do not compare again after fixing on it

The feeling that comparing tells you mostly looks right. That is because on the day you compared, it does tell you. What was not visible was that the answer lasts only until the next update, and until the next task.


## How to verify: before you compare, count how many mechanisms that comparison comes to

This is not a procedure that says stop comparing. It is a procedure for looking at the price tag before you start.

**Prerequisites**

- That you have 3 or more mechanisms put down recently (automated tests, gates, hooks) and that how many moves each one took is left in the record. If it is not left, this procedure cannot be used (count the next 3 and come back)
- That the breaks in the work are left in the record (I took mine from commits). With no breaks, neither the moves spent comparing nor the moves spent on mechanisms can be counted
- The size of 1 move needs to be the same for comparison and for mechanisms. The way out when it is not is written in "If it does not pass"

**Time required**: 20 minutes (if the records are on your machine)

**Steps**

1. Write in 1 line what you want to compare (example: "do the misses in the tests change between the expensive side and the cheap side")
2. Count the moves that comparison needs — the script that does the scoring / the prep work to line the conditions up / the number of runs / the reading. Do not forget the script and the prep work (in my own measurements, that was half of it)
3. Pick 3 mechanisms you put down recently (automated tests, gates, hooks) and count how many moves each one took
4. Divide 2 by 3. This is the price tag on this comparison ("it comes to N mechanisms")
5. Write out as many mechanisms as you can think of that you could put down now, and compare that count with N. If it is fewer than N you may compare, since it means you do not yet know where to put them
6. Once you have run the comparison, line the difference that came out up against the difference from changing "what you hand over." Put out only one of them and you cannot tell whether it is large or small

**Pass conditions**

- In step 2, the moves for the script and the prep work are not 0 (at 0 you have missed some in the count. In my own measurements, that was 9 of the 19 moves)
- In step 3, the per-move figure for the 3 stays within 2 times the per-move figure for the comparison (if they are far apart, converting by moves cannot be used. Count again by time)
- In step 5, the mechanisms you wrote out are fewer than N (by that much, the comparison wins)
- In step 6, a difference from "what you hand over" has come out (if it has not, the yardstick is not moving. The difference from the comparison cannot be read either, by that much)

**If it does not pass**

- N is less than 1: it is a cheap comparison. Run it
- You could write out N mechanisms' worth in full: put those down first, then look at whether you still want to compare. In my case, once I had put them down I no longer wanted to
- The per-move figures are 2 times apart or more: stop converting by moves, and produce it by time alone
- No difference comes out from "what you hand over": the yardstick is stuck at the ceiling or the floor. Fix that before you read the difference between the models

**Cleanup**

- Keep the classification you wrote by hand. Have the automated test watch only "whether anything has been left unclassified" (cross a day boundary and things drop out silently)
- Cut off the gaps when you were away from the desk. Leave them uncut and the hours you were asleep turn into a cost of comparing (I cut at 60 minutes. It hit 5 times)

## To anyone who skipped ahead to here

You did the right thing.

There is no need to go to the trouble of measuring what you can tell just by reading. What that procedure says comes down to this 1 line.

> **Divide the moves that comparison needs by the moves for 1 mechanism.**

It is there for when you think later, "is that really so." If you are not thinking it now, go straight on to the next one.

What decides whether to check was not how to check, but whether there is something you do not know right now.

---

Next time, the art of not letting it give up. **I stopped folding it up with "no difference came out."** Up to this article I have been cutting with numbers, but **that is because when those numbers say "there is no difference," most of the time it was another way of saying "I have not been able to measure it yet."**
