This is article 3 in the series "The Art of Not Running." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the introduction.

This time it is about how much of what comes back is the same when you throw the same question at an AI over and over. The same question means that the files you hand over and the way you ask do not differ by 1 byte. One day I handed a list of 40 findings, made by a script, to 12 workers, and had them pick just 3 of them as the ones to fix first. The 12 answers that came back did not share a single sentence. That is not the part that causes trouble. What we normally look at is 1 of those 12. But if you gather them up and take a majority vote, the majority vote erases kinds: only the answers that came back most often are left.

In article 8 of the second series, "The Art of Not Listening to the AI's Opinions," I wrote that the number of proposals is unlimited but their kinds are finite and converge. What I based that on was that my record of turning down the AI's proposals had stopped at 17.

That was a number I counted once, in my own project, though, not something I checked by throwing the same input over and over. This time I go and measure that flat assertion itself.

Start by trying just one thing.

Pick 1 question you always throw at the AI, open 2 new sessions, and throw it twice with the same wording.

When you line up the 2 answers that come back, what is the same in them, and what is different?

CC BY 4.0

I went looking for the setting that makes it return the same answer, and came up empty

For a long time I thought that a test that uses an AI reproduces if you pin the sampling temperature to 0. temperature=0, and pin seed too if you need to. That was the standard move back when I learned it, and even now, if you ask an AI, that is usually the practice that comes back.

I looked it up again, and it was gone (checked 2026-09-08).

Model generationtemperature / top_p / top_k
The current top models (Fable 5 / Opus 5 / Opus 4.8, 4.7 / Sonnet 5)Removed. Sending them returns a 400
The generation before (Opus 4.6 / Sonnet 4.6)Usable
Haiku 4.5 and earlierUsable

There is no parameter corresponding to seed in the first place. On the current top models, there is no setting you can put down that makes it return the same answer.

And this was not true even before they were removed. The official reference says this in its description of temperature.

Even at temperature 0.0, the results will not be fully deterministic

So the "pin it and it reproduces" I was carrying was wrong twice over. The setting is gone, and before it went it was not a guarantee either. At this point 1 thing swaps out for another in what there is to check. Whether the same string comes back is not worth checking. What is worth checking is what does not move while the string moves.

What moved was the items, and the kinds did not move

I made the material with a script. It is a list of 40 things that caught my eye. Each one is made of a file path, a few lines of code, and a 1-sentence explanation. 6 kinds of defect are mixed into the 40, and which entry is which kind is decided by the generator. I put that mapping table outside the directory handed to the workers. The names of the kinds appear neither in the list nor in the way I ask, not once.

What I asked for was 1 thing only. "From these 40, pick 3 to fix first, and write the ID and the reason 1 line each, in priority order." The directories handed to the 12 workers are identical down to the byte, and I match them against the hash printed at each generation.

All 12 workers wrote back 3 items. Lined up, they come out like this.

How it is summed upResult
The reason sentences36 of them, 36 distinct. Not 1 sentence is shared
The items pickedSpread over 9 of the 40. The most picked one got 9/12 workers
The kinds of item2 of the 6 kinds. ⭐ The remaining 4 kinds got 0 items
What was put 1stThe items came to 4 distinct answers. The kind was the same 1 kind for all 12 workers (⚠️ measured again in a different wave it was 29 of 30 workers. → "Caveat")

Each time you make the way of summing up 1 step coarser, the spread disappears. The strings come to 36 distinct forms, the items to 9, the kinds to 2, and the kind put 1st to 1. What I wrote in series 2, that the number of times is unlimited but the kinds converge, pointed the same way when measured with 12 workers.

It converged as far as 2 kinds, though. Not 1 kind. This is where the thing I most wanted to see in this article shows up.

KindTimes pickedWorkers that raised it
A secret leaks into a record or to the outside3012/12
What was deleted does not come back66/12

Half of the workers had raised the other kind.

And the other half did not raise it once. Had you asked only 1 worker, whether you hit this kind would be even odds.

Here, let me fold the 12 answers into 1. Take a majority vote and keep only the top 3 items by votes. It is the aggregation that gets done all the time.

The 3 that remain were, all 3 of them, "a secret leaks." 6 of the 9 items that had been mentioned fell away, and "what was deleted does not come back," which 6 workers had raised, had not 1 item left.

What I understood fits in one sentence.

Throw the same question over and over and the items that come back scatter while the kinds line up. So when you gather them into 1 by majority vote, what disappears is not the duplicates but the kinds held by the minority.

How to measure: hand the same 40 to 12 workers with the kinds hidden

This experiment has no condition. The previous 2 articles measured what changes between A and B, but this time I only throw the same thing 12 times. So making everything identical becomes the whole of the design.

I checked that they were identical with a script. That the hashes of the directories handed to the 12 workers agree on 1 value, that the count per kind across the 40 is as specified, and that not 1 of the 40 paths overlaps. If things inside the same kind look too much alike, the worker notices that the same thing is lined up here, and what is being measured changes. So I prepared 6 to 8 templates for each of the 6 kinds, redrew every file name, symbol name and number, and mixed the 40 together in the list.

The way I ask does not contain 1 word about kinds. "Kind," "category," "heavy," "serious": let even 1 of these words in and the worker starts reading by kind. From the moment the question is handed over, all that can be measured is how it answers. This is a pitfall I fell into once in an earlier experiment, and this time I put the forbidden words into an automated test and made it fail on them.

What gets counted is only the 2 columns the worker wrote, ID and Reason. The mapping to kinds is done on my side, with the mapping table.

# From n workers' answers, count the spread of the items and the convergence of the kinds separately (a shortened version of the implementation on my machine)
python3 - kind_map.tsv body_*/PICK.tsv <<'PY'
import collections, hashlib, sys
kind_of = dict(line.split("\t")[:2] for line in open(sys.argv[1]).read().splitlines() if line)
picks, reasons, first = [], [], []
for path in sys.argv[2:]:
    seen = []
    for line in open(path):
        cell = line.rstrip("\n").split("\t")
        # Leave out unknown IDs, and an ID the same worker wrote twice
        if len(cell) < 2 or cell[0] not in kind_of or cell[0] in seen:
            continue
        seen.append(cell[0])
        reasons.append(hashlib.sha256(cell[1].encode()).hexdigest())
    picks += seen
    if seen:
        first.append(seen[0])
ids = collections.Counter(picks)
kinds = collections.Counter(kind_of[p] for p in picks)
print(f"distinct reasons {len(set(reasons))} / {len(reasons)} rows")
print(f"items {len(ids)} distinct / most picked by {ids.most_common(1)[0][1]} agents")
print(f"kinds {len(kinds)} distinct {dict(kinds.most_common())}")
print(f"top pick: items {len(set(first))} distinct / kinds {len(set(kind_of[i] for i in first))} distinct")
top3 = [i for i, _ in ids.most_common(3)]
print(f"kinds left after majority-rounding to 3: {sorted({kind_of[i] for i in top3})}")
PY

Not 1 character of the reason text gets printed. All that gets printed is how many distinct hashes there were. Line the text up and you read it, and reading it lets in the impression that they are all saying roughly the same thing. That impression is exactly what this article wanted to replace with numbers.

What I stopped aggregating

Folding n workers' answers into 1 list. I used to throw the same investigation at several workers, match up what came back, remove the duplicates, and make it into 1 list. It got easier to read by exactly as much as the duplicates removed, so I thought that was the natural thing to do.

Now, before folding, I print how many workers said each kind. In the measurements above, that is the 2 lines 12/12 and 6/12. Those 2 lines are left nowhere in the folded list.

What you can read after foldingWhat you can only read before folding
Which items were raisedHow many workers raised that item
How many there areWhich kind every worker agreed on
—⭐⭐ Whether there is a kind only half of them raised

12/12 and 6/12 mean different things. The first is "with the same input anyone gets there," and the second is "whether you get there is a coin toss." Fold them and both become the same single finding.

And depending on how you fold, the second one disappears whole. That is what disappeared in the majority vote that keeps the top 3. What was meant as dropping the items with few votes was, in fact, dropping 1 kind.

The 17 I counted in series 2 is a record of my turning down the AI's proposals on typingtube, a service of mine that is in production. What let me write back then that it had "stopped at 17" was that I had written down the reasons I turned them down by kind. Had I piled them up as items, it would not have come to 17, and I would not have known whether it had stopped. The worth of counting by kind was written in series 2, but the part about the kinds shrinking if you fold n of them before that was not.

Caveat: this is not proof that "the AI's judgment is right"

This experiment has no answer key. I have not decided which of the 40 ought to be fixed first. All I measured is the degree of agreement, and nowhere am I saying that what the 12 workers raised together is right. Agreeing and being right are separate things, and mistakes agree too when they come out of the same prior distribution.

The assignment of the 6 kinds is something I made. What you count as 1 kind changes how the convergence looks. The "converged to 2 kinds" here is 2 kinds under this way of dividing them, and does not mean there are 2 of something inside the AI.

I fixed the number they had to pick at 3, so the width of the spread is tied to 3 as well. With "raise 5 items," there would have been more room for the lower kinds to get in. Whether the kinds increase when you increase the number is something I have not measured.

The denominator is 12 workers, the model is 1, and the condition is 1 as well.

And because there is a cap on how many run in parallel, I started the 12 workers in 3 waves. A wave is not a condition, but I cannot say the congestion that day was exactly the same.

After writing this article, I threw the same question at another 30 workers. In different waves, thrown again in 3 batches.

There the kind put 1st was 29 of the 30 workers, with 1 worker on a different kind. Together with the 12 workers in this article, across 4 runs it is 41 of 42.

So the "all 12 workers" in the table above holds only for that 1 measurement. The agreeing side got stronger as the runs went up, and only the "all" stayed weak. This is not a counterexample to what this article claims, it is an instance of the claim itself. Carrying a number from 1 measurement over as it stands is dangerous, which is what this article is about, and I did that myself 1 time.

Unlike the previous 2 articles, everything this time could be measured from the artifact alone. The items picked and the reasons are both written into PICK.tsv, so there was no need to open the record of the work. This is not to say that this article is an easy one: what I wanted to measure simply happened to be left in what was produced from the start.

How to run the experiment: throw the same input at n workers and count the items and the kinds separately

Prerequisites

  • That you can start n workers (10 or more if possible) that do not carry the conversation over
  • That you can make the choices with a script. Write them by hand and the writer's own ranking gets into the material
  • That you can match, by hash, that what each worker is handed is identical down to the byte

Time required: 40 minutes (if you write the generator for the material and the counter yourself. 10 minutes if you already have them)

Steps

  1. Make the choices with a script. Decide on a few kinds, prepare several templates per kind, and redraw every file name, symbol name and number. Keep the count per kind level (skew it and the skew looks like convergence)
  2. Put the mapping table of the kinds outside the directory handed to the workers. Do not let 1 word of a kind's name into the text of the list or into the way you ask
  3. Shuffle the order. If the same kind sits together in the list, the order itself becomes a clue
  4. Drop every word to do with kinds from the way you ask. Make "kind," "category," "heavy," "serious" and the like forbidden words, and look for them with a script. 1 word getting in is enough to leave you measuring only how it answers
  5. Make n copies of the same directory and check that the hashes agree on 1 value. Once that breaks, you cannot separate whether the spread came from a difference in the material or from the output moving
  6. Start n workers. Do not change 1 character of the prompt apart from the path
  7. Count at 3 levels of coarseness. How many distinct reason sentences, how many distinct items, how many distinct kinds. Always do the mapping to kinds with the mapping table (classify by reading the text and the reader's own judgment gets in)
  8. Fold it into 1 and see what disappears. Whether a kind fell away when the majority vote kept the top k items

Pass conditions (all of them have to hold)

  • The hashes of the directories handed to the n workers come to 1 distinct value (= the material is identical)
  • The reason sentences come to n × k distinct forms (= "the same string does not come back" has been caught)
  • The kinds have not collapsed into 1 kind (if they have, the way you made the choices is skewed)
  • The kinds have not spread across all of them either (if they have, the way you divided the kinds is too fine)

If it does not pass

  • Everyone agreed down to the item: 1 entry among the choices stands out. Raise the count for that same kind
  • Even the kinds came apart: the way you divided the kinds is too fine. Merge them and bring the number down
  • Only 1 kind came out: the way the choices were made is skewed. Level the counts again and shuffle the order
  • The hashes came to 2 or more distinct values: something that changes at each generation (a timestamp, a random number, a path) is in the material. Remove that first

Cleanup

  • Delete the directories handed to the workers, and the mapping table of the kinds. Leave the mapping table and the next time you rebuild the material, you will be reading it through the old table

The strings came to 36 distinct forms, the items to 9, the kinds to 2, and the kind put 1st to 1. Of these 4 numbers, the only one that survives being folded into 1 is the rightmost.


Next time, the art of not being useful. I stopped writing lines that start with "if needed" in my instruction file. Even so, work that was needed has never been left out.

CC BY 4.0 はここまで