This is article 9 in the series "The Art of Not Telling." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the introduction.

This time it is guessing. The AI sometimes talks about the contents of a file it has not opened as though it had opened it. This article deals with how you keep it from doing that.

In article 4 of the previous series, "The Art of Not Reading," I wrote that whether something was read cannot be measured. What can be measured is only whether a file that is not created unless you look at the documentation is there. The same question stands for guessing. If whether it was read cannot be measured, whether it guessed cannot be measured either.

Start by looking for just one thing.

Open your instruction file (CLAUDE.md, AGENTS.md) and search for "do not guess" and "check first." How many lines are there? Of those, how many are lines that make something fail when the check does not happen?

CC BY 4.0

My instruction file had just 1 line

Do not implement on a guess. Check the existing code before you implement

It is 1 line out of 102. Anthropic's prompting guidance gives the same sentence as well, as its example of a prompt for reducing hallucinations: never guess about code you have not opened, and when a particular file is mentioned, always read it before you answer.

It is the right line.

And it has exactly the same shape as the "always run this" I dealt with in article 1. It works only when it is read. This time it is worse than that: whether it checked cannot be measured. "I checked it" is a self-report even when it is written down. As I wrote in article 9 of the previous series, a self-report leans toward "done."

So the natural move here is to add instructions. "Always open the file before you answer." "Write down the line numbers you read." This article does not do that. Adding is article 6's job, and in this series I have settled that "tell it" gets said only 1 time.

Another place in the same repository was in a shape you cannot write without opening it

On my own machine there is a document that lays out the principles behind the experience design. It holds 18 statements of the form "this is the experience we build," written so that a program can read them. Here is 1 of them.

<!-- experience-principle
id: A1
name: give it no name
severity: critical
artifacts:                          # ← where this principle is implemented
  - app/models/character.rb
guarded_by:                         # ← the automated test that guards it
  - test/reference/a1_name_field_absence_test.rb
forbidden_patterns: []
-->

I write the principle itself in prose.

But without these 8 lines attached, it does not stand as a principle. Counted up, the 18 principles carry 53 artifacts references and 27 guarded_by references. Writing 1 principle names 4 or more files on average.

And naming them is not enough to pass.

Automated testWhat it looks at
1every principle holds at least 1 artifact and guarded_by
2the artifact it names exists
3the automated test file it names exists
4a # guards: A1 tag is there inside the automated test file
5an @experience-principle A1 marker is there inside the artifact

The AI does not pass unless it opens the file it points at and writes a marker inside it. Name a file that does not exist and it fails at 2; name one that exists but was never opened and it fails at 5.

Not 1 line of instruction was added. What was added is a requirement on the format.

What I understood fits in one sentence.

Whether it was read cannot be measured. So make the format require something that cannot be written without reading.

The mechanism: make it write a marker into what it names

What you put down is 1 automated test (the code on my machine is 6 meta-tests, but this is the whole skeleton).

# test/reference/experience_principles_test.rb (a shortened version of the code on my machine)
test "each artifact carries an `@experience-principle` marker for a known id" do
  known_ids = @principles.map { |p| p["id"] }.to_set

  @principles.each do |p|
    p["artifacts"].each do |path|
      content = File.read(REPO_ROOT.join(path))          # fails right here if the file does not exist
      markers = content.scan(MARKER_REGEX).flatten

      assert_includes markers, p["id"],
                      "principle #{p['id']}: marker not found in #{path}"
      markers.each do |id|
        assert_includes known_ids, id, "unknown principle id `#{id}` in #{path}"
      end
    end
  end
end

It also fails when a marker points at an unknown ID. Rewrite only one side, the document or the code, and it does not pass.

Anthropic's guidance on guardrails says this as well: give the model explicit permission to say "I do not know" and false information drops sharply. This structure hands that permission over from the format side. What you cannot write a marker into cannot be written as a principle. There is no route in the first place for forcing through something that cannot be written.

What I stopped making it think about

The "this file is probably like this" part. The amount of thinking has not gone down. To write 1 principle, the AI now opens 4 or more files and puts a marker in each. What went down is the route that lets it write without opening anything.

Caveat: there are more markers than names

Let me be honest. What this structure guarantees runs in one direction only.

Before I counted, I took it that the names and the markers matched 1 to 1. They did not. The real files the principles name come to 43.

Meanwhile there are 75 files with an @experience-principle marker written in them. 32 of them are named by no principle at all.

So: "what it named, it opened" is guaranteed. "everything it opened, it named" is not. Delete 1 name and the marker is left floating while the automated test still passes.

One more. This shape holds because there are 18 of them. With 200 principles, the cost of putting the markers in breaks down first. This is not a shape for every rule. It is for the side that article 6 of the previous series set aside as "cannot be translated into an automated test": the design of policy and of the experience.

How to verify: name a file that exists without opening it

Prerequisites

  • You have finished putting down the naming format (the field that lists the files it points at) and the automated test that requires a marker at the other end
  • That automated test passes right now
  • You start from a state with no work in progress (git status is empty)

Time required: 10 minutes

There are 2 ways to break it, and the 2nd is the hard one.

Steps

  1. You add 1 file that does not exist to the list of names (a path that is obviously not there, something like app/models/does_not_exist.rb). Run the automated test and note the result
  2. This time you add 1 file that exists but has no marker (a file that is certainly there, something like config/routes.rb). Run the automated test and note the result

Pass conditions (all of them have to hold)

  • In step 1, the automated test fails
  • The message in step 1 shows that this is the path that was named
  • In step 2 as well, the automated test fails
  • The message in step 2 gives a different reason, that the marker was not found

If it does not pass

  • It does not fail in step 2 — this is the real point. That automated test is looking at existence alone. The route that names a file without opening it is still left open
  • The same reason comes out in step 1 and step 2 — the 2 kinds of breakage are not being told apart. Whoever fixes it (the AI, or you) cannot tell which of them to fix

Cleanup

  • Delete the 2 lines you added and confirm that the automated test goes back to passing and that git diff is empty

Step 1 detects a mistyped path, and step 2 detects that something was named without being opened. Stop at step 1 and call it done, and you take an automated test that holds nothing but an existence check for one that holds existence plus confirmation.


Next time it is "The art of not answering questions." Even when the AI asks me "Which shall it be, A or B?" I do not answer. Even so, the work does not stop.


Series: The Art of Not Telling

End of CC BY 4.0