This is article 6 in the series "The Art of Not Listening to the AI's Opinions." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the introduction.

This time it is about the requests to check by eye that come from the AI. Checking has not gone away. What went away is me opening the screen and looking with my own eyes. Every piece of work the AI does ends with: all that is left is checking by eye on the real machine. I do not answer those requests. What I put down instead is a single smoke test where a program decides pass or fail. A smoke test is a test that runs the main operations through to the end in a real browser and sees whether anything is broken.

In article 12 of the first series, "The Art of Not Reading," I put down a shape that lets me get by without reading the finished work the AI hands over. This time it is about being asked to check that same work on the screen.

Start by counting just one thing.

The number of times in your most recent session that the AI ended its work with a check on the screen still left to do.

Of those, how many times did the changed lines come with a unit test?

CC BY 4.0

Checks left to the AI flow toward human eyes

In the production project on my own machine (typingtube, a web service for practicing typing along with music videos on YouTube), the AI's closing line is "all that is left is 2 points to check by eye on the real machine." The memo the AI hands to the next session also keeps "all that is left is checking the real screen by eye (human)." Even though I never once open it, this one line comes every time.

Human eyes are not the only place it flows to. When a color was changed, where the AI headed was toward taking a screenshot and reading the color out of the image. On the day I went back over the color scheme, I warned it first. "Do not take an approach that costs far too much, like taking screenshots of every pattern, extracting the colors from the images and recomputing them to inspect; test by computing directly from the color codes instead."

When I built the on-screen effects, what was checking the logic was a smoke test with 26 scenarios, and there was not a single unit test for the functions. All 26 passed, and on the real machine the effects did not run. Once 101 unit tests had been written and I looked again, 17 of the 26 were confirming logic those unit tests already covered.

What I understood fits in one sentence.

A request to check by eye is a way of checking that spends none of the AI's own tokens, and it hands the pass-or-fail judgment to a human. So do not answer it; replace it with a shape where a program decides pass or fail.

Why the AI flows toward checks costing it none of its own tokens

The Claude Code documentation writes that each tool use returns information that feeds back into the loop and informs the next decision. What a human sees after opening the screen is returned by no tool. With nothing on the AI's side to read, a check handed to human eyes spends not a single one of the AI's tokens.

The Claude Code documentation also writes that the model decides on its own, at each step, whether to think and how much. The same page says the setting for how much to think trades token spend against capability. Choosing how to check is one of those steps too. Human eyes cost 0, one smoke test run costs the few lines it returns, and a unit test per changed line costs the writing, the running and the reading.

The Claude Code best practices put it this way. Claude stops when the work looks done. Without an automated test it can run, "looks done" is the only signal available, and you become the verification loop. A request to check by eye comes out in this shape. On the AI's side there is no pass or fail it can run and read. Claude's prompting guidance also writes that as autonomous tasks grow longer, correctness has to be verified without continuous human feedback.

The Claude Code documentation writes that sessions are independent and do not carry the conversation history of previous sessions. Whether I opened the screen leaves nothing on disk. What reaches the AI in the next session is the other thing, the note the AI in the previous session left behind: all that is left is checking by eye.

The best practices page also spells out what the automated tests it runs can be. A test suite, a build exit code, lint, a script diffing output against a fixture, a browser screenshot compared against a design. The condition is a single one: it returns a signal the AI can read in the conversation. The AI runs it itself, reads it itself, and fixes until it passes, it says.

The mechanism: 1 smoke test, pass or fail decided by the DB

Unit tests looking at the changed lines, and mutations breaking them to make sure, were put down in article 3 of the first series. Even so, some things remain undecided. What goes here is 1 smoke test. The example is the word deck battle. 2 players enter the same room and type the words in the same order. Whether "the 2 of them are being given the words in the same order" is something I cannot tell by looking at the screen. The 2 screens each advance on their own, so the word on display drifts apart easily. In fact, the first smoke test was lining up the words showing on the 2 screens, and it returned a mismatch even though the order was the same. The side that opened its screen first had lost the 1st word to the time limit, and its display was 1 word off. The display moves ahead on the time limit, but the order typed does not change.

The smoke test drives Playwright in Docker, creates the data, plays a battle through to the end in 2 browsers, one per player, and deletes the data when it is done.

# lib/tasks/dev_smoke.rake (a shortened version of the code on my machine) — once the smoke test is over, look at the DB, not the screen
orders = VocabularyPlayLog.where(battle_room_id: room.id).map do |log|
  log.entry_play_logs.in_play_order.pluck(:vocabulary_entry_id)
end
puts "entry_order_match=#{orders.uniq.size == 1}"      # did the 2 players get them in the same order

plan = VocabularyBattleEntryOrder.call(version: record.vocabulary_deck_version,
                                       play_mode: "shuffle", shuffle_seed: record.shuffle_seed)
server_ids = plan.entries.map(&:id)
puts "server_order_match=#{server_ids == orders.first}" # is it the same order the server decided

The result comes back as 2 numbers, like "Smoke FAILED (playwright=1 verify=1)." Whether Playwright ran through to the end, and whether the DB judgments lined up. If either one drops, it is a failure. The AI runs the smoke test itself. Whether to run it is also the AI's decision. The one reading those 2 numbers is the AI. Put a smoke test down, and where the AI flows shifts from human eyes to the smoke test. The lines deciding those 2 numbers sit in a file the AI can edit. There are also situations where nothing is left in the DB. The same battle has one mode that leaves no data, and there alone, the smoke test still looks at the screen.

Whatever slips through arrives from the person who hits it in production

A smoke test only looks at the situations you wrote. The situations you did not write show up in production. I do not go and check those; I take them on a path arriving from the person who hits it. It is the bug report form. Reports are classified by area (typing / battle / lyrics display ...) and by kind (bug / request / wrong lyrics), and they carry a state of unread, read, or handled. The only thing I look at is that pile.

What I stopped checking

The screen. Local environment and production environment alike, I stopped going around opening them. When a request to check by eye comes from the AI, what I send back is the usual "go ahead with the next task." What decides pass or fail instead is the unit tests, the mutations and the smoke test. A program decides every one of them, the AI reads them, and if one fails, the AI fixes it.

How to verify: flip 1 judgment, and confirm that it fails

Prerequisites

  • You have written 1 smoke test, and it contains a judgment that decides pass or fail from the data (in the example in this article, the part that looks at the data left in the DB after a battle)
  • The smoke test runs through to the end on your machine

Time required: 20 minutes (this includes the time the smoke test takes to run)

Steps

  1. You flip 1 of the judgments that look at the data (in the example from the article, you rewrite the line that checks the order of the words so that "a mismatch is a pass"). You flip the judgment line and nothing else. You do not touch the operation side of the test
  2. You run the smoke test
  3. You check the path in production once as well. You deliberately break where the form sends its reports, and look at the number of reports

Pass conditions (all of them have to hold)

  • In step 2, the result comes out a failure
  • The failure in step 2 stands on the data judgment side, not the screen operation side (in the example in this article, the shape that corresponds to playwright=0 verify=1)
  • In step 3, the reports go to 0

If it does not pass

  • Only the screen operation side is failing — the judgment you flipped is not affecting pass or fail
  • Reports keep arriving in step 3 — the destination you thought you broke is not the path that is actually used

Step 3 is the real point. 0 reports can be read as "there are no bugs" and as "the path is severed," and the number on its own does not tell them apart. Break it once and watch it go to 0, and the next time 0 keeps coming you will know what to suspect.

Cleanup

  • Put the flipped judgment line back, run it once more, and confirm that it returns to a pass
  • Put back the destination you broke in step 3, and confirm that reports arrive

Next time, the art of not listening to what the AI says. Even when the AI says something slightly off the mark, I do not explain. Even so, the AI comes back around to the right direction.


Series: The Art of Not Listening to the AI's Opinions

End of CC BY 4.0