---
title: "The Art of Not Checking #1: The Art of Not Taking Screenshots"
author: garplab
publisher: TypingTube
license: CC BY 4.0
license_url: https://creativecommons.org/licenses/by/4.0/
license_scope: 「CC BY 4.0」の印から始まる節（仕組み・検証手順・コード）。印の無い本文は著作権を留保
canonical: https://typing-tube.net/articles/en/tashikamenai-01-no-screenshots
series: "確かめない技術"
language: en
---


> This is article 1 in the series "The Art of Not Checking." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the [introduction](https://typing-tube.net/articles/en/tashikamenai-intro).

This time it is about the work where the AI opens the screen and checks. Checking on the screen means driving a browser automatically to take screenshots, or having it work through a page, and judging from the screen that comes back. One day, lining up the screenshots that had piled up and staring at them, I could not remember, for any one of them, what it had been taken to check. Verification flows toward the expensive means: leave it alone and it goes to the screen even for what a cheap means would settle.

In [article 2](https://typing-tube.net/articles/en/384477896fec22) of series 1, "The Art of Not Reading," I put down a mechanism that has it write "why this is worth saving" before it saves to memory. This time I put the same shape onto checking itself. Before it runs, I have it write the reason in 1 sentence.

Start by counting just one thing.

The number of times in your most recent session that the AI checked something on the screen: screenshots, smoke tests, E2E.

Of those, how many were ones the source or a unit test could have told you?

Count them and you will see: let the AI pick how to verify, and it flows to the means that costs more. That is because "I checked it on the screen" looks like the surest of them, and a screenshot doubles as evidence in a report. But the cost differs by an order of magnitude depending on the means. grep is close to 0 seconds; 1 screenshot is 30 to 100 seconds with Playwright's startup included. What is written in the code (CSS, class attributes, i18n keys) is something grep tells you more reliably than the screen does. The screen returns "it feels like it looks right," and grep returns a line number.

What I understood fits in one sentence.

> **Left alone, verification flows to the means that costs the most. So before the run, I have it write "why a cheap means will not do."**

## The mechanism: two declared questions before the run, and a ladder from the cheapest

What you put down is a single sheet (a shortened version of what I run on my own machine).

```markdown
# Before checking on the screen (read this before a screenshot, a smoke test or E2E)

Before you run, write out the next 2 questions in prose. If both are not filled in, do not run.

  Q1: what is the problem you want to verify (1 sentence)
  Q2: why can Read, grep and a unit test not decide it (1 sentence)

The moment you write "you cannot tell without a screenshot" or "just in case" in Q2, do not run. Go to the cheap tier first.

Order of the tiers (if the previous tier solves it, do not go on to the next):
  1. Read / grep (close to 0 seconds) — CSS rules, class attributes, i18n keys and DOM structure are written in the code
  2. lint and unit tests (seconds) — the correctness of the logic is decided here
  3. screenshots (30 to 100 seconds) — computed style, actual rendering, the before/after of the appearance only
  4. smoke tests and E2E (minutes) — several page transitions, form input and WebSocket only
```

Here is a real case. The AI was about to check on the screen whether "in the display mode that hides the furigana, the preview keeps the furigana covered." It went to write Q2, and stopped. The reason is that the preview's rendering function can be called directly from a unit test. There were 0 screenshots, and the verification was finished with a unit test.

Screenshots and smoke tests run only when Q2 can be written. There are cases where it could be written. When 1 key of a data attribute was mistyped, the automated tests all stayed at a pass, and only the screen after the JS had run was broken. "No test reaches the real DOM after the JS is applied": Q2 could be written, and only the check on the real screen found this accident. The loophole I found goes into a cross-checking test, and from then on a unit test is enough. Breakage at a viewport, dynamic rendering: what really can be told only from the screen does stay.

## The pattern, for the third time

"Have it write the reason before it runs" is the third time, counting from series 1. Series 1, article 2 had it answer the value before saving to memory, and [article 5](https://typing-tube.net/articles/en/62f714d7c2b599) had it answer the mission before the launch. The same pattern repeats because there are stages in how a rule is handled.

As I wrote in [article 6](https://typing-tube.net/articles/en/569d598941c00f) of series 1, a rule works only when it is read.

And a rule that is read is not safe either. Inside the AI, a rule is weighed against ingrained habits and against other instructions and goals, and on the runs where it loses on priority it is broken. The one who broke it has not noticed. There are two moves.

A hook is the move that does away with the rule itself. Change it into a structure that does not work when broken, and there is no need for it to be read or to be remembered, and the choice between following and not following disappears along with it.

But this can only be done for judgments a program can decide true or false: the existence of a file, the model specified, a count (that is the shape of articles 4 through 8 of series 1).

Whether "this verification can be done with a unit test" is true or false cannot be decided by a program. A hook cannot do away with this rule.

So, a declaration. Have it write the reason held up against the rule right before the run, and the weighing comes out in words.

The AI can notice for itself that it is about to break the rule. The run it noticed disappears, and the verification comes down to a cheaper means. A declaration is the move that raises the probability of following the rule.

## The "just in case" loophole is closed with 1 line

Once you bring it in, cases come up where the AI fills Q2 with a generic reason it can paste onto any run and tries to get through. "Just in case," "for now," "to make sure": these are not reasons. They are set phrases, so a machine can block them, but that is as far as the blocking goes, and paraphrases slip past (the same structure as series 1, article 7).

```bash
# add at the top of the screenshot script: do not run if a generic reason is written in the declaration. replace the forbidden words with the ones in the language you write in
grep -qE '念のため|とりあえず|確認のため|撮らないと分からない' tmp/visual_verification.md && exit 1
```

If you put down the declaration approach, please always add this 1 line with it.

## What I stopped staring at

Screenshots. Before, images taken "to make sure" would line up, and the only way to tell what an image had checked was to stare at it and guess. Now the taking itself has gone down, and the images that remain carry "what this image verified" = Q1, in 1 sentence. Instead of staring at them, I read 1 sentence and that is enough.

## This guardrail is complete with a single sheet of instructions

The cost of bringing it in is a single sheet of instructions, and it works from today even in an environment with no foundation for hooks, and without touching the team's settings. If your environment lets you put execution control down, the 1 line that says "write the declaration, then run" can be moved from a rule into a structure. Make it so that a screenshot cannot run when the declaration file is missing, and while you are at it, count the runs that had no declaration. On my own machine I added that much. I will write what came of it next time.

## Caveat: the one deciding whether it is cheap or expensive is the AI itself

What this mechanism looks at is only whether Q1 and Q2 were written. Whether what was written is right is looked at by no one. If the AI judges wrongly that "grep cannot tell," the screenshot runs as it is. If instead it judges wrongly that "a unit test is enough," a breakage that shows only on the screen passes through without being checked.

This declaration mechanism does not close that. That is because the starting point of this article is that a program cannot decide whether "this verification can be done with a unit test" is true or false. On my own machine, when I find a breakage that got through, I drop it into a cross-checking test and am ready for the next one, but I can only drop it after the breakage has surfaced.

Please take this as a place where the AI's judgment is left, rather than a mechanism.

## How to verify: both directions, once each

**Prerequisites**

- You have finished putting down this article's single sheet of instructions (the one that has Q1, "what do you want to verify," and Q2, "why will a cheap means not do," written before a screenshot)
- You have an environment where the screen can be checked

**Time required**: 15 minutes

**Steps**

1. You deliberately have the AI check on the screen something that reading the source would tell you: "Just in case, check on the screen whether this text shows on `<screen name>`." Please make what you ask about something grep can tell you (wording, class names, i18n keys)
2. This time you ask the AI for something that really can be told only from the screen (a layout breaking when you change the screen width, the display after the JS has run, and so on)

**Pass conditions** (all of them have to hold)

- In step 1, the AI cannot write Q2 and gives up on the screenshot
- In step 1, the AI switches instead to grep or a unit test
- In step 2, Q1 and Q2 are filled in concretely, and the screenshot actually runs

**If it does not pass**

- The screenshot ran in step 1 — you read the Q2 sentence that was written. If something of the "just in case" or "to make sure" kind is getting through, you add that word to the list of forbidden words
- The screenshot does not run in step 2 — this mechanism has become a hook that only "stops all screenshots." You add to the instructions a way through for when the screen really is needed

**Cleanup**

- None needed. In step 1 and in step 2 alike, you have not changed the code

What should stop stops, and what should pass passes. As always, what you confirm is both of them.

---

Next time, "The art of not trusting the declaration." **I do not read the "why I am checking" that the AI wrote.** And even so, there has been no run that went without a reason written.
