---
title: "The Art of Not Checking #2: The Art of Not Trusting the Declaration"
author: garplab
publisher: TypingTube
license: CC BY 4.0
license_url: https://creativecommons.org/licenses/by/4.0/
license_scope: 「CC BY 4.0」の印から始まる節（仕組み・検証手順・コード）。印の無い本文は著作権を留保
canonical: https://typing-tube.net/articles/en/tashikamenai-02-declaration-guard
series: "確かめない技術"
language: en
---


> This is article 2 in the series "The Art of Not Checking." It lays out symptoms that go wrong and their remedies, one at a time. Each article is finished once you put down a single file or script. Why that mechanism is needed becomes clear when you read the explanation afterward. The whole picture and the list of articles are in the [introduction](https://typing-tube.net/articles/en/tashikamenai-intro).

This time it is the declaration. A declaration means having the 2 questions put down last time, what you want to verify and why a cheaper means will not do, written out in prose before anything runs. After I put it down, I was never once able to count how many runs had gone ahead without one being written.

In [article 4](https://typing-tube.net/articles/en/8f9ae8791b448d) of the first series, "The Art of Not Reading," I wrote that you do not measure whether it was read. You look only at the state where nothing can move unless it has been written. This time I turn that measure on last time's declaration itself.

Start by looking for just one thing.

Pick 1 run from your most recent session where the AI ran a screenshot or a smoke test.

Is there a sentence left anywhere, from just before that run, saying why it was needed?

Last time the mechanism was "have Q1 (the problem you want to verify) and Q2 (why a cheaper means will not do) written before a screenshot." The moment it goes to write them, the AI notices for itself that it is about to break a rule — it works as a guardrail. One issue is left: a run that skips the declaration step goes straight through. Going straight through leaves no record, so how many times it was happening also stayed a guess.

What I understood fits in one sentence.

> **A declaration works. But it works only when it is written. So make the trace that a declaration was written a precondition for the run.**

## The mechanism: make the existence of the declaration file a precondition for the run

What I put down is a single PreToolUse hook (a shortened version of how it runs on my own machine).

```bash
# picks up only the entrances for screenshots and smoke tests (SKIP_VISUAL_GUARD=1 goes through as the explicit escape hatch)
case "$command" in *SKIP_VISUAL_GUARD=*) exit 0 ;; esac
printf '%s' "$command" | grep -qE 'screenshot\.sh|smoke_[a-z_]+\.js' || exit 0

# 1. it cannot move if the declaration file is not there
[ -f tmp/visual_verification.md ] || { guide; exit 2; }

# 2. reject the forbidden words mechanically (this is the only part a program can decide; replace them with the ones in the language you write in)
grep -qE '念のため|とりあえず|確認のため' tmp/visual_verification.md && exit 2

# 3. are Q1 / Q2 filled in with 1 sentence each (rejects blanks and untouched templates)
# 4. freshness of 30 minutes (rejects a declaration reused from the previous check)
```

The line is drawn the same way as last time's dichotomy. Whether "a unit test will do" is true or false is not something a program can decide, so the hook does not measure whether what is inside is sound. What it measures is only the trace that the judgment happened before the run: the existence of the file, the 2 questions being filled, the forbidden words, the freshness. In last time's terms, a declaration is a move that raises the odds of a rule being followed, and the hook is a move that does away with the rule. What I did away with here is only the 1 line "write the declaration, then run," and I have not dropped the declaration approach. That a list of forbidden words gets slipped past by rephrasing is also as I wrote last time, and what I closed off here is only the easiest route.

## A prohibition with no replacement produces the abandonment of verification

Once I had built it, I found that stopping is not enough. At the moment of the block, "check it with a screenshot" is stacked on the AI's TODO list. Stop only the run, and the task is left hanging, or verification is abandoned along with it. Work that closes with "I could not check it" is a worse ending than a screenshot running.

So I spelled the replacement out in the guidance. "Rewrite 'check with a screenshot' on the TODO list as 'check with grep / a unit test' and work it off. If you really do need the screen, write Q1/Q2 and run it again." It is a variation on the pattern from [article 8](https://typing-tube.net/articles/en/f7c9c30f567aee) of the first series: a policy is read only when it is inserted at the moment of the problem. If you put down a prohibition, write the next move on the very screen where you stopped it.

## Measured: on the day I put it in, 2 were caught

The 1st: an automated test found a bug in the hook itself. In bash, `${#var}` returns bytes rather than characters depending on the locale. "見たい" ("I want to see it") is 3 characters but 9 bytes. A lazy 3-character declaration was going straight through the check that "fewer than 8 characters counts as blank." It was found because I had written a test item for the blank check, and I changed the character count to be done with python3.

The 2nd: the hook stopped me, the one who had put it in. It is a false positive: the command I was checking the behavior with was stopped just because it contained the string `screenshot.sh`. I could get through with the escape-hatch environment variable I had written into the guidance. A guardrail is bound to misfire. What matters is that the way out is written somewhere you can read it the moment it misfires.

The hook is a guardrail and a measuring instrument at the same time. Every time it judges, it appends 1 line to a record file (the date and time, the reason it stopped, the first 120 characters of the command), so I have started counting "how many runs there are without a declaration," which had stayed a guess up to last time. In 4 days it went through this hook 22 times, and it stopped 7 of them. All 7 were the same false positive as the 2nd case. The name of an entrance happened to be in the command string, and not 1 of them was a screenshot run. 2 of them are the command that writes the declaration file itself. The reason left in the record is "reuse of the previous declaration," but that is the name of the check, not what it stopped. Neither a screenshot without a declaration nor a forbidden word has come up once in these 4 days.

## What I stopped reading

What is inside the declaration. I have not once read the "why check it" that the AI wrote.

The hook is set not to measure whether what is inside is sound, so there is no one anywhere left to read it and decide whether it is any good.

What I read is only the "when, and what was stopped" that adds 1 line to the record when something stops. Even so, there is no run that went ahead without a reason written. If it is not written, the run itself does not start.

## Caveat: I am looking at the form, but not at what is inside

A declaration that got through this hook is not necessarily a correct declaration. It goes through if Q1 and Q2 each hold 1 plausible-looking sentence, there is no forbidden word, and it was written within 30 minutes. Write "this is a screen that cannot be called from a unit test" and it goes through. Whether that is true is not something a program can decide.

This hook does not close that off. Trying to close it off means building one more stage that judges whether what is inside is sound, and that stage ends up being someone's judgment again. What I got here is only the trace that "the judgment happened before the run," and the quality of the judgment itself stays on the AI's side, as it was last time.

One more thing. What this hook can pick up is only the names of the entrances I wrote down as things to pick up. That is exactly what happened on my machine: what I wrote in the hook was the file name of the script that runs on the inside.

The wrappers I actually type are another 14, and those have not been looked at once since the day I put them down. Add an entrance, and add it on the hook's side too. Mistake a name, forget to add one, and the hook lets it through without saying anything.

## How to verify: what should stop stops, and what should pass passes

**Prerequisites**

- You have registered this hook, and you have finished deciding where the declaration file goes and what is on the list of forbidden words
- You have the screenshot command on your machine

**Time required**: 20 minutes (if you write the 9 items into automated tests all at once, it takes a little longer)

**Steps**

You hand the 6 cases below to the AI one at a time and note down how the hook reacts. Not finishing after checking only the side that stops is the whole of this section. The 4 that stop and the 2 that pass are set side by side.

| # | What you hand over | The result you expect |
|---|---|---|
| 1 | Have it run the screenshot command **without putting down** a declaration file | It stops. The guidance says "write Q1 and Q2," and **what to rewrite it as** comes out (screenshot → grep / a unit test) |
| 2 | A declaration with **only a forbidden word** written in Q2 ("just in case") | It stops. **Which word is forbidden** comes out |
| 3 | A declaration where Q1 or Q2 is **empty** | It stops |
| 4 | An **old** declaration (left over from the previous piece of work) | It stops. **When the declaration is from** comes out |
| 5 | A declaration where Q1 and Q2 are filled in concretely | It **passes** (the screenshot runs) |
| 6 | The escape hatch written into the guidance (an environment variable or the like) | It **passes** |

**Pass conditions** (all of them have to hold)

- 1 through 4 stop in the way you expect
- 5 and 6 pass without stopping
- Delete the 1 line of the forbidden-word check from the hook and 2 gets through (= the automated test that watches that item notices the deletion)

**If it does not pass**

- 5 and 6 stopped — this hook has become "a thing that only stops all screenshots." An automated test that checks only that things stop cannot show the way the hook breaks when it stops everything: the same pitfall as the "automated test that has never failed" seen in [article 6](https://typing-tube.net/articles/en/569d598941c00f) of the first series
- You delete 1 line of a check and the automated test does not notice — that item has been looking at nothing since the day you put it down

**Cleanup**

- You put back the 1 line of the check you deleted, and clear away the declaration file you put down

---

Next time, "The art of not re-running tests." **I stopped re-running tests in order to check.** And I still know, every time, whether it is passing right now.
