Verification Testing 8 Oct 2026

Sabotage-Check Your Tests

Three verification asks worth making by default. Break the code and watch the test fail, run the control case, and count instead of eyeballing.


I've been running the same real production tickets from work through different models side by side. What separated the runs most wasn't the code. It was the account of whether the work had been verified.

On one tightly specified ticket, two models produced byte-identical diffs. The whole difference was in the account of whether it worked.

An agent can find the fact that breaks the spec, understand it, and still report against the spec. Finding it and telling you are separate behaviours.

So I treat verification as a first-class piece of work. An agent producing a patch isn't the same as a verifiably correct change. These are the three asks that have earned their place.

Break the code, watch the test fail

Sabotage-checking is the strongest verification technique I've seen so far: mutate the source, watch the named test fail, restore it in the same script.

In one run the agent reversed the direction of an event scan and six tests failed. That was the only evidence in the whole series that a test bites rather than merely passes.

sabotage.sh · illustrative
#!/bin/sh
# Break the code on purpose, prove the named test notices, put it back.
set -u
FILE=src/events.ts
cp "$FILE" "$FILE.bak"
trap 'mv "$FILE.bak" "$FILE"' EXIT

sed -i.tmp 's/a.at - b.at/b.at - a.at/' "$FILE" && rm -f "$FILE.tmp"

if npx vitest run src/events.test.ts; then
  echo "SABOTAGE NOT CAUGHT: the test passes with the code broken"
  exit 1
fi
echo "Test bites"
The ask

For each test you add or change, break the code it covers on purpose, run that test and show it failing, then restore the code in the same script.

Ask for the control case

Asking for the control case is worth more than paying for a better model.

The single most valuable thing one run did was re-run the scenario without the change, which turned an apparent regression into a disclosed pre-existing quirk. Cheap, mechanical, and it shouldn't depend on model choice. So don't leave it to the model. Ask for it.

The ask

Before calling anything a regression, re-run the same scenario without your change. Say whether the behaviour was already there.

Count, don't eyeball

Counting beats eyeballing for a UI claim.

"The reason appears exactly once on the page" was only sayable because the run counted occurrences of the string across the whole document.

reason.spec.ts · illustrative
// Not "is it visible", but "how many times is it on the page".
await expect(page.getByText(reason, { exact: true })).toHaveCount(1);
The ask

For any claim about what's on a page, count it across the whole document and report the number.

Make them the default

None of these should depend on someone remembering to ask. Put them in CLAUDE.md and every session gets them.

CLAUDE.md · suggested starting point
## Verification

A passing test isn't evidence on its own. Before finishing a task:

- Sabotage-check new or changed tests. Break the code the test covers,
  run that test and show it failing, then restore the code in the same script.
- Run the control case. If something looks wrong, re-run the same scenario
  without your change and say whether the behaviour was already there.
- Count, don't eyeball. For claims about what's on a page, count occurrences
  across the whole document and report the number.

In your summary, say what you ran and what it showed. If you found anything
that contradicts the spec, say so, even if you built to the spec.
Note

This is a starting point, not a copy of anyone's real file. Adjust the wording to your test runner and your repo.

Same family as knip: give the agent a way to check its own work while it still has the context to fix it.