Files
agent-skills/skills/verification-before-completion/SKILL.md
T
iWP Claudy 39a637410c Initial skill library: 9 skills for delegate briefs
WordPress plugin rebrand/conventions/remote-CLI patterns, Gitea release
workflow, bastille jail provisioning, remote shell quoting safety, server
fleet map, delegate brief writing, and verification discipline -- all
derived from real incidents this session, plus two skills adapted (MIT
license, attributed) from obra/superpowers and andrej-karpathy-skills.
2026-08-02 20:08:29 +02:00

66 lines
2.9 KiB
Markdown

---
name: verification-before-completion
description: Use before reporting any task as done, especially when the work was done by a delegate/subagent. Never relay a self-report as fact.
license: MIT
source: adapted from obra/superpowers (MIT license), condensed for delegate-brief reuse
---
# Verification Before Completion
**A delegate's "done" is a claim, not a fact.** This applies whether the
delegate is another AI agent, a subagent, or your own past self in an
earlier step. Re-derive the result yourself before reporting completion
to whoever is waiting on it.
## The rule
Never report "done" based solely on:
- A delegate saying it's done.
- A test that wasn't actually re-run by you.
- Code that "should work" based on reading it, when running it was possible.
Instead, for each claim, do the cheapest check that would actually catch
a wrong claim:
| Claim | Real check |
|---|---|
| "The file was created" | `ls`/`cat` it yourself |
| "The syntax is valid" | Actually run the linter yourself, don't trust a self-report of having run it |
| "The tests pass" | Run them yourself, read the actual output |
| "The API returns X" | Make the actual call yourself |
| "The deploy succeeded" | Check the live system, not the deploy log's exit code alone |
| "The old code is gone" | grep for it yourself |
## Why this matters (not just process theater)
Across a real multi-delegate session, every single delegate's first
self-report claimed success. Independent re-verification (re-running the
linter, re-grepping, and — critically — testing the actual live
end-to-end pipeline rather than just checking that files existed) is
what actually confirmed correctness. In an earlier phase of the same
session, a wrapper-agent pattern (spawn an agent whose only job is to
invoke one CLI tool and report back) repeatedly ended its own turn the
instant it *started* a background process, before that process had
produced any real result — producing a loop of confident-sounding "still
running" reports that carried zero information. The fix was to stop
trusting the wrapper's self-reported status and directly monitor the
actual underlying process.
## Escalating verification for higher-stakes changes
- Low-stakes (a text file, a doc change): read it back.
- Medium-stakes (application code): lint + a targeted functional test.
- High-stakes (something that serves real traffic, handles money, or
changes shared/production state): the *actual* end-to-end path a real
user would hit — not a unit test in isolation, the real live pipeline.
If you cannot exercise the real path yourself (e.g. no browser
available for a checkout flow), say so explicitly rather than reporting
success on a partial check.
## What NOT to do
Don't pad out a report with hedging language as a substitute for actually
checking ("this should work", "this is likely correct"). Either verify
it, or clearly state it's unverified and why — don't blur the line
between the two.