Agent was deferring framework feedback and state writes to closeout, then
reconstructing what happened from conversation memory after compression —
the exact mechanism that fabricates false positives.
- Guard now self-logs: _check_command_internal appends a Guard Block/Review
row to state/framework_feedback.md at the moment check_command rejects a
command (only when the file exists, i.e. benchmark engagements; no-op
otherwise). Friction is captured with zero agent bookkeeping.
- SKILL.md Operational Contract: 'Record as you go' — after EVERY
violin_review_batch, immediately update hypothesis board + coverage matrix
+ PTT in the same turn; never batch state writes to closeout, never
reconstruct tests from memory. State files are the only source of truth.
- 5 new tests: block-row append, no-op without file/errors, dedupe, pipe
escaping.
243 tests pass, ruff clean.