Run 20260813_101504 burned ~15 min and 19 close attempts on the
undispositioned-coverage gate: the error listed failures without naming
the fix, so the agent flailed (even reading guard source) and wrote
invalid reasons ('evidence: none else'). The error now spells out the
exact remediation per cell status (tested/not_applicable/blocked/missing).
VIOLIN_BENCHMARK_RECEIPT_KEY -> VIOLIN_RECEIPT_KEY. The all-caps
'BENCHMARK' slipped past the original case-sensitive de-cheat grep in
plugins/violin_guard. No consumers hardcode the literal; all use the
RECEIPT_KEY_ENV constant.
Run 20260812_184313 scored 0/20 formalized despite 7 FIND files and 8
Validated hypotheses: the agent canonized findings carrying 'Linked
Hypothesis: H-00N' but never wrote the forward 'Linked findings' board
field (it closed PT-102 by editing state/ptt.md directly, bypassing the
record_ptt handler gates). The scorer now:
- parse_findings extracts reverse 'Linked Hypothesis' references
- validated_challenge_ids falls back to reverse links / shared evidence
when a Validated hypothesis has no forward links
- paper trail is unchanged: Validated hypothesis + substantive finding
over real evidence, in either direction
Re-scoring 184313: formalized 0% -> 30% (6 confirmed). Regression test
test_validated_hypothesis_confirms_via_find_reverse_link added.
- check_command now accepts hypothesis_id (schema + handler + args)
- Exploit-phase research gate checks only the named hypothesis when
provided; otherwise all candidates (previous behavior)
- Error message names the remediation tool: violin_record_hypothesis
id=H-00N cve_research=... exploit_research=...
- Fixes silent bypass: normalized id comparison (002 vs 2) so an
unresearched hypothesis is actually blocked
- New test: named-hypothesis passes, unresearched H-002 blocked,
all-candidates mode still enforces full rows
- Remove challenge_ids seeding (vulnerability-name leak) from scope;
seed engagement.coverage_obligations (client-style in-scope endpoints)
+ generic engagement.audit_mode flag instead
- plugins/violin_guard now contains zero benchmark/run-id references;
gates are framework-owned and audit-mode-gated
- Cross-engagement guard blocks ANY foreign engagement dir, not just
benchmark-run-*
- Anti-cheat regression test asserts no vuln names in generated scope
The advisory contract rule was not enough: flash-class agents defer
hypothesis/coverage/PTT writes to closeout, then reconstruct results from
conversation memory after compression — the false-positive factory.
check_hypothesis_freshness now hard-blocks further target commands when the
newest execution evidence is newer than the hypothesis board's last update
(15 min grace for burst timing). The agent MUST record the batch result via
violin_record_hypothesis before proceeding. Bursts preflight all commands
before any executes, so no mid-burst false blocks.
Also fixes a silent test-suite gap: tests/__init__.py and the four
subpackage __init__.py files were missing, so 171 integration tests
(incl. test_plugin_guard.py which pins freshness behavior) NEVER ran.
Suite went from 248 passed / 7 collection errors to 419 passed, 0 errors.
Agent was deferring framework feedback and state writes to closeout, then
reconstructing what happened from conversation memory after compression —
the exact mechanism that fabricates false positives.
- Guard now self-logs: _check_command_internal appends a Guard Block/Review
row to state/framework_feedback.md at the moment check_command rejects a
command (only when the file exists, i.e. benchmark engagements; no-op
otherwise). Friction is captured with zero agent bookkeeping.
- SKILL.md Operational Contract: 'Record as you go' — after EVERY
violin_review_batch, immediately update hypothesis board + coverage matrix
+ PTT in the same turn; never batch state writes to closeout, never
reconstruct tests from memory. State files are the only source of truth.
- 5 new tests: block-row append, no-op without file/errors, dedupe, pipe
escaping.
243 tests pass, ruff clean.
- _engagement_needs_closeout no longer fires on leftover unchecked PTT rows:
a single '[ ]' task marker triggered a continuation pass that re-ran the
entire assessment (255 msgs / 20m50s) instead of closing out. Closeout is
now gated on the artifacts the scorer consumes (FIND files + report.md +
retrospective.md) only.
- Closeout goal rewritten to be surgical: assessment is COMPLETE, evidence is
final, no new tests/scans/verification, no evidence modification; only
generate missing report/retrospective (prefer generate-closeout) and update
PTT statuses.
- Friction-log instruction moved out of the general pentest SKILL.md (it is a
benchmark-only artifact) into the benchmark goal prompt so the agent
actually sees it.
- New pinned tests: leftover-PTT-row closeout detection, closeout-goal
wording.
* fix: bump distribution version to 3.0.1
* feat(benchmark): introduce benchmark runner engine, Docker containerization, and calibration datasets
* Fix URL/path scope exclusions and update duck-store template
- Modified targets.py to ensure that excluding a specific URL or path does not result in the entire host being blocked. Implemented a strict command payload check for excluded endpoints.
- Updated duck-store benchmark scope template to explicitly block the /vulnerabilities endpoint (which leaks intentional challenges) while allowing access to endpoint documentation.
* fix: correct scope.yaml indentation and refine guard blocking messages
* fix: resolve ModuleNotFoundError when score.py is executed directly
---------
Co-authored-by: Dan <dan@strategicautomation.local>