Files
Merlin's CatGitHubCopilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>MerlinHCopilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
9fa9ac25a4 feat: add workflow eval framework (#32)
* feat: add workflow eval framework

Add workflow evaluation scenarios, rubrics, schemas, and runner scripts for installed Truthmark workflows.

Move research notes under docs/research and migrate tests from Vitest to node:test.

* Potential fix for pull request finding 'CodeQL / Replacement of a substring with itself'

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>

---------

Co-authored-by: MerlinH <merlinh221@gmail.com>
Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
2026-06-30 20:57:55 +10:00

1.3 KiB

Failure Taxonomy

Use stable labels when reviewing workflow eval failures.

  • skill-not-triggered: the expected workflow skill or prompt was not used.
  • wrong-workflow: the agent selected a different Truthmark workflow.
  • over-triggered-workflow: the agent ran a write workflow for a read-only or general task.
  • missing-route-read: the agent did not inspect required route/config evidence.
  • wrong-route-owner: the agent updated or cited the wrong bounded truth owner.
  • bootstrap-route-misuse: the agent treated a bootstrap catch-all as behavior truth.
  • forbidden-code-write: a documentation workflow changed functional code.
  • forbidden-doc-write: a code-realization workflow changed truth docs or routes.
  • missing-evidence: a changed truth claim lacks checkout evidence.
  • unsupported-claim: the result preserves or adds a claim without evidence.
  • stale-truth-followed: Realize followed stale truth instead of blocking.
  • report-invalid: the workflow report is absent or fails validation.
  • verification-skipped-without-rationale: required checks were neither run nor explained.
  • token-bloat: the agent loaded or repeated unnecessary context.
  • agent-runner-failed: the configured runner command failed before behavior could be evaluated.
  • judge-not-evaluable: judge output was missing, malformed, or insufficient.