Files
truthmark/docs/research/2026-06-29-manual-agent-skill-quality-eval-framework.md
Merlin's CatGitHubCopilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>MerlinHCopilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
9fa9ac25a4 feat: add workflow eval framework (#32)
* feat: add workflow eval framework

Add workflow evaluation scenarios, rubrics, schemas, and runner scripts for installed Truthmark workflows.

Move research notes under docs/research and migrate tests from Vitest to node:test.

* Potential fix for pull request finding 'CodeQL / Replacement of a substring with itself'

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>

---------

Co-authored-by: MerlinH <merlinh221@gmail.com>
Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
2026-06-30 20:57:55 +10:00

1.4 KiB

status, doc_type, last_reviewed, source_of_truth
status doc_type last_reviewed source_of_truth
draft research-index 2026-06-29
../../workflow-eval-framwork/README.md
../../workflow-eval-framwork/catalog.yaml
../../scripts/workflow-eval-framwork/run-agent-scenario.mjs
../../tests/evals/workflow-eval-framwork.test.ts

Manual Agent Skill And Prompt Quality Eval Framework

The implemented manual evaluation framework lives under workflow-eval-framwork/.

Use these artifacts:

The requested framework folder name is intentionally spelled workflow-eval-framwork.

The key invariant is that Truthmark's generated-surface tests prove injection and freshness, while this framework tests actual agent behavior when workflow skills and prompts are used.

The framework is manual and token-expensive. It must not become a default CI gate, hosted service, daemon, database, hidden memory layer, or required downstream runtime.