The path fix has no regression guard: nothing in CI resolves references/
links, so all 18 were broken while CI stayed green. validate-artifact-
paths.js is scoped to spec/plan/todo artifacts and says in its own header
that it is not a general markdown path linter.
Add validate-reference-links.js, which resolves every `references/*.md`
link in skills/*/SKILL.md against that skill's own directory. This accepts
both conventions in CLAUDE.md: shared checklists reached via
../../references/, and a skill's own colocated references/ directory.
Scope stays narrow on purpose. Skills legitimately name paths that do not
exist yet -- tasks/todo.md, PERF.md, docs/ideas/[idea-name].md -- and a
general markdown linter would fail the build on them. A test pins that.
Proven against the pre-fix tree: 18 error(s), exit 1, matching the 18
links fixed in the previous commit. After the fix: 0 error(s), exit 0.
7 unit tests cover the regression itself, colocated references/,
markdown-link syntax, a renamed target, multiple violations in one skill,
and the non-reference paths that must be ignored. Wired into the
validate-skills job, alongside the other skill-content checks.
Known limitation: fenced code blocks are not stripped, so a SKILL.md that
documents the anti-pattern inside a fence would be flagged. Nothing does
today. Sharing stripFencedCodeBlocks looks right once #444 lands.
Refs #468
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The /spec and /plan commands write spec, plan, and todo artifacts to paths that /build and the spec/plan skills read back. When a producer moves an artifact without updating the consumers, the pipeline breaks and nothing in CI catches it: the command-parity check only compares descriptions, not paths. PR #93 hit exactly this, pointing /spec and /plan at docs/features/[name]/ while /build still required SPEC.md and tasks/plan.md.
validate-artifact-paths.js enforces one canonical set of artifact paths across every file in the pipeline (the spec/plan/build commands, the spec-driven-development and planning-and-task-breakdown skills, and the getting-started and adoption guides). Any spec/plan/todo artifact path outside the allowlist fails CI, so changing the convention has to touch the allowlist and every guarded file in the same change.
Scope is deliberately narrow: only spec/plan/todo artifacts, only the pipeline files; it is not a general markdown path linter. Wired into the validate-commands CI job alongside its test.
materializeWorkspace() creates directories in os.tmpdir() but
runBehavioral() never cleans them up on success or failure. If
execFileSync throws (timeout, signal, OOM), the workspace persists
in /tmp indefinitely with fixture data. The test file
run-evals-test.js uses try/finally cleanup (line 234), proving
the pattern was known but omitted in production code.
Fix: Wrap eval execution in try/finally with best-effort
fs.rmSync cleanup.
The skillName argument from --behavioral has zero validation and is
used directly in path.join() for reading .json files, reading
SKILL.md, and WRITING .grading.json results. A value like
'../../etc/passwd' traverses outside the project directory for
both reads and writes.
Fix: Validate skillName against /^[a-z0-9]+(-[a-z0-9]+)*$/ before
any filesystem operations. This matches the existing KEBAB_CASE
validation in skill-lint.js.
Two heuristic bugs in the section and trigger checks, surfaced in
addyosmani#387 and triaged by @addyosmani and @nucliweb:
Section check: content.includes('## Overview') matches headings inside
fenced code blocks and sub-headings (### Overview) because substring
matching can't distinguish prose from examples. Replace with a
line-anchored regex match against content that has had fenced code
blocks stripped, so only real top-level headings satisfy the check.
Trigger check: 'Do not use when testing' satisfied the DESCRIPTION_TRIGGER
regex because it contains 'use … when'. Add a negation pattern and reject
descriptions where the only trigger match is negated.
Verified: all 24 skills pass (output unchanged — no current skill has
either bug pattern). Crafted edge cases confirm the fixes catch both
issues. Eval suite green, rank-1 rate unchanged at 86%.
Co-authored-by: Cursor <cursoragent@cursor.com>
Per review: the policy collections (REQUIRED_SECTIONS, SECTION_EXEMPT_SKILLS,
SKILL_REF_PATTERNS, regexes) were exported by reference, so a consumer could
mutate shared state and change lint results process-wide. Export only the
linting functions; keep the policy collections private. No behavior change
(validator output byte-identical).
Split the skill validation rules out of the validate-skills.js CLI into a
shared, importable scripts/lib/skill-lint.js: a pure lintSkillContent() with
no I/O plus a thin lintSkill() file wrapper. validate-skills.js becomes a thin
CLI over the lib. No behavior change to the validator's output or exit codes;
this only makes the rules unit-testable for a follow-up test battery.
Two latent Tier-3 bugs caught in review before the path was ever exercised:
- The grader prompt embeds the full stream-json trace (up to megabytes) and
was passed as an argv entry, which would fail with E2BIG on any real run.
It now goes to `claude -p` over stdin; the executor prompt moves to stdin
for the same reason.
- The executor ran headless with no permission mode, so file edits and
command runs could be denied, forcing the narrate-instead-of-perform
failure mode that trace grading exists to catch. It now runs with
--permission-mode acceptEdits and a pre-approved tool list
(Read,Glob,Grep,Edit,Write,Bash), documented in the README.
Per review from @federicobartoli and @nucliweb on #342:
Tier 3 (behavioral):
- Grade the execution trace, not the final output: executor runs with
--output-format stream-json --verbose so the grader judges tool calls
and file edits rather than the model's self-reporting.
- Run each eval in a throwaway workspace; files[] fixtures materialize
from evals/fixtures/ so evals can operate on real code.
- Node-level timeouts on executor and grader calls; grader output parsed
and shape-validated before writing (raw saved on failure); the trace is
fenced as untrusted data in the grader prompt.
- All 24 behavioral evals flagged trust_level: "provisional" until they
gain fixtures; the runner surfaces this and exits nonzero on failed
expectations.
Tier 2 (deterministic):
- Negative triggers accept an "owner" skill that must outrank this one,
turning them into pairwise routing tests that cannot pass vacuously;
37 of 48 negatives now declare owners (the rest are tracked in #351).
- Warn when a case file is below the documented minimums (3 positive /
2 negative / 1 behavioral); promotion to error tracked in #352.
- Stemmer: cluster trailing y/i ("simplify"/"simplifies").
Baseline holds: 120 checks, 0 errors, 85% trigger rank-1 rate.
There was no way to measure whether skills trigger correctly, stay
distinct, or change agent behavior. This adds evals, aligned with what
the community has converged on, with a deterministic CI tier on top:
- evals/cases/<skill>.json for all 24 skills. The evals[] block uses
Anthropic skill-creator's evals.json schema verbatim (id, prompt,
expected_output, expectations[]) so its runner, benchmarks, and eval
viewer work against our files unmodified. A trigger block (this
repo's extension) adds positive/negative routing prompts per skill.
- scripts/run-evals.js, zero-dependency runner:
Tier 2 (CI): trigger evals via stemmed TF-IDF ranking over skill
descriptions (positive prompts must rank top-k, negative prompts
must not rank first), catalog collision detection between skill
descriptions, schema and coverage checks.
Tier 3 (opt-in): --behavioral <skill> executes each eval through
headless claude -p and grades the transcript against expectations[]
(superpowers-style); --dry-run previews without spending tokens.
- CI: run the deterministic tier in the validate-skills job.
- Docs: evals/README.md defines the framework and prior art;
CONTRIBUTING requires an eval file for new skills (warning-level in
the runner until in-flight skill PRs clear); CLAUDE.md pointers.
Current baseline: 120 checks pass, 85% trigger rank-1 rate across 72
positive prompts, zero catalog collisions.
The validator claims to check skills "against the rules in
docs/skill-anatomy.md", but two rules that doc marks as Required were
never enforced:
- Directory names must be lowercase-hyphen-separated (Naming Conventions).
Previously only `name === dirName` was checked, so `My_Skill/` passed.
- Descriptions must say *when* to use the skill, not just what it does
(Required vs Recommended). Formalized in #167 but the validator was
never updated to match.
Adds both as blocking checks. All 24 existing skills pass, so this is a
non-breaking guardrail. Closes the remaining gap in #233 (frontmatter,
name-match, and section checks already shipped in the original validator).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Support single-quoted TOML description strings (previously silently returned null)
- Wrap file reads in try/catch to produce actionable CI errors instead of raw stack traces
- Remove dead tomlStems variable; declare allTomlStems once and compute allCanonicalStems for an accurate command count in the summary
- Simplify description mismatch output to always show all three values, removing a redundant conditional branch
- scripts/validate-skills.js: zero-dependency Node.js validator that
checks every skill for valid frontmatter, description length (≤1024),
required sections (Overview, When to Use, Common Rationalizations,
Red Flags, Verification), and dead cross-skill references
- Skills with type:meta or exempt:sections in frontmatter skip section
checks; applied to using-agent-skills (meta) and idea-refine (legacy
structure predating the anatomy spec)
- CI: validate-skills job runs before plugin-manifest validation and
blocks merge on any error; uses Node 20, no npm install required