Per review: the new-skill 'ships as a set' block restated CONTRIBUTING's
frontmatter, anatomy, and eval-count rules instead of pointing at them,
which is the 'don't duplicate, reference' rule this guide champions. Keep
the shape (SKILL.md + eval case + optional scripts) and defer the exact
requirements to CONTRIBUTING.md and skill-anatomy.md so they can't drift.
Per review: the severity-label taxonomy's point is the labels that gate a
merge (Critical and Required), with Nit/Optional/FYI being the don't-block
side. In a doc about protecting a legacy codebase, name the blocking half
instead of only the optional one; let the skill remain the source of the
full taxonomy.
Add docs/developer-onboarding.md: a guided tour for people working on the
repo itself (the five layers, local setup, the verification loop, the
contribution paths, and a suggested reading order), complementing the
authoritative rules in CONTRIBUTING.md, skill-anatomy.md, and evals/README.md.
Link it from the top of CONTRIBUTING.md as the map to its rulebook.
Add docs/adoption-guide.md covering two rollout paths: full lifecycle
from day one for a greenfield project, and an incremental,
verification-first path for an established codebase. Link it from the
README (new Adoption section) and from getting-started.md's Recommended
Setup as the in-depth companion to the quick setup.
Bring docs/comparison.md up to date and make it more useful for people
choosing between the packs:
- agent-skills: add the three-tier eval framework as the current point of
difference, plus current tooling (Codex, Kiro, the npx skills CLI),
/build auto, and the 24-skill / 7-checklist / Definition-of-Done facts.
- Superpowers: correct to ~14 inner-loop skills, the consolidated single
task reviewer, the worst-case-executor plan standard, and its main
unmet ask (agent teams); drop the stale Gemini reference.
- Matt Pocock's skills: reframe around the grilling primitive and the
grown ~30-skill toolkit (in-progress/deprecated dirs, wayfinder,
seam-based TDD), not a "tight set".
- Add a much fuller "How to decide what to use" section after the table:
by shape of work, by what you optimize for, concrete scenarios, solo
vs team, and an honest shared-frontier note on cross-session memory.
Keeps the fair-not-flattering stance and the Om Mishra head-to-head.
State in CONTRIBUTING.md that translations of docs and skills are not
accepted: translated copies drift as content evolves and can't be
maintained long-term without leaning on agent translations plus
community corrections, for limited value.
Add a Scope banner to AGENTS.md and CLAUDE.md stating they configure
agents working on the addyosmani/agent-skills repository itself, not
users' own projects, referencing the repo by its canonical GitHub URL
so the scope is unambiguous when the file is read out of context. Add a
matching "Repo-scoped files" note to CONTRIBUTING.md so setup-guide
authors don't instruct users to copy these files.
In Codex CLI 0.122 the marketplace subcommand moved from
`codex marketplace` to `codex plugin marketplace`. Update README and
docs/codex-setup.md so install snippets work on current Codex (verified
on 0.128.0). Keep one historical reference in the v0.122 callout.
Register the repo as a Codex plugin so `codex marketplace add addyosmani/agent-skills`
installs it in a single step. The plugin reads the existing `skills/` directory via
a symlink — no files are copied, and `skills/<name>/SKILL.md` remains the single
source of truth shared with Claude Code.
- codex/.codex-plugin/plugin.json — Codex manifest, skills: "./skills/"
- codex/skills → ../skills — symlink so the plugin dir stays self-contained
while git tracks a single canonical skills/ at the repo root
- .agents/plugins/marketplace.json — local marketplace entry, plugin at ./codex
- docs/codex-setup.md — install and usage guide
Verified end-to-end with codex-cli 0.121.0: skills appear in the plugin's
Skills list after install.
The pack has been installable via `npx skills add` all along (all 24
skills are indexed on skills.sh) but the README never mentioned it. Add
it as the fastest cross-agent path at the top of Quick Start, with
individual install examples for three signature skills.
Two latent Tier-3 bugs caught in review before the path was ever exercised:
- The grader prompt embeds the full stream-json trace (up to megabytes) and
was passed as an argv entry, which would fail with E2BIG on any real run.
It now goes to `claude -p` over stdin; the executor prompt moves to stdin
for the same reason.
- The executor ran headless with no permission mode, so file edits and
command runs could be denied, forcing the narrate-instead-of-perform
failure mode that trace grading exists to catch. It now runs with
--permission-mode acceptEdits and a pre-approved tool list
(Read,Glob,Grep,Edit,Write,Bash), documented in the README.
Wire the framework docs to the tracking issues: #351 (description
vocabulary gaps) and #352 (Tier 3 graduation + deterministic ratchets),
so warning-level checks have an explicit promotion path instead of
becoming permanent.
Per review from @federicobartoli and @nucliweb on #342:
Tier 3 (behavioral):
- Grade the execution trace, not the final output: executor runs with
--output-format stream-json --verbose so the grader judges tool calls
and file edits rather than the model's self-reporting.
- Run each eval in a throwaway workspace; files[] fixtures materialize
from evals/fixtures/ so evals can operate on real code.
- Node-level timeouts on executor and grader calls; grader output parsed
and shape-validated before writing (raw saved on failure); the trace is
fenced as untrusted data in the grader prompt.
- All 24 behavioral evals flagged trust_level: "provisional" until they
gain fixtures; the runner surfaces this and exits nonzero on failed
expectations.
Tier 2 (deterministic):
- Negative triggers accept an "owner" skill that must outrank this one,
turning them into pairwise routing tests that cannot pass vacuously;
37 of 48 negatives now declare owners (the rest are tracked in #351).
- Warn when a case file is below the documented minimums (3 positive /
2 negative / 1 behavioral); promotion to error tracked in #352.
- Stemmer: cluster trailing y/i ("simplify"/"simplifies").
Baseline holds: 120 checks, 0 errors, 85% trigger rank-1 rate.
Extends the existing adopt gate with the missing upgrade workflow: read the
changelog over the version number, one package per change, verify via tests,
review the lockfile/transitive diff. Cross-links security-and-hardening for
npm audit and supply-chain rather than duplicating it.
Adds a Database Schema Migrations section to deprecation-and-migration covering
expand/contract, dual-write + batched backfill, additive-first/destructive-last,
tested down paths, and non-blocking index builds. References incremental-implementation
for slicing; reuses the skill's existing Feature Flag Migration pattern.
There was no way to measure whether skills trigger correctly, stay
distinct, or change agent behavior. This adds evals, aligned with what
the community has converged on, with a deterministic CI tier on top:
- evals/cases/<skill>.json for all 24 skills. The evals[] block uses
Anthropic skill-creator's evals.json schema verbatim (id, prompt,
expected_output, expectations[]) so its runner, benchmarks, and eval
viewer work against our files unmodified. A trigger block (this
repo's extension) adds positive/negative routing prompts per skill.
- scripts/run-evals.js, zero-dependency runner:
Tier 2 (CI): trigger evals via stemmed TF-IDF ranking over skill
descriptions (positive prompts must rank top-k, negative prompts
must not rank first), catalog collision detection between skill
descriptions, schema and coverage checks.
Tier 3 (opt-in): --behavioral <skill> executes each eval through
headless claude -p and grades the transcript against expectations[]
(superpowers-style); --dry-run previews without spending tokens.
- CI: run the deterministic tier in the validate-skills job.
- Docs: evals/README.md defines the framework and prior art;
CONTRIBUTING requires an eval file for new skills (warning-level in
the runner until in-flight skill PRs clear); CLAUDE.md pointers.
Current baseline: 120 checks pass, 85% trigger rank-1 rate across 72
positive prompts, zero catalog collisions.
The /plan command specifies saving the plan to tasks/plan.md and
tasks/todo.md, but the planning-and-task-breakdown skill (the canonical
source) had no file path instructions. When spec-driven-development
transitions to Plan phase, it references the skill directly, bypassing
the /plan command, causing plans to be written to the wrong location.
- Add Output Files section to planning-and-task-breakdown with explicit
tasks/plan.md and tasks/todo.md paths
- Add path instruction to Step 1 (Enter Plan Mode)
- Add Output convention note to spec-driven-development Phase 2