347 Commits
Author SHA1 Message Date
Addy OsmaniandGitHub 98967c45a4 Merge pull request #396 from nucliweb/docs/adoption-guide
docs: add adoption guide for greenfield vs brownfield rollout
0.6.4
2026-07-12 10:58:04 -07:00
Addy OsmaniandGitHub 3517e07228 Merge pull request #397 from nucliweb/docs/developer-onboarding
docs: add developer onboarding guide for repo contributors
2026-07-12 10:54:26 -07:00
Joan Leon 7209bd12d8 docs(developer-onboarding): point to CONTRIBUTING for new-skill rules
Per review: the new-skill 'ships as a set' block restated CONTRIBUTING's
frontmatter, anatomy, and eval-count rules instead of pointing at them,
which is the 'don't duplicate, reference' rule this guide champions. Keep
the shape (SKILL.md + eval case + optional scripts) and defer the exact
requirements to CONTRIBUTING.md and skill-anatomy.md so they can't drift.
2026-07-12 19:48:34 +02:00
Joan Leon 27dcc32b83 docs(adoption-guide): name the merge-blocking review labels
Per review: the severity-label taxonomy's point is the labels that gate a
merge (Critical and Required), with Nit/Optional/FYI being the don't-block
side. In a doc about protecting a legacy codebase, name the blocking half
instead of only the optional one; let the skill remain the source of the
full taxonomy.
2026-07-12 19:47:07 +02:00
Joan Leon 695505ddac docs: add developer onboarding guide for repo contributors
Add docs/developer-onboarding.md: a guided tour for people working on the
repo itself (the five layers, local setup, the verification loop, the
contribution paths, and a suggested reading order), complementing the
authoritative rules in CONTRIBUTING.md, skill-anatomy.md, and evals/README.md.
Link it from the top of CONTRIBUTING.md as the map to its rulebook.
2026-07-12 15:50:11 +02:00
Joan Leon ffd7db67e3 docs: add adoption guide (greenfield vs brownfield rollout)
Add docs/adoption-guide.md covering two rollout paths: full lifecycle
from day one for a greenfield project, and an incremental,
verification-first path for an established codebase. Link it from the
README (new Adoption section) and from getting-started.md's Recommended
Setup as the in-depth companion to the quick setup.
2026-07-12 14:16:49 +02:00
Addy OsmaniandGitHub 849850e039 Merge pull request #392 from federicobartoli/docs/package-manager-supply-chain-hardening
docs(security): harden package-manager supply-chain guidance
2026-07-11 10:52:07 -07:00
Federico Bartoli 509281da06 docs(security): clarify npm install-script versions
Date the version matrix and distinguish fallback, npm 11.18.x, and npm 12.x behavior using verified lifecycle-script probes.
2026-07-11 14:44:03 +02:00
Federico Bartoli 175013ab41 docs(security): harden package-manager supply-chain guidance 2026-07-11 09:34:05 +02:00
Addy Osmani 4e8bd9fde4 docs(readme): add Team section 2026-07-10 16:42:04 -07:00
Addy Osmani 7ce442de03 docs(comparison): refresh and expand the three-way comparison
Bring docs/comparison.md up to date and make it more useful for people
choosing between the packs:

- agent-skills: add the three-tier eval framework as the current point of
  difference, plus current tooling (Codex, Kiro, the npx skills CLI),
  /build auto, and the 24-skill / 7-checklist / Definition-of-Done facts.
- Superpowers: correct to ~14 inner-loop skills, the consolidated single
  task reviewer, the worst-case-executor plan standard, and its main
  unmet ask (agent teams); drop the stale Gemini reference.
- Matt Pocock's skills: reframe around the grilling primitive and the
  grown ~30-skill toolkit (in-progress/deprecated dirs, wayfinder,
  seam-based TDD), not a "tight set".
- Add a much fuller "How to decide what to use" section after the table:
  by shape of work, by what you optimize for, concrete scenarios, solo
  vs team, and an honest shared-frontier note on cross-session memory.

Keeps the fair-not-flattering stance and the Om Mishra head-to-head.
2026-07-10 10:01:07 -07:00
Addy OsmaniandGitHub 6bcfeb9dae Merge pull request #218 from nguyenducthaonguyen/docs/update-cursor-setup
docs: update Cursor setup guide for rules and skills layout
2026-07-09 19:59:20 -07:00
Addy OsmaniandGitHub 3780895567 Merge pull request #368 from mvanhorn/fix/249-docs-document-https-plugin-install-workaround
docs: document the HTTPS git-config workaround for the /plugin install SSH clone failure
2026-07-09 19:27:53 -07:00
Addy OsmaniandGitHub e6bab792a5 Merge pull request #372 from ShiroKSH/fix/eval-validation-windows
fix: harden eval and command validation
2026-07-09 19:27:43 -07:00
Addy OsmaniandGitHub 4848a9cfc2 Merge pull request #357 from orbisai0security/fix-v-002-prevent-hardcoded-credentials-webperf
fix: references to api key usage in configuration fi... in webperf.toml
2026-07-09 19:27:34 -07:00
Addy OsmaniandGitHub dfb9bf2f33 Merge pull request #346 from HMAKT99/feat/dependency-upgrades
docs(code-review): add dependency upgrade workflow to Dependency Discipline
2026-07-09 19:27:25 -07:00
Addy OsmaniandGitHub 8ddfe89eb8 Merge pull request #345 from HMAKT99/feat/db-schema-migrations
docs(deprecation): add database schema migration patterns (expand/contract)
2026-07-09 19:27:16 -07:00
Addy OsmaniandGitHub 2373dd6284 Merge pull request #358 from ZhiyaoWen999/codex/description-vocabulary-gaps
fix(evals): cover description vocabulary gaps
2026-07-09 19:27:06 -07:00
Addy OsmaniandGitHub 8d74a4d604 Merge pull request #375 from nucliweb/docs/no-translations-policy
docs: decline documentation and skill translations
2026-07-09 12:49:37 -07:00
Joan Leon 9ad0fbd0ef docs: decline documentation and skill translations
State in CONTRIBUTING.md that translations of docs and skills are not
accepted: translated copies drift as content evolves and can't be
maintained long-term without leaning on agent translations plus
community corrections, for limited value.
2026-07-09 21:46:56 +02:00
Joan LeónandGitHub d6da74760e Merge pull request #374 from nucliweb/docs/scope-note-repo-agent-files
docs: clarify AGENTS.md and CLAUDE.md are repo-scoped, not for user projects
2026-07-09 21:36:49 +02:00
Joan Leon b665c5c596 docs: clarify AGENTS.md and CLAUDE.md are repo-scoped
Add a Scope banner to AGENTS.md and CLAUDE.md stating they configure
agents working on the addyosmani/agent-skills repository itself, not
users' own projects, referencing the repo by its canonical GitHub URL
so the scope is unambiguous when the file is read out of context. Add a
matching "Repo-scoped files" note to CONTRIBUTING.md so setup-guide
authors don't instruct users to copy these files.
2026-07-09 21:22:08 +02:00
Addy OsmaniandGitHub edaa28a229 Merge pull request #359 from debs-obrien/improve-playwright-role-locators
Improve Playwright locator examples
2026-07-09 10:23:44 -07:00
ShiroKSH eb5acf1d46 chore: normalize text line endings 2026-07-09 19:48:28 +03:00
ShiroKSH df8eda8fbe fix: harden eval and command validation 2026-07-09 18:03:46 +03:00
Matt Van Horn 3f08584f6e docs: document the HTTPS git-config workaround for the /plugin install 2026-07-08 22:16:48 -07:00
Addy OsmaniandGitHub 0e63d8ea96 Merge pull request #88 from federicobartoli/feat/codex-plugin-support
feat: add Codex plugin support without duplicating skills
2026-07-08 15:27:36 -07:00
Federico Bartoli 9c0edb07df fix: prevent Codex from loading Claude hooks 2026-07-09 00:03:10 +02:00
Federico Bartoli 97522bb69b fix: use root Codex plugin layout 2026-07-08 15:57:32 +02:00
Federico Bartoli 41f6a03a59 docs(codex): update install command for Codex CLI v0.122+
In Codex CLI 0.122 the marketplace subcommand moved from
`codex marketplace` to `codex plugin marketplace`. Update README and
docs/codex-setup.md so install snippets work on current Codex (verified
on 0.128.0). Keep one historical reference in the v0.122 callout.
2026-07-08 15:39:36 +02:00
Federico Bartoli 217a7e723d feat: add codex quick start 2026-07-08 15:39:36 +02:00
Federico Bartoli 1afa9ef325 feat: add codex cli reference in codex-setup 2026-07-08 15:39:36 +02:00
Federico Bartoli 3a76ead2ef feat: add Codex plugin support without duplicating skills
Register the repo as a Codex plugin so `codex marketplace add addyosmani/agent-skills`
installs it in a single step. The plugin reads the existing `skills/` directory via
a symlink — no files are copied, and `skills/<name>/SKILL.md` remains the single
source of truth shared with Claude Code.

- codex/.codex-plugin/plugin.json — Codex manifest, skills: "./skills/"
- codex/skills → ../skills — symlink so the plugin dir stays self-contained
  while git tracks a single canonical skills/ at the repo root
- .agents/plugins/marketplace.json — local marketplace entry, plugin at ./codex
- docs/codex-setup.md — install and usage guide

Verified end-to-end with codex-cli 0.121.0: skills appear in the plugin's
Skills list after install.
2026-07-08 15:39:36 +02:00
Debbie O'Brien 6ace60a354 Improve Playwright locator examples 2026-07-07 10:21:00 +02:00
Zhiyao 13ca1dea5c fix(evals): cover description vocabulary gaps 2026-07-07 14:57:16 +08:00
Addy OsmaniandGitHub 70b7506ce9 Merge pull request #342 from addyosmani/feat/skill-evals
feat(evals): three-tier skill eval framework (trigger, routing, behavioral)
2026-07-06 17:59:41 -07:00
orbisai0security 6045554bce fix: V-002 security vulnerability
Automated security fix generated by OrbisAI Security
2026-07-07 00:34:38 +00:00
Addy Osmani 076af17226 docs(readme): add skills CLI install path with per-skill examples
The pack has been installable via `npx skills add` all along (all 24
skills are indexed on skills.sh) but the README never mentioned it. Add
it as the fastest cross-agent path at the top of Quick Start, with
individual install examples for three signature skills.
2026-07-06 13:21:46 -07:00
Addy Osmani dd142f1da3 fix(evals): pipe grader prompt via stdin; grant executor tool permissions
Two latent Tier-3 bugs caught in review before the path was ever exercised:

- The grader prompt embeds the full stream-json trace (up to megabytes) and
  was passed as an argv entry, which would fail with E2BIG on any real run.
  It now goes to `claude -p` over stdin; the executor prompt moves to stdin
  for the same reason.
- The executor ran headless with no permission mode, so file edits and
  command runs could be denied, forcing the narrate-instead-of-perform
  failure mode that trace grading exists to catch. It now runs with
  --permission-mode acceptEdits and a pre-approved tool list
  (Read,Glob,Grep,Edit,Write,Bash), documented in the README.
2026-07-06 12:38:26 -07:00
Addy Osmani fe13251e20 docs(evals): document trace grading, owners, trust levels, and follow-ups
Wire the framework docs to the tracking issues: #351 (description
vocabulary gaps) and #352 (Tier 3 graduation + deterministic ratchets),
so warning-level checks have an explicit promotion path instead of
becoming permanent.
2026-07-06 11:36:29 -07:00
Addy Osmani 194b2099c6 feat(evals): harden Tier 3 and make negatives pairwise routing tests
Per review from @federicobartoli and @nucliweb on #342:

Tier 3 (behavioral):
- Grade the execution trace, not the final output: executor runs with
  --output-format stream-json --verbose so the grader judges tool calls
  and file edits rather than the model's self-reporting.
- Run each eval in a throwaway workspace; files[] fixtures materialize
  from evals/fixtures/ so evals can operate on real code.
- Node-level timeouts on executor and grader calls; grader output parsed
  and shape-validated before writing (raw saved on failure); the trace is
  fenced as untrusted data in the grader prompt.
- All 24 behavioral evals flagged trust_level: "provisional" until they
  gain fixtures; the runner surfaces this and exits nonzero on failed
  expectations.

Tier 2 (deterministic):
- Negative triggers accept an "owner" skill that must outrank this one,
  turning them into pairwise routing tests that cannot pass vacuously;
  37 of 48 negatives now declare owners (the rest are tracked in #351).
- Warn when a case file is below the documented minimums (3 positive /
  2 negative / 1 behavioral); promotion to error tracked in #352.
- Stemmer: cluster trailing y/i ("simplify"/"simplifies").

Baseline holds: 120 checks, 0 errors, 85% trigger rank-1 rate.
2026-07-06 11:36:29 -07:00
Arun Kumar Thiagarajan e270415226 docs(code-review): add dependency upgrade workflow to Dependency Discipline
Extends the existing adopt gate with the missing upgrade workflow: read the
changelog over the version number, one package per change, verify via tests,
review the lockfile/transitive diff. Cross-links security-and-hardening for
npm audit and supply-chain rather than duplicating it.
2026-07-05 19:30:20 +05:30
Arun Kumar Thiagarajan 5a4a69adfc docs(deprecation): add database schema migration patterns (expand/contract)
Adds a Database Schema Migrations section to deprecation-and-migration covering
expand/contract, dual-write + batched backfill, additive-first/destructive-last,
tested down paths, and non-blocking index builds. References incremental-implementation
for slicing; reuses the skill's existing Feature Flag Migration pattern.
2026-07-05 19:28:17 +05:30
Addy Osmani 45e1449138 feat(evals): add a three-tier skill eval framework
There was no way to measure whether skills trigger correctly, stay
distinct, or change agent behavior. This adds evals, aligned with what
the community has converged on, with a deterministic CI tier on top:

- evals/cases/<skill>.json for all 24 skills. The evals[] block uses
  Anthropic skill-creator's evals.json schema verbatim (id, prompt,
  expected_output, expectations[]) so its runner, benchmarks, and eval
  viewer work against our files unmodified. A trigger block (this
  repo's extension) adds positive/negative routing prompts per skill.
- scripts/run-evals.js, zero-dependency runner:
  Tier 2 (CI): trigger evals via stemmed TF-IDF ranking over skill
  descriptions (positive prompts must rank top-k, negative prompts
  must not rank first), catalog collision detection between skill
  descriptions, schema and coverage checks.
  Tier 3 (opt-in): --behavioral <skill> executes each eval through
  headless claude -p and grades the transcript against expectations[]
  (superpowers-style); --dry-run previews without spending tokens.
- CI: run the deterministic tier in the validate-skills job.
- Docs: evals/README.md defines the framework and prior art;
  CONTRIBUTING requires an eval file for new skills (warning-level in
  the runner until in-flight skill PRs clear); CLAUDE.md pointers.

Current baseline: 120 checks pass, 85% trigger rank-1 rate across 72
positive prompts, zero catalog collisions.
2026-07-03 23:46:30 -07:00
Addy OsmaniandGitHub 8c65303053 Merge pull request #334 from HMAKT99/feat/release-versioning
docs(git-workflow): add release & versioning (semver, tags, changelogs)
0.6.3
2026-07-02 15:05:04 -07:00
Addy OsmaniandGitHub eae843fcde Merge pull request #337 from hiyochi/fix/plan-output-path-in-skill
fix: add tasks/plan.md and tasks/todo.md output paths to planning skills
2026-07-02 15:04:51 -07:00
hiyochi 7d36add8cf fix: add tasks/plan.md and tasks/todo.md output paths to planning skills
The /plan command specifies saving the plan to tasks/plan.md and
tasks/todo.md, but the planning-and-task-breakdown skill (the canonical
source) had no file path instructions. When spec-driven-development
transitions to Plan phase, it references the skill directly, bypassing
the /plan command, causing plans to be written to the wrong location.

- Add Output Files section to planning-and-task-breakdown with explicit
  tasks/plan.md and tasks/todo.md paths
- Add path instruction to Step 1 (Enter Plan Mode)
- Add Output convention note to spec-driven-development Phase 2
2026-06-30 05:16:40 +08:00
Addy OsmaniandGitHub aba7c4e969 Merge pull request #323 from An-idd/feat/validate-naming-and-trigger
feat(scripts): enforce naming + description-trigger rules in skill validator
2026-06-28 11:11:20 -07:00
Addy OsmaniandGitHub b01a145389 Merge pull request #325 from nucliweb/docs/pr-overlap-guideline
docs: add PR-overlap guideline to CLAUDE.md
2026-06-28 11:08:05 -07:00
Addy OsmaniandGitHub 30e55cb060 Merge pull request #313 from nucliweb/docs/skill-contributing-guardrails
docs: route new-skill work through CONTRIBUTING pre-flight checks
2026-06-28 02:51:20 -07:00