Phases are the release, so AGENTS.md stops burying them. The six sections that were `###`s inside "Task file format" — where they had split that section's own prose in half, header table above and Assignee below — become a `## Phases` section of their own, between "Pull requests" and "Stages" because that is the order they are learned in. Nothing in them changed: the diff is a move plus the section's opening paragraph. Which gives the site three pages to cut rather than a tail nobody would find at the bottom of "Task files": /concepts/phases/ (the model), /concepts/running-a-phase/ (the branch, the beat, the halt, the ending) and /concepts/phases-on-the-board/ (what the Board stops drawing, and the Phases view that draws it instead). They sit after "PRs and review" in the Concepts flow, which is the order AGENTS.md now reads in — the manifest and the source cannot disagree about that without the build saying so. The landing page's third reason says it too, since "parallel work, zero collisions" was only half of what the board now does, and the README's opening paragraph gains the sentence it was missing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
4.6 KiB
Bench
A live kanban for coding-agent work: task files in stage directories are
the only source of truth; a stdlib-only board narrates everything that
happens to them — agents working in git worktrees, PRs opening on review,
CI and Copilot state on the cards, drives of the app from a task's own
branch, and an archive that is never a delete. Work that is a run of
related cards rather than one is a phase: a card that lists its
cards, an integration branch of its own, and members branched from it,
run one at a time and merged back on green — ending in a single PR into
main for a human. Turn on team mode
(BOARD_SYNC=1) and the truth is origin/main: moves commit and push
themselves, every board pulls on a beat, and the person who claimed a
card first keeps it.
Docs: bench.12vectors.com — guides, concepts, and a settings reference, every page of it generated from this repository's own markdown, so what the docs say and what bench does cannot drift apart.
bench's own board, running on this repository. Three agents at work, each on its own branch — and every one of them stops at review.
Install into a repo
mkdir .task-manager && curl -L \
https://github.com/12vectors/bench/releases/latest/download/bench.tar.gz \
| tar -xz -C .task-manager
./.task-manager/start.sh # wires the project (idempotent) and serves
No token, no clone: releases are curated artifacts that never contained bench's own cards or settings, so the board starts empty by construction.
The first run asks the two questions it cannot answer for you — solo or
team, and which agent adapter — and writes manager/local/.env from the
documented example, so every other setting is discoverable in your own
copy. What runs your tests is read off the project rather than asked
(package.json → npm test, and so on); nothing recognisable leaves
BOARD_AGENT_COMMANDS empty, and that is the key to set before an agent
can run them. Bare Enter takes the default
throughout; with no terminal (CI, a script) it asks nothing and
install.py --setup asks later.
Commit .task-manager/ into the host repo — core is vendored on purpose,
so clones work offline and updates show up in the host's own diffs.
The workflow brief ships as .task-manager/AGENTS.md — the cross-vendor
name coding agents read natively — with CLAUDE.md beside it as a one-line
compatibility pointer. Both live inside .task-manager/, so a host repo's
own root AGENTS.md is never touched.
Update
./.task-manager/update.sh # latest release
BENCH_REF=v2 ./.task-manager/update.sh # an exact release tag
The artifact is stamped with the repo it was built from, so updating
needs no configuration; BENCH_SOURCE=<owner/repo> in
manager/local/.env overrides the stamp. Updates replace
manager/core/ and the top-level scripts wholesale and touch nothing
else — tasks, plans, reference, and everything under manager/local/
(your driver, commands, prompt overrides, settings, state) survive every
update. Then python3 .task-manager/install.py and restart the board.
If the source repo has no published release yet, update.sh says so and
changes nothing.
Working on bench itself
git clone git@github.com:12vectors/bench.git && cd bench && ./start.sh
A clone carries bench's own cards and local/ content — that is dev mode,
not an install. (Installing from a clone anyway works: install.py
clears the shipped cards on its first boot in a host repo.) Releases are
built by ./release.sh from the manifest at
manager/core/release-manifest: tag = v<VERSION>, one stable asset
name (bench.tar.gz), contents at the tarball root — the two things the
install one-liner above depends on.
The three-layer law
Core knows about tasks, worktrees, PRs and events. It knows nothing about
any particular app (drivers do: manager/local/driver/start), agent
vendor (adapters do: manager/core/adapters/), or project
(manager/local/ does).
The consequence is what makes an update safe: update.sh replaces
manager/core/ wholesale, and everything a project taught bench about
itself lives outside it. Full docs in AGENTS.md — the file
the docs site is cut from, and the one an agent working in your repo
reads; the adapter contract in
manager/core/adapters/README.md.
License
MIT.
