istosandClaude Opus 5 da8984d0e6 A phase runs itself, on a branch of its own
Starting a phase cuts phase/<stem> from the newest main it can see and
works the list into it: each member branched from the phase's tip, run
headless, merged back when its checks are green, the next one started.
At the end one PR into main, for a human. The human gate moves from
every card to the phase boundary, and the promise survives: the board
merges into a branch it created, inside a scope you opened.

The runner is a beat, not an agent — everything it decides is already
structured state, and an agent paid to poll would be the wrong tool at
the wrong price. It holds no registry of where a phase is. Two durable
things carry the memory, and the board already writes both: git, where
a member is finished when its branch is contained in the phase branch,
and the card, which grows a ## Phase log the runner adds one line to
per decision. The log is what tells "this member has run and it ended
badly" from "the phase has not reached it yet" — without it a
restarted board would relaunch a run that died.

Containment alone is not enough to call a member merged: a clean exit
that committed nothing leaves an empty branch that is contained. The
card has to have settled into review/ too, or a broken launch would
hide exactly where it always tries to.

Five conditions halt, each already a visible state on the card, and a
halt is written once and then held. Running the phase again is the
person's decision and is what clears it — the run is scoped to its own
log line, so a member whose run died is launchable again. A dependency
that has not landed is a wait, not a halt.

Merges are additive throughout: main into the phase branch on every
beat so a long run does not drift into one enormous conflict, members
into it as they go green, nothing rebased and nothing force-pushed. A
conflict aborts, leaves the branch as it was, and halts naming the
files that collided.

The actor rule decides who runs it, written where it already lives:
the phase card's assignee. A replica renders the phase and advances
nothing. Reachable through /api/phase/run and the ticker; the header
chip and the card actions are a separate card.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 09:43:57 +02:00
2026-08-01 09:12:34 +02:00
2026-07-31 15:09:02 +02:00
2026-07-30 07:40:17 +02:00

Bench

A live kanban for coding-agent work: task files in stage directories are the only source of truth; a stdlib-only board narrates everything that happens to them — agents working in git worktrees, PRs opening on review, CI and Copilot state on the cards, drives of the app from a task's own branch, and an archive that is never a delete. Turn on team mode (BOARD_SYNC=1) and the truth is origin/main: moves commit and push themselves, every board pulls on a beat, and the person who claimed a card first keeps it.

Docs: bench.12vectors.com — guides, concepts, and a settings reference, every page of it generated from this repository's own markdown, so what the docs say and what bench does cannot drift apart.

The bench board: cards in backlog, to-do, in-progress, review and done, three agents working, each card showing the branch it is on and the command its agent is running.

bench's own board, running on this repository. Three agents at work, each on its own branch — and every one of them stops at review.

Install into a repo

mkdir .task-manager && curl -L \
  https://github.com/12vectors/bench/releases/latest/download/bench.tar.gz \
  | tar -xz -C .task-manager
./.task-manager/start.sh        # wires the project (idempotent) and serves

No token, no clone: releases are curated artifacts that never contained bench's own cards or settings, so the board starts empty by construction.

The first run asks the two questions it cannot answer for you — solo or team, and which agent adapter — and writes manager/local/.env from the documented example, so every other setting is discoverable in your own copy. What runs your tests is read off the project rather than asked (package.jsonnpm test, and so on); nothing recognisable leaves BOARD_AGENT_COMMANDS empty, and that is the key to set before an agent can run them. Bare Enter takes the default throughout; with no terminal (CI, a script) it asks nothing and install.py --setup asks later.

Commit .task-manager/ into the host repo — core is vendored on purpose, so clones work offline and updates show up in the host's own diffs.

The workflow brief ships as .task-manager/AGENTS.md — the cross-vendor name coding agents read natively — with CLAUDE.md beside it as a one-line compatibility pointer. Both live inside .task-manager/, so a host repo's own root AGENTS.md is never touched.

Update

./.task-manager/update.sh              # latest release
BENCH_REF=v2 ./.task-manager/update.sh # an exact release tag

The artifact is stamped with the repo it was built from, so updating needs no configuration; BENCH_SOURCE=<owner/repo> in manager/local/.env overrides the stamp. Updates replace manager/core/ and the top-level scripts wholesale and touch nothing else — tasks, plans, reference, and everything under manager/local/ (your driver, commands, prompt overrides, settings, state) survive every update. Then python3 .task-manager/install.py and restart the board. If the source repo has no published release yet, update.sh says so and changes nothing.

Working on bench itself

git clone git@github.com:12vectors/bench.git && cd bench && ./start.sh

A clone carries bench's own cards and local/ content — that is dev mode, not an install. (Installing from a clone anyway works: install.py clears the shipped cards on its first boot in a host repo.) Releases are built by ./release.sh from the manifest at manager/core/release-manifest: tag = v<VERSION>, one stable asset name (bench.tar.gz), contents at the tarball root — the two things the install one-liner above depends on.

The three-layer law

Core knows about tasks, worktrees, PRs and events. It knows nothing about any particular app (drivers do: manager/local/driver/start), agent vendor (adapters do: manager/core/adapters/), or project (manager/local/ does).

The consequence is what makes an update safe: update.sh replaces manager/core/ wholesale, and everything a project taught bench about itself lives outside it. Full docs in AGENTS.md — the file the docs site is cut from, and the one an agent working in your repo reads; the adapter contract in manager/core/adapters/README.md.

License

MIT.

S
Description
Bench is a local task manager that integrates with coding agents to make it easier to manage work in a repo.
Readme MIT
1.7 MiB
Languages
Python 78.2%
HTML 17.5%
CSS 2.1%
Shell 2%
JavaScript 0.2%