istosandClaude Opus 5 d25221bb74 The drawer reads a wrapped list as one item
md() split a block into physical lines and made each one a unit. Task
files are hard-wrapped at ~74 columns, so the second line of an item
became its own bullet, `- [ ]` rendered as a literal bracket pair,
nested lists flattened, and prose kept the author's ragged edge as
<br>. "Enough for these task files" was exactly what it was not.

Lists are now grouped into logical items before rendering: a new item
begins only at a marker, and a line without one is continuation text
joined with a space. Indentation is honoured — a marker past its level
opens a nested list, a shallower one closes back to the level that
fits — and one entry point serves both bullets and ordered lists, so
an <ol> nests under a <ul> the same way. Task-list items render as a
glyph in a span, never an <input>: the file is the source of truth and
the drawer is not an editor. A ticked box reads as settled (--calm);
an open one stays neutral. Paragraphs and blockquotes join their
source lines with a space, so prose reflows to the drawer's width.

Fences, tables, headings and rules are untouched, including the fence
state machine that spans blocks.

The tests lift esc() and md() out of the page and run them under node,
because the renderer is a pure function and its output is what to
assert on; node is not a bench dependency, so those checks skip when
it is absent and source-level invariants cover the shape of the fix.
One check renders every card on the board plus AGENTS.md and asserts
one bullet per source marker — the acceptance criterion applied to the
whole corpus.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 15:01:26 +02:00
2026-07-31 14:49:15 +02:00
2026-07-30 07:40:17 +02:00

Bench

A live kanban for coding-agent work: task files in stage directories are the only source of truth; a stdlib-only board narrates everything that happens to them — agents working in git worktrees, PRs opening on review, CI and Copilot state on the cards, drives of the app from a task's own branch, and an archive that is never a delete. Turn on team mode (BOARD_SYNC=1) and the truth is origin/main: moves commit and push themselves, every board pulls on a beat, and the person who claimed a card first keeps it.

Install into a repo

mkdir .task-manager && curl -L \
  https://github.com/12vectors/bench/releases/latest/download/bench.tar.gz \
  | tar -xz -C .task-manager
./.task-manager/start.sh        # wires the project (idempotent) and serves

No token, no clone: releases are curated artifacts that never contained bench's own cards or settings, so the board starts empty by construction.

The first run asks three questions it cannot answer for you — solo or team, which agent adapter, what command runs your tests — and writes manager/local/.env from the documented example, so every other setting is discoverable in your own copy. Bare Enter takes the default throughout; with no terminal (CI, a script) it asks nothing and install.py --setup asks later.

Commit .task-manager/ into the host repo — core is vendored on purpose, so clones work offline and updates show up in the host's own diffs.

The workflow brief ships as .task-manager/AGENTS.md — the cross-vendor name coding agents read natively — with CLAUDE.md beside it as a one-line compatibility pointer. Both live inside .task-manager/, so a host repo's own root AGENTS.md is never touched.

Update

./.task-manager/update.sh              # latest release
BENCH_REF=v2 ./.task-manager/update.sh # an exact release tag

The artifact is stamped with the repo it was built from, so updating needs no configuration; BENCH_SOURCE=<owner/repo> in manager/local/.env overrides the stamp. Updates replace manager/core/ and the top-level scripts wholesale and touch nothing else — tasks, plans, reference, and everything under manager/local/ (your driver, commands, prompt overrides, settings, state) survive every update. Then python3 .task-manager/install.py and restart the board. If the source repo has no published release yet, update.sh says so and changes nothing.

Working on bench itself

git clone git@github.com:12vectors/bench.git && cd bench && ./start.sh

A clone carries bench's own cards and local/ content — that is dev mode, not an install. (Installing from a clone anyway works: install.py clears the shipped cards on its first boot in a host repo.) Releases are built by ./release.sh from the manifest at manager/core/release-manifest: tag = v<VERSION>, one stable asset name (bench.tar.gz), contents at the tarball root — the two things the install one-liner above depends on.

The three-layer law

Core knows about tasks, worktrees, PRs and events. It knows nothing about any particular app (drivers do: manager/local/driver/start), agent vendor (adapters do: manager/core/adapters/), or project (manager/local/ does).

The consequence is what makes an update safe: update.sh replaces manager/core/ wholesale, and everything a project taught bench about itself lives outside it. Full docs in AGENTS.md; the adapter contract in manager/core/adapters/README.md.

License

MIT.

S
Description
Bench is a local task manager that integrates with coding agents to make it easier to manage work in a repo.
Readme MIT
1.7 MiB
Languages
Python 78.2%
HTML 17.5%
CSS 2.1%
Shell 2%
JavaScript 0.2%