Commit Graph
5 Commits
Author SHA1 Message Date
Renn F cc4ccb7ea3 fix(usage): capture agent token usage from the Claude Code transcript
The token-usage pipeline was fully built — per-session SDK counters,
/usage/status, the orchestrator finalize-fetch that writes token columns and
estimated cost to the spawn-session row, the daily rollup, and the dashboard
— but nothing ever populated the counters. /usage/report had zero callers, so
every session reported zero tokens and the cost dashboard rendered all-zeros.
A redeploy could not fix code that was never written.

Close the loop with the producer that was missing. Claude Code does not pass
token counts to hooks, but it does pass the session transcript path, and each
assistant entry records its API call's usage. Add:

- POST /usage/sync, which parses the transcript and *sets* the cumulative
  totals absolutely (idempotent — re-syncing the same or a grown transcript
  overwrites, never double-counts), with a (size, mtime) short-circuit so an
  unchanged transcript skips the re-parse.
- usage-report-hook.sh, which hands the SDK the transcript path. Registered on
  PostToolUse (keeps mid-run snapshots and reaped-agent sessions accurate) and
  Stop (guarantees a final sync at turn end before finalize reads the totals).

Field mapping verified against a real Claude Code transcript:
message.usage.{input_tokens, output_tokens, cache_read_input_tokens,
cache_creation_input_tokens}. Unit tests cover summation, idempotency, growth,
a missing transcript, and malformed lines.
2026-06-11 07:46:19 +02:00
9f8834155a Feature: prompter gold upgrade (#84)
* feat(prompter): make the assistant a RoboCo insider and fully wire launch

The Prompter's intelligence lived in two thin static prompts, so it asked
generic checklist questions and produced a flat task. The launch path was
also only half-wired: the panel called the generic task-create endpoint with
no project, bypassing the Prompter's own confirm flow.

Interview brain
- Rewrite the chat system prompt with RoboCo's org model, the task-spec
  standard, a dimensions playbook, and a reflect-back, 1-2-questions-per-turn,
  auto-stop discipline.
- Inject the live projects/products list each turn so the assistant grounds
  questions in real surfaces and resolves the target itself.
- Replace the brittle phrase-match readiness with a parsed roboco-meta control
  block (parse_readiness); the block is stripped from the visible reply and the
  turn now returns draft_ready + scale.

Structured GOLD draft
- Add first-class draft fields (objective, what_this_builds, the_work, notes)
  carried in the existing draft_data JSONB — no migration.
- Compose the GOLD markdown description deterministically from those fields
  (compose_description); the model never hand-formats the body.

Adaptive routing + wired launch
- Confirm now runs through the Prompter confirm endpoint with the human's
  project/product choice and edited structured draft.
- Single-cell targets a project and the cell team; a multi-cell feature targets
  a product and becomes a Main-PM coordination root that fans out.

Frontend
- Turn-envelope draft_ready (drop the duplicated phrase-match), structured
  draft card, confirm dialog with a project/product picker and a per-cell
  The Work editor, and the corrected priority labels (0 highest .. 3 lowest).

* fix(prompter): commit session writes so they survive across requests

Session create returned 201 but the row was never durably committed, so the
immediately-following /messages call could not find it and 404'd. The prompter
routes were the only write surface that never called db.commit() — every other
write route (tasks, a2a, groups, docs, product) commits explicitly rather than
rely on the request-teardown auto-commit, which is sensitive to middleware and
teardown ordering under the production server.

- Commit explicitly in all four prompter write routes (create session, send
  message, get/generate draft, confirm).
- Fix _get_session's NotFoundError: it passed a full sentence as resource_type,
  producing the doubled "... not found not found" message; now uses the
  (resource_type, resource_id) signature.
- Panel: when a message hits a session the server no longer has, start a fresh
  session and retry once instead of dead-ending on a stale id.

Add a regression test that gives each request its own non-committing session —
the real cross-request boundary the shared-session integration tests never
crossed. It reproduces the production 404 without the route commit and passes
with it.

* refactor(prompter): drop the "GOLD" jargon for plain wording

"GOLD" was informal shorthand for "a good/well-formed spec" that should never
have been baked into the LLM prompts, comments, and docstrings as if it were a
defined term. Replace it everywhere with plain language ("a well-formed task",
"a complete task spec", "the markdown description", "structured spec fields").
No behaviour change.

* feat(intake): add the intake interviewer agent role (static definition)

Phase 1 of the intake-agent feature: a new first-class `prompter` role — the
intake interviewer the CEO chats with to draft a task. This commit defines the
role across every foundation layer (no runtime yet); spawning + the live
session come next.

- identity: Role.PROMPTER, RoleLevel.INTAKE (lowest authority), an AGENTS row
  (intake-1) on the board team, ROLE_LEVEL entry. Deliberately NOT in
  BOARD_ROLES — it interviews, it does not review.
- lifecycle: gets i_am_idle like every agent (its only verb); no
  delivery-lifecycle intents.
- journaling: ReadTier.OWN — isolated, reads only its own journal.
- role_config: human-only manifest — note + evidence only, no say/dm/notify/
  channels; allows_subagent=True (research), allows_write=False.
- agents_config derives it automatically and correctly excludes it from
  TASK_CREATOR_ROLES (it drafts, it never creates tasks).
- seed presentation ("Intake"); regenerated lifecycle artifacts.
- role system prompt: read the code first, single-CEO awareness, propose
  rather than interrogate — written against the failures we saw.
- docs: roster count 19 -> 20, org charts, verb-surface table, usage roster.

All foundation drift checks pass; role/manifest/permission tests green.

* feat(intake): migrate agentrole enum to add 'prompter'

ALTER TYPE agentrole ADD VALUE IF NOT EXISTS 'prompter' so the intake agent
row seeds/spawns against a migrated production DB. Forward-only (postgres
can't drop enum values), guarded for offline mode — matches migration 012.

* style(intake): ruff format the role additions

* fix(intake): unguard the agentrole migration so it renders offline

The enum-migration-parity test renders 'alembic upgrade head --sql' (offline)
and greps for ALTER TYPE ... ADD VALUE. The is_offline_mode() guard skipped
emitting it, so the parity check couldn't see 'prompter'. Drop the guard —
PG16 permits ADD VALUE in a transaction, same as migration 020's backfill.

* feat(intake): the live-session driver (Claude Agent SDK loop)

Phase 2 begins. The intake agent isn't a one-shot `claude -p`; it's a live
Claude Code session the human chats with. This driver is the container's loop:
open one long-lived claude-agent-sdk ClaudeSDKClient, then per human message
run a turn (query + receive_response) and stream its events out, keeping
conversation context in-process — verified against the real SDK (v0.2.94).

- StreamChunk + normalize(): map SDK messages (StreamEvent text deltas,
  AssistantMessage text/thinking/tool_use blocks, ResultMessage→session_id) to
  panel-facing chunks. Duck-typed, so it works on real SDK objects and on test
  fakes alike — SDK-free, fully unit-tested.
- IntakeDriver.run(): the loop, with injected session/source/sink seams; a turn
  failure surfaces as an error chunk without killing the session.
- SdkIntakeSession + build_intake_options: the only SDK-coupled code (lazy
  import; needs the live claude binary, so excluded from coverage).
- Add claude-agent-sdk dependency + mypy ignore-missing-stubs.

Relay, panel SSE, and the persistent on-demand spawn are the next steps.

* Updated uv.lock

* feat(intake): the panel<->agent live bridge (registry, routes, entrypoint, image)

Wires the live intake chat end to end (Phase 2 integration layer):

- prompter_live.py: the orchestrator-side per-session registry — open/close,
  push (agent->panel), stream (SSE drain), deliver (panel->container). In-process
  (the orchestrator is single-process). 7 unit tests.
- routes/prompter_live.py: GET /live/{id}/stream (SSE), POST /live/{id}/messages
  (deliver), POST /live/{id}/events (relay in); registered under /api/prompter.
  5 integration tests.
- agent_sdk/intake_main.py: the container entrypoint — a POST /turn receiver
  (the driver's MessageSource) + a relay-poster EventSink + the ClaudeSDKClient
  session, run concurrently. 4 unit tests on the wiring helpers.
- docker/agent-prompter.Dockerfile: FROM base, ENTRYPOINT = the driver (not the
  one-shot `claude` the other agents use).

Remaining for Phase 2: the orchestrator persistent-spawn path (scope->workspace
clone, CMD = driver, registry.open on spawn, reap-on-confirm) — the deploy-side
piece, best finalized against a buildable image.

* feat(intake): orchestrator persistent spawn + start/stop for the live chat

Add the task-free spawn path for the intake (prompter) agent: one fixed
intake-1 container running the Agent-SDK driver (image ENTRYPOINT, not
claude -p), one live session at a time.

- spawn_intake_session clones the scope's repo(s) via WorkspaceService
  (project -> one; product -> each distinct project, primary first),
  composes the intake-1 prompt, resolves the model, and builds docker run
  via _build_intake_run_cmd: no settings/hook mount (driver owns 9000),
  no MCP config, no -w; registers the live relay and best-effort delivers
  the opening message once the receiver is up.
- reap_intake_session closes the relay and stops the container.
- Routes: POST /live/start (project XOR product) and POST /live/{id}/stop.
- ROLE_MODEL_MAP[prompter]=opus; intake-1 -> roboco-agent-prompter image map.
- Replace the budget-sweep try/except/continue with _fetch_budget_status,
  which logs the swallow at debug instead of silently dropping it.

25 new tests; docker + the clone are mocked. End-to-end container spawn is
pending a built image and the stack.

* feat(intake): wire /prompter to the live agent — scope form + SSE chat

Replace the Ollama chat loop on /prompter with the spawned-agent flow.

- IntakeForm: pick scope (project XOR product) + opening message + Start
  before the chat; the agent clones that scope and reads the real code.
- use-prompter rewritten as the live brain (lib/api/prompter-live.ts): Start
  spawns via POST /live/start, then an EventSource on /live/{id}/stream
  streams the agent working — token deltas fill the assistant bubble,
  tool_use/thinking drive a live activity line, a draft event renders the
  existing DraftProposalCard. Messages go via POST /live/{id}/messages.
- Chat UX unchanged (Keep Chatting / Review & Confirm / ConfirmDialog reused);
  reap-on-confirm and reap-on-leave call POST /live/{id}/stop.
- Drop the dead Ollama prompterApi client; trim prompter.ts to shared types.

Frontend gate green (tsc --noEmit, lint, build). The draft event + the
/live/{id}/confirm endpoint are the Phase 4 backend seam.

* feat(intake): confirm draft -> backlog task + agent draft emission

Complete the live intake vertical: the agent proposes a structured draft and
Review & Confirm turns it into a task.

- Draft emission: the prompter prompt instructs the agent to emit a fenced
  roboco-draft JSON block when the spec is ready; the driver parses it into a
  'draft' event over the existing relay -> the panel's DraftProposalCard. The
  panel strips the raw block from the chat bubble.
- Fix a double-text bug: with include_partial_messages the reply arrives as
  both StreamEvent deltas and the final AssistantMessage; the driver now takes
  text from deltas only and the AssistantMessage for thinking/tool_use/draft.
- POST /live/{id}/confirm -> confirm_live_draft, reusing a draft->task core
  extracted from confirm_draft; reaps the session on success.
- Both prompter confirm paths create at BACKLOG, not pending: backlog is the
  holding area a draft waits in until it's reviewed and promoted to pending
  (TaskService.activate). The legacy Ollama confirm was creating at pending,
  skipping that gate — fixed.
- Remove the dead 'context' bootstrap param from the Ollama session-create
  chain (schema + route + method + tests), superseded by the live scope form.
- No suppressions: replace every type:ignore/noqa across the intake surface
  with a real fix (ORM .id -> UUID(str(x)); fakes -> monkeypatch.setattr;
  lazy imports -> pyproject per-file ignore; union-attr -> recipients[0]).

Full make quality green; frontend tsc + lint green.

* build(intake): add the agent-prompter image builder to compose

The orchestrator references roboco-agent-prompter (AGENT_IMAGES + the
_ensure_agent_image dockerfile map) and docker/agent-prompter.Dockerfile
exists, but docker-compose.yml built every other agent image up front and
left this one out — so the image wasn't pre-built for a stack bring-up.

Mirror the other specialized agent-*-image builders: build from
docker/agent-prompter.Dockerfile, tag roboco-agent-prompter, depend on
agent-base-image.

* Created docker-compose.yaml for the NAS

* fix(intake): non-blocking /live/start so spawn never times out

The start POST awaited the whole spawn — workspace clone + first-time image
build + docker run — which blew past the panel's 60s HTTP timeout ('Request
timed out. The server may be busy.') and triggered a duplicate send. Found on
the 2026-06-09 NAS smoke.

- start_intake_session opens the live relay synchronously, then spawns the
  container in the background (_spawn_intake_container_guarded). The route
  returns the session id immediately; the panel opens the SSE stream right away.
- A background spawn failure is pushed onto the relay as an 'error' event and
  closes the session, so the panel shows it instead of hanging.
- spawn_intake_session stays as the synchronous variant for direct callers/tests.
- Panel shows a 'Preparing the agent…' indicator until the first event arrives.

18 intake-spawn tests green; tsc + lint green. E2E re-validates on next smoke.

* fix(intake): propose_draft MCP tool + lock the agent down

Smoke 2026-06-09 exposed two compounding problems: the agent never reliably
emitted the draft (it narrated the spec instead of typing the magic fence), and
it had inherited the CEO's entire Claude Code env — Write/Edit/Bash + Gmail/
Notion/Calendar/Drive MCP — because bypassPermissions ignored the allowlist and
the mounted ~/.claude leaked the host MCP config.

- propose_draft: build_intake_options now registers an in-process SDK MCP tool
  (create_sdk_mcp_server + @tool). The agent calls it to submit the draft; the
  driver turns that ToolUseBlock into a 'draft' event (_is_propose_draft /
  _draft_from_tool_input, tolerant of nested/flat/JSON-string input). The fenced
  roboco-draft block stays as a fallback.
- Lockdown: strict_mcp_config=True + setting_sources=[] (ignore host MCP +
  settings); permission_mode 'dontAsk' + a can_use_tool gate enforcing a hard
  allowlist (Read/Grep/Glob/Task + propose_draft) replaces bypassPermissions.
- Prompt: call propose_draft (not a fence); the draft's downstream chain is
  backlog -> Board (PO + HoM) -> CEO approve -> Main PM, and the agent's job ends
  at the draft (it never routes or hands off).

SDK API verified against the installed claude-agent-sdk. Driver detection unit-
tested; the SDK-construction is validated on the next NAS smoke (incl. that
setting_sources=[] doesn't break the mounted-~/.claude auth).

* fix(intake): panel UX cluster from the smoke (#3/#4/#6/#12)

- #3 message boundaries: a tool call now ends the current text bubble, so the
  agent's words before and after a tool render as separate messages instead of
  one merged wall (the 'two waves merged into one bubble' the CEO saw).
- #4 activity indicator: promoted from tiny grey text to a prominent primary-
  tinted pill so 'watch it work' is actually visible.
- #12 End chat: a header button (any chat state) reaps the agent and resets to
  the form, reusing startAnother (which already stops the session). Backend
  POST /live/{id}/stop already existed.
- #6 log noise: the opening-message delivery retry logs at debug, not error —
  those failures are expected until the container receiver is up.
- Also fix a latent test gap from the #1 commit: the live-route test's fake
  orchestrator now exposes start_intake_session (the route's non-blocking entry).

Frontend tsc + lint green; live-route + prompter_live tests green.

* fix(intake): render markdown in the chat bubbles (#8)

The agent emits rich markdown (### headers, **bold**, tables, lists) but the
bubble rendered raw text, so it was illegible (CEO-flagged on the smoke). Render
assistant content with react-markdown + remark-gfm (GFM tables) in a prose
container. Adds react-markdown + remark-gfm to the panel.

* feat(intake): #14 — two start routes (Board review vs straight to Main PM)

Per the CEO spec, the draft confirm now starts the task at PENDING with an
explicit assignment instead of parking it at backlog:

- route="board" (Board review & Start): assigned to the Product Owner, so the
  orchestrator dispatches the full Board review (PO + Head of Marketing) before
  the Main PM picks it up.
- route="main_pm" (Approve & Start): assigned straight to the Main PM, who
  delegates to the cells (Board review skipped).

create_task_from_draft gains status + assigned_to params (default BACKLOG, so the
legacy confirm_draft is unchanged); confirm_live_draft + the /live/{id}/confirm
request carry the route. Service tests cover both routes.

* feat(intake): #14 draft-card buttons — Board review vs Approve & Start

Three buttons on the draft card now (CEO spec): Keep chatting / Board review &
Start / Approve & Start. The two action buttons confirm directly with their
route — launchTask(route) sends route to POST /live/{id}/confirm, which starts
the task at pending assigned to the Board (PO+HoM) or straight to the Main PM.

Supersedes the ConfirmDialog review step (scope is chosen up front in the form),
so it's removed from the page flow. The ConfirmDialog component + its sub-editors
are now unused — flagged for a follow-up cleanup, left in place to avoid churn.

tsc + lint green.

* fix(intake): keep the live SSE stream bound to its relay session

The orchestrator opened the relay session twice per live chat — once on the
request path (before the start call returns) and again inside the background
container spawn. The SSE stream binds to the session's queue the moment the
panel connects, so the second open swapped in a fresh queue and stranded the
stream: the agent replied normally, but its events went to the new queue while
the panel kept reading the old one, so the chat looked frozen on "Preparing…".

The second open was always redundant (the relay is opened by the caller before
the spawn). Remove it, and make open() idempotent so a live session is never
replaced out from under a stream that is already connected to it.

* fix(intake): draft-card launch buttons silently did nothing

The launch path required a `description` field, but the prompter draft schema
intentionally has none — it sends `objective` + the structured spec and the
backend composes the description (compose_description). `editableDraft.description`
was therefore undefined, so `description.trim()` inside launch validation threw a
TypeError that propagated out of the button's onClick. Clicking "Board review &
Start" / "Approve & Start" did nothing, with no feedback — the wall blocking the
whole confirm → task → reap flow.

- Map a proposed draft's description from `objective` as a fallback.
- Make launch validation null-safe.
- Replace the silent early-return with a toast that names what's missing, so a
  blocked launch is never a dead, feedback-less button again.

* fix(intake): steer the agent to ask inline, not via AskUserQuestion

The intake's job is to ask clarifying questions, so it reached for the
AskUserQuestion tool — which isn't wired to the live chat panel and isn't in its
allowlist. The bare deny left it to stumble ("let me clarify… — no worries, let
me just lay it out") and waste a visible turn.

- Prompt: spell out that it asks by writing in the chat (the human reads every
  message live) and that no question/prompt tool is available to it.
- Gate: give AskUserQuestion a specific deny message that nudges it to ask inline,
  so even a reflex attempt degrades gracefully.

Also refresh the now-stale "what happens after propose_draft" section: the draft
card has three choices (Keep chatting / Board review & Start / Approve & Start)
and produces a pending task — not the old two-button "backlog" description.

* feat(intake): copy buttons on agent messages and the draft card

The CEO asked for a way to save the agent's plan/spec elsewhere "just in case" —
a cheap manual backstop until refresh-durability lands.

- New CopyButton: async Clipboard API when available, plus a legacy
  textarea+execCommand fallback. The fallback is load-bearing — the panel is
  served over plain http on a LAN IP, where navigator.clipboard is absent
  (clipboard needs a secure context), so the modern API alone would never copy.
- Copy button under each assistant message (copies its text).
- Copy button on the draft card (copies the full spec as markdown: title,
  objective, what-this-builds, the-work per cell, notes, success criteria).

* feat(intake): unbuffer logs + log each turn so the container isn't a black box

Debugging the intake smoke was painful for two reasons: (a) the orchestrator
block-buffered stdout, so `docker logs` lagged minutes behind reality, and (b)
the intake container logged only "session opened" then went silent for the whole
conversation (the chat streams to the relay, not stdout).

- Set PYTHONUNBUFFERED=1 on the orchestrator and agent-base images so structured
  logs reach `docker logs` in real time instead of in large delayed chunks.
- Log each intake turn: "turn received" (with char count) and "turn streamed"
  (chunk count + whether a draft was emitted), so the container logs show the
  conversation's shape at a glance.

* chore(intake): remove the dead ConfirmDialog draft editor

The three-button draft card (Keep chatting / Board review & Start / Approve &
Start) replaced the old review-modal confirm flow, leaving ConfirmDialog and its
sub-editors (StringListEditor, TheWorkEditor) referenced by nothing but the
barrel export. Remove the three files and the export — typecheck + lint confirm
no remaining references.

* fix(intake): coerce bad draft enums on confirm instead of hard-failing

The intake agent is an LLM and will emit off-enum values — e.g. task_type="feature",
which is not a valid TaskType (code/documentation/research/planning/design/
administrative). `_coerce_draft_enums` called `TaskType(value)` directly, which
raised, and the confirm 400'd with "Draft has invalid or missing required fields:
'feature' is not a valid TaskType". That forced the agent to discover the valid
values and self-correct in-chat — unacceptable: clicking "Approve & Start" must
never blow up on a cosmetic enum guess.

Coerce each enum to a sane default on invalid/missing (task_type→code,
nature→technical, complexity→medium); team falls back to the first valid cell in
the_work, then backend. `_lead_cell_team` now skips invalid cell names too. The
confirm/launch action no longer hard-fails on an enum the model got wrong.

* fix(intake): draft card no longer renders above the user's latest message

attachDraft fell back to "the last assistant message anywhere" when the current
turn had no streamed text yet (propose_draft called first). That last message was
often the PREVIOUS turn's — sitting above the user's "Yes, propose it" — so the
draft card rendered above the user's message. Attach only to the current turn's
streaming message; otherwise append a fresh assistant message so the card always
lands at the bottom of the thread.

* test(intake): guard draft enum coercion + invalid-cell skipping

Regression tests for the confirm-time enum coercion: an off-enum task_type
("feature") / nature / complexity coerce to code/technical/medium instead of
raising, and _lead_cell_team skips invalid cell names. Locks in that a bad enum
guess from the agent can never 400 the launch again.

* fix(intake): stop the agent fumbling through Claude Code meta-tools

In smoke it reflexively probed CC built-ins before reaching propose_draft —
plan mode + ExitPlanMode (it announced a written plan and waited instead of
emitting the draft), ToolSearch, Write — each correctly denied by the lockdown
but stumbly, and it only proposed after explicit CEO nudges.

- Gate: ExitPlanMode now gets a specific deny nudge ("you don't use plan mode;
  call propose_draft"), and the generic deny names the actual toolset instead
  of a bare "not available", so any probe degrades into guidance.
- Prompt: forbid plan mode/ExitPlanMode/ToolSearch explicitly and spell out
  "you do not plan and wait — call propose_draft directly when the spec is
  ready," plus an anti-pattern bullet.

* feat(intake): make the container logs transparent mid-turn

`docker logs` on the intake container was a black box: only turn start/end, while
the agent read the codebase and spawned 20+ subagents invisibly (the conversation
streams to the relay, not stdout), and the benign 3x ~/.claude.json warning was
the only thing visible.

- Driver logs each tool call mid-turn ("Intake tool use" with the tool name) and
  the draft emission, plus a tools count in the turn-streamed summary. Text deltas
  stay unlogged (they'd spam). Now the logs show the turn's real shape.
- Pre-create ~/.claude.json ({}) at container boot so the CLI's "config not found"
  warning (printed 3x, self-healed anyway) stops drowning the real logs.

* fix(intake): render markdown in user messages + scope copy to code blocks

Two display fixes from the smoke:
- User messages collapsed newlines (plain {content} in a div) and rendered no
  markdown — a "1.\n2.\n3." answer showed as one run-on line. Render user AND
  assistant bubbles through a shared GFM markdown body that inherits the bubble's
  text color, so lists / newlines / styling render correctly on both.
- Copy was blanketed on every assistant message; scope it to KEY parts — a copy
  button on fenced code blocks (the draft card keeps its own). Removed the
  per-message button.

* fix(intake): prevent duplicate tasks from a double-click on launch

Clicking a draft launch button twice fired two confirms and created duplicate
tasks. Add a synchronous re-entry guard (a ref — no stale-closure window) at the
top of launchTask so a second click returns immediately, and disable + spin the
draft-card buttons while a launch is in flight so it's visually clear it's working.

* docs(how-to): lead task creation with the Task Assistant flow

Rewrite "1 · It starts with you" to walk the Prompter/Task Assistant path —
scope form, the agent reading the codebase, its grounded analysis, the draft
card, and the created task — then flow into the Board review. Replaces the old
manual task-definition form shots.

Image placeholder: images/prompter_draft_card.png (the 3-button card) is
referenced but not yet captured — TODO comment marks it for the next smoke run.
A second comment flags an optional re-capture of prompter_run_2 after the
markdown-rendering fix.

* fix(intake): restore assistant message text contrast

The markdown refactor dropped `dark:prose-invert` and made text inherit the
bubble's color, but the assistant bubble had no explicit text color — so its text
rendered near-invisible (dark-on-dark on bg-muted). Give the assistant bubble an
explicit text-foreground; the user bubble already carries text-primary-foreground,
and [&_*]:!text-inherit now resolves to a readable color on both.

* fix(intake): coerce draft priority too — confirm 500'd on priority="high"

The enum-coercion fix covered task_type/nature/complexity/team, but priority is a
non-enum int field handled by `int(draft_data.get("priority", 2))`, and the agent
guesses a word ("high") as readily as a number — so int("high") raised ValueError
and the confirm 500'd. Same class of bug, one field missed.

Add _coerce_priority: map words (urgent/high/medium/low → 0/1/2/3), clamp numbers
to 0-3, default to 2 (medium) on anything else. The launch can no longer crash on
any field the LLM guessed. + regression test.

* fix(intake): draft card shows distinct cells, not one badge per work item

the_work has one entry per work item, so a cell with several items rendered its
badge repeatedly ("Board-led across Backend Backend Backend Frontend Frontend
…"). De-dupe to distinct teams so the card reads "Board-led across Backend
Frontend" — and the "Cell:" vs "Board-led across" label keys off distinct count.

* docs(how-to): hero the teaser gif + resolve the Prompter/Task Assistant thread

- Move the 12s teaser gif to the top as the hero — it was buried between the
  "prefer video" link and the first screenshot.
- Name the connection: the Task Assistant IS the Prompter, so section 1 (using
  the tool) and the rest (RoboCo building it) read as one story — you use the
  tool the company built for itself, then watch the build.
- Re-anchor the section 1 → Board transition to follow the Prompter's own
  journey, instead of implying section 1's example task is the one reviewed next.

* Included images for how-to.md

* docs(how-to): align agent count to 20 (matches README + CLAUDE.md)

The how-to said "18 agents" with UX/UI at one dev and no Intake — stale against
the authoritative count. Bump 18→20 (prose + spelled-out eighteen→twenty), give
UX/UI 2 devs, and add the Intake line to the org tree (Intake leads section 1, so
it belongs in the tree). README + CLAUDE.md already say 20.

* ci(release): publish all RoboCo images to GHCR + Docker Hub

The release published only the orchestrator to GHCR. Build and push the full set
the stack needs — agent-base, the 8 agent images, orchestrator, and panel — to
BOTH ghcr.io/rennf93/* and docker.io/renzof93/*, at :<version> and :latest, so
consumers can pull instead of compose-building.

- agent-base builds first (the agent images build FROM roboco-agent-base, a local
  tag), then the rest; push only after every build succeeds.
- Image names mirror the docker-compose `image:` values 1:1.
- Free disk on the runner first (11 images is space-heavy).
- Needs a DOCKERHUB_TOKEN repo secret for the Docker Hub login.
- SECURITY.md updated to reference both registries.

* ci(release): use short SHA as the image tag on manual dispatch

A workflow_dispatch runs against a branch, and the branch name (e.g.
feature/prompter-gold-upgrade) was used verbatim as the image tag — but "/" is
illegal in a Docker tag, so the first build failed instantly with "invalid
reference format". Releases still tag from the release tag; manual dispatch now
always uses the short SHA, which is a valid tag.

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-06-09 17:08:34 +02:00
Renn F d4126ffb7f fix(sdk): _TERMINAL_TOOLS uses current gateway verb names
Smoke-7 evidence: every agent's first successful i_am_idle was followed
by a stop-hook error "you stopped without calling a terminal tool"
even though i_am_idle had just succeeded.

Root cause: agent_sdk._TERMINAL_TOOLS still held pre-gateway names
(roboco_agent_idle, roboco_task_submit_qa, ...). /terminal/tool_recorded
strips the `mcp__roboco-flow__` prefix and stores 'i_am_idle' — the
membership check against {'roboco_agent_idle', ...} never matched, so
had_terminal_recently() always returned False, and the stop hook nagged
every clean exit. Wasted ~2-3 turns per agent + burned stop_allowance.

Fix: rebuilt _TERMINAL_TOOLS with the current gateway verb names:
i_am_idle, i_am_done, i_am_blocked, i_documented, unclaim, pass, fail,
complete, submit_up, escalate_up, escalate_to_ceo.

7 new tests pin:
- every role's terminal verbs are recognized
- pre-gateway names are NOT in the set
- _SessionState.had_terminal_recently() returns True after i_am_idle
2026-05-15 03:20:43 +02:00
207aaecd72 Feature: lifecycle canonical spec (#14)
* chore: clean make quality baseline on feature/lifecycle-canonical-spec

Three classes of pre-existing issues blocking `make quality`:

1. Alembic migrations 002/009/011 used runtime introspection
   (op.get_bind() + inspect / bind.execute) without guarding for
   offline (--sql) mode. `alembic upgrade head --sql` is part of
   `make quality`; in offline mode `op.get_bind()` returns a
   MockConnection with no inspection system, so the migrations
   crashed before emitting their SQL stubs. Each migration now
   short-circuits or simplifies in `context.is_offline_mode()` —
   live-DB behavior is unchanged.

2. ruff format drift on three files left over from prior in-flight
   edits (choreographer/_impl.py, content_actions.py, and one test
   file). `ruff format` applied.

3. vulture flagged two unused `tb` parameters in async __aexit__
   stubs in test_task_service_lifecycle_misc.py. The parameter is
   protocol-required but unused by the body — renamed to `_tb`
   (vulture treats underscore-prefixed names as intentionally unused).

`make quality` is now green from this branch's HEAD; subsequent
lifecycle-spec work can use it as the per-task gate.

* feat(lifecycle): canonical spec package + Role/Status/TaskType enums

Foundation for the canonical lifecycle/permissions module. Enums
mirror docs/internal/old/workflows/STATUS_TRANSITIONS.md +
PERMISSIONS.md. Tests pin enum membership against both the
predecessor canon and roboco.models.base.TaskType.

* feat(lifecycle): Decision dataclass with allow/reject/tracing_gap constructors

Single rejection shape every consumer maps to its native format
(Envelope, HTTP code, prompt hint). __post_init__ enforces the
allowed/rejection_kind invariants so a malformed Decision can't reach
a consumer.

* fix(lifecycle): tighten Decision invariants per Task 2 review

Two reviewer findings on the Task 2 Decision dataclass, addressed
in one commit:

1. The docstring promised `allowed=True ⇒ rejection_kind is None
   AND missing == [] AND remediate is None`, but __post_init__ only
   checked the rejection_kind half. A caller could construct an
   allow-shaped Decision with stale missing/remediate fields and
   sneak it past validation. Tighten __post_init__ to enforce the
   full invariant. Add a regression test.

2. tracing_gap defensively copies the missing list (`list(missing)`)
   to isolate the stored list from later caller-side mutation, but
   no test pinned this. Add a regression test that mutates the source
   list after construction and asserts the stored list is unchanged.

Issue 2 from the same review (mutable list vs tuple for `missing`)
is a broader design call deferred until consumers exist; the
defensive copy is sufficient until then.

* feat(lifecycle): Precondition/ActionSpec/IntentSpec/StatusTransition dataclasses

The four dataclasses that hold the canonical tables. ActionSpec and
StatusTransition are direct ports of pre-gateway PERMISSIONS.md +
STATUS_TRANSITIONS.md rows. IntentSpec is the gateway-only addition:
each gateway intent verb declares which atomic actions it composes.

* feat(lifecycle): _STATUS_TRANSITIONS table + STATUS_GRAPH view

Direct port of STATUS_TRANSITIONS.md. Every transition records its
trigger action and (optionally) a role constraint. STATUS_GRAPH is
the precomputed source→{targets} view callers use for reachability
checks.

* fix(lifecycle): pin role_constraint values + clarify Task-5 handoff

Two reviewer findings on Task 4 _STATUS_TRANSITIONS, addressed in
one commit:

1. The original Task-4 tests verified (source, target) pairs but
   not role_constraint contents. A typo in a single role name (e.g.
   forgetting MAIN_PM from escalate_to_ceo) would have slipped past
   them silently. Add test_status_transitions_role_constraints_match_canon
   pinning every non-None constraint and the cancel-block invariant.

2. role_constraint=None on the `claim` rows from PENDING and
   NEEDS_REVISION was load-bearing — it is the explicit handoff
   point between the StatusTransition table (state machine layer)
   and CLAIM_RULES (per-role claim authority, lands in Task 5).
   The original inline comment said this in passing; expand it so
   the design choice is unmissable for a stranger reading just
   spec.py.

* feat(lifecycle): _ATOMIC_ACTIONS + CLAIM_RULES + ROLE_TEAM_RULES tables

Direct port of PERMISSIONS.md. Every task management tool gets an
ActionSpec with allowed_roles, source_statuses, target_status,
self_review_block, and needs_team_match flags. CLAIM_RULES maps each
Role to the statuses they can claim from. ROLE_TEAM_RULES is the
per-slug team restriction.

* fix(lifecycle): tighten ActionSpec contracts per Task 5 review

Three reviewer findings on Task 5's _ATOMIC_ACTIONS table, addressed
in one commit:

1. set_plan.source_statuses widened to {CLAIMED, IN_PROGRESS} but
   every existing caller (i_will_work_on / i_will_plan compositions)
   runs set_plan while CLAIMED, between claim and start. Narrow to
   {CLAIMED} only. If a future "edit plan mid-flight" feature lands,
   widen explicitly with test coverage at that time.

2. needs_team_match was set True only on claim/qa_pass/qa_fail/
   docs_complete. Defense-in-depth says every role-scoped task
   action should re-assert team match (don't rely on the inheritance
   chain through assigned_to alone). Flip to True on: start,
   set_plan, block, pause, submit_verification, submit_qa,
   submit_pm_review, complete, create_subtask. Leave False on
   board/CEO actions and PM cross-cell interventions (unblock,
   resume, cancel) where the cross-cell semantics are intentional.

3. claim.source_statuses is intentionally a SUPERSET of any single
   role's CLAIM_RULES allowance (the table holds the union; CLAIM_RULES
   holds the per-role authority). Add an inline comment above the
   claim ActionSpec so a future reader doesn't conclude the two
   tables disagree — they don't, they encode overlapping facts at
   different grains.

* feat(lifecycle): _INTENT_VERBS table — every gateway verb declared

Each gateway intent verb is now a named composition of atomic actions
plus optional side effects. i_will_work_on = (claim, set_plan, start);
i_am_done = (submit_verification, submit_qa); open_pr is pure side
effects (push_branch, create_pr); etc.

* fix(lifecycle): widen block.allowed_roles to include QA + Documenter

Task 6 review caught a role-set inconsistency: i_am_blocked.allowed_roles
admits dev/QA/doc, but the underlying block.allowed_roles only allowed
dev+PM. Result: a QA or documenter calling i_am_blocked would pass the
IntentSpec gate and then be rejected by the composed ActionSpec gate
when Task 7 wires can_invoke_intent.

Widen block to include QA + Documenter. The semantic case is sound: a
QA reviewing a task can discover an external blocker; a documenter
writing docs may need PM intervention. Predecessor PERMISSIONS.md
restricted block to dev+PM, but with the gateway exposing i_am_blocked
to all worker roles, the underlying atomic must agree.

The deeper unclaim/escalate_up "imperative verb" concern from the same
review (composes=() but mutates state) is deferred to Task 8 where the
validator design lands.

* feat(lifecycle): public lookup functions + Context + preconditions

can_claim, can_invoke_action, can_invoke_intent, valid_next_verbs,
composed_actions_for, intents_for_role, status_after — the entire
public surface every consumer will use. Context carries the
caller-supplied state preconditions need (plan, journal-decision
flag, etc.). Preconditions for plan/commits/no_pr/ownership are
declared once and wired into the relevant IntentSpecs.

* fix(lifecycle): wire PRECONDITION_OWNERSHIP through Context.actor_id

Task 7 review found _p_owns_task reads agent.id but every call site
passes None for the agent arg. Result: getattr(None, "id", object())
returns a fresh sentinel, task.assigned_to == <sentinel> is always
False, and open_pr / i_am_done would reject every owner the moment
Task 9 wires consumers.

Fix: thread identity through Context.actor_id (new UUID field) and
rewrite _p_owns_task to read from the context. Both call sites already
pass the Context — no signature changes elsewhere. Add green-path
test exercising the owner-can-open-pr case the existing tests
missed (the Task 7 plan only tested precondition-failure paths,
which masked the bug).

Plus surface hygiene: STATUS_GRAPH, CLAIM_RULES, ROLE_TEAM_RULES,
and the four PRECONDITION_* constants are now in
roboco.lifecycle.__init__.__all__ so consumers in Tasks 8/9 don't
depend on the implicit `from roboco.lifecycle.spec import ...`
backdoor.

* feat(lifecycle): import-time self-consistency validators

10 validators run at module import; first failure raises
LifecycleSpecError and prevents the package from loading. Covers
status enum coverage, reachability, terminal exits, intent
compositions, status chain consistency, claim-rule role/status
coverage, self-review symmetry, team-rule slug existence, and
StatusTransition action references.

* fix(lifecycle): close validator gaps; resolve BACKLOG-claim and submit_qa IN_PROGRESS-shortcut ambiguity

Three reviewer follow-ups on Task 8's _validate.py, plus two real
data corrections the new action-target-reachability validator
surfaced.

1. Design spec §9 calls for "every ActionSpec.target_status, when
   set, is reachable from each source_status via STATUS_GRAPH" —
   missing from Task 8's 10 validators. Add
   _check_action_target_reachable_from_source.

2. _check_role_team_rules_slugs verified slug existence in
   AGENT_UUIDS but NOT that the cell team in ROLE_TEAM_RULES
   matches the seed. Add _check_role_team_rules_team_match,
   scoped to non-None entries only — None means "exempt from
   team-match enforcement" (cross-cell roles), not "no team in
   org chart".

3. test_validators_pass_on_real_spec was ceremonial. Add
   test_run_all_validators_raises_on_unknown_intent_action,
   a deliberate-break regression that monkeypatches _INTENT_VERBS
   to inject a fake action and asserts LifecycleSpecError raises.

The new action-target-reachability validator caught two real
data inconsistencies between the predecessor canon docs and the
spec tables:

A. claim.source_statuses listed BACKLOG and CLAIM_RULES[*PM]
   listed BACKLOG, but STATUS_GRAPH[BACKLOG] = {PENDING, CANCELLED}
   only. Resolution: PMs use the explicit \`activate\` action to
   move BACKLOG → PENDING, then claim from PENDING. Drop BACKLOG
   from claim.source_statuses and CLAIM_RULES.

B. submit_qa.source_statuses listed IN_PROGRESS, but
   STATUS_GRAPH[IN_PROGRESS] does NOT include AWAITING_QA. The
   intent verb i_am_done composes (submit_verification, submit_qa)
   which forces IN_PROGRESS → VERIFYING → AWAITING_QA — no
   shortcut. Drop the stale IN_PROGRESS entry from
   submit_qa.source_statuses.

Both corrections tighten the canonical state machine to a strict
no-skip transition graph. Pre-gateway PERMISSIONS.md/STATUS_TRANSITIONS.md
disagreements are resolved here; spec.py is the canon now.

* feat(gateway): Envelope.from_decision maps lifecycle Decisions to envelopes

Single shape adapter so verb bodies stop hand-composing rejection
envelopes. Each rejection_kind maps to a specific envelope flavor;
'self_review' folds into 'not_authorized' with a parenthetical hint;
constructing from an allow Decision raises (programmer error).

* feat(gateway): VerbRunner for atomic composed-action dispatch

Wraps spec.composed_actions_for(intent) in session.begin_nested()
so mid-sequence failures roll the DB back. Side effects run AFTER
the savepoint commits. Each atomic action name dispatches to a
TaskService method via a single, exhaustive _dispatch_atomic
mapping. New verbs slot in by adding an IntentSpec entry + a
_dispatch_atomic case if a new atomic is needed.

* refactor(gateway): i_will_work_on uses spec.can_invoke_intent + VerbRunner

Replace the bespoke status-branch dispatcher in i_will_work_on with the
spec-driven flow: load task -> load agent -> build spec.Context ->
spec.can_invoke_intent (and spec.can_claim for per-role status authority)
-> Envelope.from_decision on rejection -> VerbRunner.run_intent on success.

The _i_will_work_on_pending, _i_will_work_on_claimed,
_i_will_work_on_needs_revision, and _start_failed_envelope helpers are
removed; the runner replaces them. Two narrow verb-body re-entry blocks
remain for behaviors the spec does not yet model:

  1. in_progress + same agent -> idempotent heartbeat-only return
  2. claimed + same agent -> _resume_from_claimed (set_plan + start)
     to recover from a stuck mid-claim crash without re-running claim
     against a state the spec excludes.

The behavioral claim guards (already_active / paused / sibling_sequence)
also stay imperative for now -- they're not in the spec yet and migrate
into spec.extra_preconditions in a later task. Per-role claim authority
is enforced via spec.can_claim because the atomic claim action's
source_statuses are the union across roles; CLAIM_RULES narrows.

Parity test in tests/lifecycle/test_consumer_parity.py runs the verb
against every (role x status x task_type='code') combo (112 rows) and
asserts the envelope error matches the spec's Decision (or can_claim's
Decision when the intent gate passes but per-role claim authority does
not). This is the contract that makes spec/verb drift impossible.

Existing tests updated where rejection-message text changed (the spec
now produces the messages, e.g. "role 'cell_pm' may not call
'i_will_work_on'" instead of "PM cannot execute code") or where the
spec's stricter view ("invalid_state" -> "not_authorized" for a dev
trying to claim awaiting_qa) is more accurate. Test fixtures were
updated to wire task.session.begin_nested as a proper async context
manager (required by VerbRunner) and to set agent_for().id so runner-
driven calls line up with assert_awaited_with(task_id, agent_id).

* refactor(lifecycle): push CLAIM_RULES enforcement into can_invoke_action

Task 11's i_will_work_on migration had to call spec.can_claim()
separately after spec.can_invoke_intent() because the claim action's
source_statuses is the union across all claim-eligible roles —
can_invoke_intent alone would let a developer pass for claiming
awaiting_qa (a QA-only state).

The retrofit pattern would repeat in every claim-composing verb
(i_will_plan, claim_review, claim_doc_task). Push the per-role
narrowing inside can_invoke_action when the action is "claim",
using the same not_authorized vs invalid_state disambiguation
can_claim already implemented (status-reserved-for-another-role
returns not_authorized; status-no-role-can-claim returns
invalid_state). Extracted the body to _check_claim_rules_narrow
to keep can_invoke_action under xenon's complexity threshold.

Update _i_will_work_on_gate to drop the redundant spec.can_claim
call. Update test_consumer_parity.py to assert only against
can_invoke_intent's Decision.

Tasks 12-22 will inherit the cleaner pattern: spec.can_invoke_intent
is the single gate; verb bodies don't need per-action retrofits.

* refactor(gateway): i_will_plan uses spec.can_invoke_intent + VerbRunner

Migrates i_will_plan to the spec-driven pattern Task 11 set up for
i_will_work_on. The verb body now: (1) loads task + agent, (2) builds
Context, (3) checks idempotent/recovery re-entry, (4) calls
spec.can_invoke_intent, (5) returns Envelope.from_decision on
rejection, (6) delegates composition to VerbRunner. The
_i_will_plan_* helpers are removed — the runner replaces them.

Parity test in tests/lifecycle/test_consumer_parity.py runs the verb
against every (role × status × task_type) combo and asserts the
envelope matches spec.Decision.

* refactor(gateway): delegate uses spec.can_invoke_intent for role/state gate

Migrates delegate to the spec-driven role/state gate. The chain
validation (main_pm->cell_pm, cell_pm->its team's devs), the
assignee-vs-task_type rule (Cell PMs receive planning-typed only),
the enum coercion, and the parent-lifecycle/cap guards STAY in the
verb body — they encode delegate-specific semantics the spec
doesn't model.

Parity test in tests/lifecycle/test_consumer_parity.py asserts the
spec's role+state rejection is correctly surfaced. Chain/assignee
rejections continue to be tested in test_choreographer_pm_extras.

* refactor(gateway): open_pr uses spec.can_invoke_intent + VerbRunner

Migrates open_pr to spec-driven gating. The spec's
extra_preconditions (PRECONDITION_OWNERSHIP, PRECONDITION_COMMITS,
PRECONDITION_NO_PR) handle all three precondition checks; the verb
body delegates side-effect dispatch (push_branch, create_pr) to
VerbRunner.

Idempotent re-entry retained: an open_pr call against a task that
already has a PR (and the caller owns it) returns OK without
re-opening, rather than the tracing_gap the spec would otherwise
produce. This preserves agent ergonomics — two calls in a row
shouldn't surface a misleading "no_prior_pr" hint.

Parity test in tests/lifecycle/test_consumer_parity.py runs the verb
against representative (status x commits x pr_number) combos and
asserts the envelope matches spec.Decision.

* refactor(gateway): i_am_done uses spec.can_invoke_intent + VerbRunner

Migrates i_am_done to spec-driven gating. The spec's
extra_preconditions (PRECONDITION_OWNERSHIP, PRECONDITION_COMMITS)
handle ownership and commit-count checks; VerbRunner dispatches
the (submit_verification, submit_qa) atomic chain.

The tracing-gate preconditions (progress entry, journal:reflect,
acceptance criteria) and the field-level submit-qa gates stay in
the verb body — they model gates the spec doesn't yet cover.
Defense-in-depth: those gates run after the spec accepts the
ownership/commits checks.

Parity test in tests/lifecycle/test_consumer_parity.py runs the
verb against (role × status × ownership × commits) and asserts
the envelope matches spec.Decision.

* refactor(gateway): i_am_blocked uses spec.can_invoke_intent + VerbRunner

Migrates i_am_blocked to spec-driven gating. The journal:struggle
write stays in the verb body (it's a side effect outside the
lifecycle action). VerbRunner dispatches the `block` atomic action
via task_service.escalate.

Parity test in tests/lifecycle/test_consumer_parity.py.

* refactor(gateway): unclaim and resume use spec.can_invoke_intent

Migrates both verbs to the spec-driven gate. unclaim's verb body
keeps its dispatch (task.unclaim_for_agent) because composes=();
resume goes through VerbRunner with composes=("resume",).

The reassignment-rejection branch (introduced in 19f27b4 for the
2026-05-08 trace's "not your claim" case) stays - the spec doesn't
model "task got reassigned out from under you by an upstream verb,"
and the existing envelope text ("current owner: X - call
give_me_work() to find your current work") is the load-bearing
hint that fixed the original bug. Extracted the shared branch into
_reassigned_rejection / _ReassignedCtx so both verbs reuse it
without duplicating the envelope construction.

Parity tests in tests/lifecycle/test_consumer_parity.py.

* refactor(gateway): complete uses spec.can_invoke_intent at the dispatcher

Migrates the top-level `complete` dispatcher to gate role/state via
spec.can_invoke_intent before routing to cell_pm_complete or
main_pm_complete. The two lower-level methods keep their existing
PR-merge / CEO-escalation logic and pre-flight guards (those model
journal:decision preconditions and PR-mergeability checks the spec
doesn't model yet).

The runner pattern is NOT applied here — `complete` has two divergent
runtime paths (Cell PM merges leaf into parent branch; Main PM opens
master PR + escalates to CEO) that don't fit the runner's
single-composition model. Verb-body-owns-dispatch is the right
pattern.

Parity test in tests/lifecycle/test_consumer_parity.py runs the
verb against (role × status) combos and asserts the dispatcher's
spec rejection is correctly surfaced.

* refactor(gateway): escalate_up, escalate_to_ceo, submit_up use spec.can_invoke_intent

Migrates the three PM-side escalation/submission verbs to
spec-driven role/state gating. The verb-specific guards
(journal:decision, escalation_target configured, _submit_up_guard's
ownership + notes-length + subtasks-terminal) STAY in the verb body
- the spec doesn't model these.

escalate_up has composes=() so the verb body owns dispatch via
task.escalate. escalate_to_ceo and submit_up route their
compositions through VerbRunner.

Parity tests in tests/lifecycle/test_consumer_parity.py.

* refactor(gateway): qa.py + doc.py role mixins use spec.can_invoke_intent

Migrates the five QA + Documenter verbs (claim_review, pass_review,
fail_review, claim_doc_task, i_documented) to spec-driven gating.

The self-review block lives at the atomic-action layer
(_ATOMIC_ACTIONS["qa_pass"|"qa_fail"|"docs_complete"].self_review_block=True)
and naturally fires when the verb body builds a Context with
actor_slug==original_developer_slug. No verb-body retrofits needed.

The verb-specific helpers (_verify_qa_owner, _qa_pass_gate_check,
_check_i_documented_inputs) STAY — they encode notes-length /
journal:learning / files-list / qa_evidence_inspected gates the
spec doesn't model.

claim_review and claim_doc_task own dispatch via task.qa_claim /
task.doc_claim respectively (not the runner) because those
specialized claim methods keep status at AWAITING_QA /
AWAITING_DOCUMENTATION, which is what the downstream qa_pass /
qa_fail / docs_complete source-status requirement expects.
The spec gate still validates role + claim source-status + task_type
before dispatch.

pass_review / fail_review / i_documented route their compositions
(qa_pass / qa_fail / docs_complete) through VerbRunner.run_intent
inside a savepoint.

Parity tests in tests/lifecycle/test_consumer_parity.py for all
five verbs.

* fix(lifecycle): claim_review and claim_doc_task have empty composes

Tasks 21-22 surfaced a real spec/runtime mismatch: both verbs were
declared composes=("claim", "start"), but the actual implementation
uses task.qa_claim / task.doc_claim which intentionally keep status
at AWAITING_QA / AWAITING_DOCUMENTATION. If the runner ever ran the
declared composition, it would transition the task to CLAIMED then
IN_PROGRESS, breaking the source-status invariants of qa_pass,
qa_fail, and docs_complete.

The spec is the canon — align it to the runtime. composes=() means
"verb body owns dispatch" (same pattern as escalate_up and unclaim).
The spec gate still validates role + AWAITING_QA / AWAITING_
DOCUMENTATION source-status via the role's CLAIM_RULES narrowing,
enforced through special handling in can_invoke_intent, so role/state
safety is preserved.

* refactor(gateway): role_config flow lists derived from spec.intents_for_role

Hand-maintained _DEV_FLOW etc. tuples replaced with calls into the
spec. Adding/removing a role from an IntentSpec.allowed_roles now
automatically updates the MCP manifest. The spec is the canon;
role_config becomes a thin shim that adds the do-tool / write /
subagent / description metadata the spec doesn't carry.

* feat(lifecycle): generators + make lifecycle for deterministic artifact regen

Renders intent-verbs.md, status-transitions.md, panel/lib/lifecycle.json,
and per-role agents/prompts/_generated/lifecycle-{role}.md fragments
from the canonical spec. `make lifecycle` runs the regenerator;
deterministic output enables CI to gate on `git diff --exit-code` after
running it. The agent prompt fragments will be injected at the top of
each role's system prompt (Task 25) so agents see the same verbs the
gateway accepts.

* feat(lifecycle): inject generated prompt fragments + CI drift gate

Each agent's system prompt now starts with the spec-generated
'verbs available to your role' fragment. CI runs make lifecycle
and fails if regeneration produces a diff — drift between spec
and artifacts cannot land on master.

* refactor(gateway): delete verb_gates.py — superseded by lifecycle.spec

verb_gates.is_verb_allowed and verb_gates.valid_next_verbs are now
spec.can_invoke_intent(...).allowed and spec.valid_next_verbs.
Importers updated to consume the canonical spec module directly.
tests/unit/gateway/test_verb_gates.py removed — coverage lives in
tests/lifecycle/test_spec.py.

envelope.with_introspection wraps spec.valid_next_verbs with role-string
coercion + best-effort try/except so malformed task fixtures (AsyncMock
status) and unknown role strings still yield [] instead of raising —
preserves the legacy verb_gates contract.

content_actions content-tool RBAC (commit/notify) is now a pair of
explicit role frozensets in this file. These are content tools, not
lifecycle intents, so they intentionally do NOT live in spec._INTENT_VERBS.

Two existing introspection tests asserted "commit" in valid_next_verbs;
fixed to assert open_pr/i_am_done — commit is correctly absent under
the canonical spec because it is a do-server content tool, not a flow
intent verb.

* refactor(gateway): collapse scattered role constants into spec

The pm_cannot_execute_code_guard and role_typed_claim_guard guards
both modeled rules the spec now handles via can_invoke_action's
CLAIM_RULES narrowing and ActionSpec.allowed_task_types. Drop them
from claim_guards.py — the choreographer's existing skip-flags on
_run_claim_guards are now permanent: those guards no longer fire.
Simplify _run_claim_guards's signature accordingly.

The concurrency-invariant guards (already_active_guard,
paused_tasks_guard, sibling_sequence_guard) STAY — the spec doesn't
model these system-level invariants. sibling_sequence_guard's loop
body extracted into _earlier_blocking_sibling helper to keep the
slimmed module under xenon's --max-modules A average.

* refactor(enforcement): task_lifecycle becomes a thin view of lifecycle.spec

VALID_TRANSITIONS and ROLE_RESTRICTED_TRANSITIONS are now derived
from roboco.lifecycle.spec — no independent tables. The 433-line
file collapses to ~30 lines of view definitions; future changes
go in spec.py. Helper functions exported by the legacy module are
preserved as thin wrappers so existing consumers don't need to
change their imports today.

A small _LEGACY_OPERATIONAL_EDGES table sits alongside the
spec-derived view to cover transitions the runtime exercises but
the spec has not yet absorbed (voluntary unclaim, reaper sweep,
PM-direct completes from in_progress, parallel-doc-PR developer
trigger). It is fenced and clearly documented; once those callers
are migrated to spec-driven dispatch the constant goes empty and
the file collapses to a pure view.

A test in test_task_service_lifecycle_misc.py was rewritten: the
predecessor asserted CEO-only authority over awaiting_ceo_approval
cancels (legacy table behavior), but the canonical spec authorizes
{CELL_PM, MAIN_PM, CEO} uniformly across all non-terminal cancel
sources. The test now exercises the broader spec-defined cascade.

* feat(lifecycle): UNMIGRATED guard pins known-debt consumers

Two pieces of debt surfaced during Task 28's collapse of
enforcement/task_lifecycle.py: (1) ~11 operational edges still in
the shim's _LEGACY_OPERATIONAL_EDGES because the spec's
_STATUS_TRANSITIONS doesn't yet model them; (2) role-gate
disagreements in _LEGACY_ROLE_GATES that the spec disagrees with.

UNMIGRATED is the named-debt set; KNOWN_UNMIGRATED_CONSUMERS pins
the catalog so a contributor adding a new entry must update both
sides. Validator (_check_unmigrated_is_subset) fires at import if
they drift. Test pins the current entries.

Phase 3's terminal invariant is `UNMIGRATED == frozenset()` —
expected when both legacy data carriers fold into spec, at which
point the assertion becomes a permanent regression guard.

* test(lifecycle): tier 3 end-to-end real-DB happy paths

Eight integration tests covering every major lifecycle path:
dev (pending → awaiting_qa), QA pass, QA fail, doc handoff,
Cell PM complete, Main PM escalate-to-CEO, block+unblock,
pause+resume. Each test drives the spec → choreographer →
TaskService → DB stack with only the git layer mocked. Catches
"spec says X, DB constraint says Y" mismatches the unit-tier
parametrized parity suite cannot detect.

* test(lifecycle): tier 4 smoke replay — pin known-bug shapes after spec migration

Synthesized fixture covering the 9 bugs from the 2026-05-08
audit-log trace + the 2 from the 2026-05-09 follow-up trace. Each
record documents (verb, role, task setup, expected post-fix
envelope shape, fix commit, spec invariant). The replay test
parametrizes over the records and asserts the spec / choreographer
behavior now matches the post-fix expectation — locks in the
fixes as permanent regressions.

The original audit log was wiped during cleanup; the fixture is
a documented synthesis, not a verbatim capture. The bug list is
faithful to the prior session's analysis of the trace.

* fix(orchestrator): silence dev-dispatcher noise for non-dev-lane tasks

Dev dispatcher fetched all pending/claimed/in_progress tasks regardless
of assignee role and warned 'role/task_type mismatch' on each pass when
it found cell_pm/main_pm/product_owner/etc. tasks — those belong to
_dispatch_pm_work, not this lane. The 30s warning loop showed up
prominently in the 2026-05-10 smoke run.

Filter at the lane boundary: silently skip when assignee role is not
developer/documenter/unknown. The D-49 misassignment warning still
fires for the legitimate cases (developer assigned a documentation
task, etc.).

* fix(gateway,prompts): unblock the three smoke-run dead-ends

Three issues surfaced by the 2026-05-10 smoke run, fixed together
because they're all blockers for end-to-end task completion:

1. Acceptance-criteria tracing gate was unsatisfiable. Nothing in the
   codebase writes to task.acceptance_criteria_status, so
   _check_acceptance_criteria always returned every criterion as
   missing. Treat a reflect note as the addressing artifact: when the
   agent has written one, the gate clears. Per-criterion citation via
   acceptance_criteria_status is still honored when populated, so the
   schema stays available for future per-criterion tracking.

2. Cell PM runaway re-decomposition. On every wake-up be-pm
   re-decomposed its parent task without checking for existing
   children, producing duplicate dev subtasks. cell_pm.md now teaches
   'list children before delegating' and 'one dev subtask is usually
   enough — QA/Documenter/PM-merge engage automatically'. Added
   anti-pattern entries for re-decomposition and over-decomposition.

3. Main PM exit/respawn loop on claimed-state tasks. The model
   cycled through delegate/resume/escalate/unblock looking for a verb
   that worked on 'claimed', and got cleanly rejected by every one.
   The right verb is i_will_plan (it composes claim+set_plan+start
   and resumes from claimed). main_pm.md now spells this out
   explicitly with a worked example of which verbs reject and why.

* feat(prompts): restore pre-gateway lifecycle scaffolding across all 6 roles

The gateway migration shrank role prompts from ~50 lines to ~15
(commit 534152c for dev; analogous shrinks for qa/doc/cell_pm/main_pm/board
in e12a596, 05ac832, 8dc381b). The verb surface got cleaner but the
prescription for using verbs through the lifecycle disappeared. The
2026-05-10 smoke run surfaced the regression: agents thrash through
verbs hoping one fits, journal sparsely, skip the dev reflect note,
and (for cell PMs) re-decompose on every wake-up.

Each role prompt now restores three sections that the pre-gateway
versions had:

1. State -> Verb table — what to call when respawned in each
   lifecycle status. Eliminates the verb-cycling antipattern: the
   agent looks up its current status and calls the one verb that
   transitions out of it.

2. Mandatory pre-handoff checklist — explicit walk-through of the
   gates the next verb will check, ordered so the agent fixes the
   missing piece before retrying:
   - developer: 7 items before i_am_done
   - qa: 8 items before pass/fail (incl. self-review forbidden,
     read dev journal not just diff, name artifact per criterion)
   - doc: 7 items before i_documented
   - cell_pm: 7 items before submit_up (incl. integration green)
   - main_pm: 7 items before complete(root)
   - board: separate checklists for escalate_to_ceo (PO/HoM) and
     reflect-note quality (Auditor — its only output)

3. Journaling cadence — when to use each of the five scopes
   (note/decision/struggle/learning/reflect). The pre-gateway
   prompts named all five scopes with role-specific examples;
   the post-gateway prompts mention 'reflect' once and skip the
   rest. Restored across every role.

Plus restored the load-bearing rules that got dropped:

- Cell PM: 'A SINGLE subtask flows through dev -> QA -> doc ->
  PM-merge. DON'T split into per-role subtasks.' This is exactly
  what be-pm violated in the smoke run, creating duplicate
  'branch naming subtask' / 'PR workflow subtask' / etc.
- QA + Doc: 'read the dev's journal, not just the diff' — pre-
  gateway forced this via roboco_journal_read_team; post-gateway
  the inline data exists but the agent isn't told to use it.
- Developer: 'every acceptance criterion gets a citation in the
  reflect note' — pairs with the tracing-gate change in 75b667d
  where the reflect note is treated as the addressing artifact.

* feat(foundation): bootstrap foundation/identity.py with Role/Team/RoleLevel

Phase 1 task 1 of the foundation canonicalization plan
(docs/superpowers/specs/2026-05-10-foundation-canonicalization-design.md).

Three enums, no consumers yet — separate tasks migrate the existing
forks (models.base.AgentRole, lifecycle.spec.Role, agents_config role
sets, services/permissions.PM_ROLES) onto this canonical surface.

* feat(foundation/identity): add AGENTS catalog (single source for slug->role+team+UUID)

Resolves head-marketing.team drift (spec §5.1) by setting Team.BOARD
authoritatively. Team.MARKETING remains in the enum for legacy seed
data but no agent claims it; flagged for removal in cleanup.

* feat(foundation/identity): add role-sets + ROLE_LEVEL hierarchy

* feat(foundation/identity): add lookups + public API re-exports

* feat(foundation): import-time validators (uniqueness, role coverage, role-level)

* chore(foundation): verify+align postgres agentrole/team enums with foundation/identity

scripts/verify_postgres_enums.py reads the live agentrole+team enums
from postgres (via asyncpg using roboco.config.settings.database_*)
and compares them against the foundation Role+Team enums. Exits 0 on
match, 1 on drift (with a per-side diff), and 1 with a clear message
if postgres is unreachable so callers like make foundation-check can
treat that as a skip.

alembic/versions/012_align_agentrole_team_with_foundation.py is the
forward-only safety-net migration. It runs ALTER TYPE agentrole ADD
VALUE IF NOT EXISTS 'system' (idempotent on postgres >= 9.6) so any
DB without the recently-added Role.SYSTEM sentinel gets it on next
upgrade. Postgres has no DROP VALUE primitive without a destructive
type recreation, so foundation keeps legacy values (e.g. Team.MARKETING)
to absorb the inverse direction; the migration's downgrade is
intentionally a no-op.

Local verification deferred: postgres is not reachable from this
workstation (role 'roboco' does not exist), so the script could not
confirm the live enum shape. The migration is idempotent and runs
unconditionally on the next alembic upgrade head, and whoever next
runs make foundation-check against a live DB will get the post-migration
proof of alignment.

* refactor(lifecycle): re-export Role from foundation.identity (single source)

* refactor(models): re-export AgentRole and Team from foundation.identity

Removes the parallel Team and AgentRole StrEnum definitions in
models/base.py. They are now bound to roboco.foundation.identity.Role
and roboco.foundation.identity.Team respectively, so AgentRole IS
identity.Role (same Python class object). SQLAlchemy column types
bound as sa.Enum(AgentRole, name='agentrole') continue to work because
identity is preserved across import paths.

Note: foundation.Team drops the legacy 'fullstack' member that lived
on models.base.Team. The two _resolve_team_dir tests that used
Team.FULLSTACK to exercise the 'fullstack' branch now pass the literal
string 'fullstack' instead — same code path, no enum-membership coupling.

Adds two identity assertions to tests/foundation/test_role_reexport.py
verifying AgentRole is identity.Role and Team is identity.Team.

* fix(foundation): correct Team enum — add FULLSTACK, remove QA

The original plan's audit incorrectly identified the models/base.Team
membership. Actual original was 7 values: backend, frontend, ux_ui,
fullstack, main_pm, board, marketing. My plan replaced fullstack with
qa and added system — but qa was never a team (only a role).

Postgres team enum has fullstack (alembic 009), and services/task.py:675
+ services/git.py:779 branch on the literal "fullstack". Without
foundation.Team.FULLSTACK, any Project row with assigned_cell="fullstack"
would fail to round-trip through the SQLAlchemy ORM.

This correction:
- Adds FULLSTACK; removes QA from foundation.Team
- Updates the 8-value test expected set
- Restores tests/integration/test_task_service_misc.py to use Team.FULLSTACK

* refactor(agents_config): derive AGENT_ROLE_MAP/AGENT_TEAM_MAP/CELL_MEMBERS from foundation

* refactor(roles): canonicalize role-sets via foundation.identity

- agents_config.PM_ROLES (5-role: PMs + board + CEO) renamed to
  TASK_CREATOR_ROLES; the name PM_ROLES is reserved for the canonical
  2-role set (CELL_PM + MAIN_PM) defined in foundation.identity.
- agents_config._BOARD_ROLES aliased to foundation.BOARD_ROLES (drops
  main_pm from the set; board A2A handler updated to keep allowing
  board -> main_pm direct messaging via explicit branch).
- services/permissions.PM_ROLES (2-role) re-exported from foundation.

Closes the silent semantic divergence flagged in spec section 3 (HIGH severity).

* refactor(seeds,orchestrator): derive agent catalogs from foundation

- seeds/initial_data.AGENT_UUIDS derived from foundation.AGENTS.
- DEFAULT_AGENTS row generation pulls slug+role+team+id from foundation;
  per-agent presentation strings (display name) stay in this file in
  _AGENT_PRESENTATION dict. The system sentinel remains a literal with
  team=None because the postgres `team` enum has no 'system' value.
- runtime/orchestrator._AGENT_TEAM_MAP and the cell-prefix table replaced
  with foundation.team_for_slug. _AGENT_TEAM_MAP is now a derived ClassVar
  covering every slug (not just management).
- head-marketing.team resolved to "board" (was "marketing" in seed +
  orchestrator, "board" in agents_config — three-way drift, now unified).
- ceo.team resolved to "board" (was None in seed; foundation declares
  board membership so the seed-bootstrapped DB row now reflects that).
- Adds tests/foundation/test_seed_orchestrator_parity.py — gate against
  future drift between seed/orchestrator and foundation.

Closes the identity sub-phase. Adding an agent edits exactly one file:
foundation/identity.py:AGENTS.

* feat(foundation/policy): task_completeness rules + denylist

Implements spec §5.2: field-level completeness rules at create/delegate
time, plus the denylist that catches the literal placeholder string from
the deleted services/task.py:5061-5062 silent fallback ("completed and
reviewed by assignee" — agents copy-paste this from old logs).

CompletenessSpec is data; check() is a pure function; field_hints map
gives the agent the literal answer key for each missing field.

* feat(envelope): add incomplete_input envelope kind for interrogation pattern

Sister to tracing_gap; distinct error code lets agent prompts teach
incomplete_input handling separately from tracing-gap recovery. Carries
missing + field_hints + remediate for the spec §5.2.1 interrogation
pattern; Task 19 will wire the gateway delegate verb to use it.

* feat(foundation/task_completeness): auto-fill helpers (team, priority, parent)

* feat(api/schemas): DelegateRequest enforces TASK_AT_CREATE constraints

Removes silent defaults for nature/task_type/estimated_complexity;
adds min_length=20 to description; requires non-empty acceptance_criteria.
Mirrors foundation.policy.task_completeness.TASK_AT_CREATE so under-filled
delegate calls fail at the request boundary (422) instead of being silently
papered over downstream.

Tests touching delegate calls updated to pass the now-required fields.

* feat(models/task): TaskCreate + TaskCreateRequest enforce TASK_AT_CREATE

* feat(api/schemas): TaskUpdate rejects blanking acceptance_criteria

Golden Rule preservation — acceptance_criteria cannot be set to []/None
via PATCH. Pydantic field min_length doesn't catch explicit None, so a
model_validator(mode='before') guards the patch payload.

* fix(services/task): delete silent acceptance_criteria fallback (skeleton-task root cause)

The fallback at services/task.py:5061-5062 silently replaced empty
acceptance_criteria with ['completed and reviewed by assignee'] -
the proximate cause of every skeleton task in the 2026-05-10 smoke
run. Removed; create_subtask now invokes foundation.policy.task_completeness
and raises TaskCompletenessError on missing fields (spec section 5.2).

Two existing transition tests relied on a 1-char description default
that the new completeness check rejects (description min_length=20);
both updated to pass an explicit valid description. New integration
test pins the rejection contract (empty list + legacy phrase both
raise).

Companion code path at services/gateway/choreographer/_impl.py:1852
(the upstream `or []` collapse) is fixed in the next task.

* fix(gateway/delegate): use task_completeness + Envelope.incomplete_input

Replaces the `acceptance_criteria=inputs.acceptance_criteria or []`
collapse at _impl.py:1852 with a foundation.policy.task_completeness
check that rejects empty / placeholder input via
Envelope.incomplete_input — the spec section 5.2.1 interrogation pattern.

Auto-fill helpers (fill_team_from_assignee + fill_priority_from_parent)
fill the unambiguous fields before the check, then anything still
missing surfaces as a structured rejection with field_hints; the agent
gets a literal answer key for what to provide on retry.

DelegateInputs gains an explicit `nature` field (no default) so the
HTTP boundary can thread DelegateRequest.nature through to the
choreographer. Route handlers (flow_cell_pm, flow_main_pm) forward it.
The hardcoded TaskNature.TECHNICAL fallback in _create_subtask_from_inputs
is removed; the helper now coerces inputs.nature to the enum or raises
TaskCompletenessError if a non-gateway caller bypassed the check.

Closes the gateway-side path to skeleton tasks. Service-layer raise
(Task 18) remains as defense-in-depth for non-gateway callers.

Existing delegate-guard tests updated to pass full payloads — the
prior `title='x', description='y'` minimal stubs now hit the
completeness gate first; the full payloads still exercise the
auth/chain/cap guards downstream.

* feat(api/routes/tasks): POST /tasks uses foundation.task_completeness check

Replace the hand-rolled acceptance_criteria non-empty check in the
POST /tasks handler with a call to task_completeness.check(TASK_AT_CREATE,
data). Route, schema (TaskCreate), and service (TaskCreateRequest) now
all share one canonical notion of 'complete' — the fourth and final
create path is now strict.

Pydantic still rejects structurally invalid payloads (empty AC list,
short title/description, missing enums) with 422. The TC check at the
route boundary additionally rejects denylisted placeholder phrases
('completed and reviewed by assignee', etc.) that pass schema validation
but signal a stub task.

Add tests/integration/test_post_tasks_completeness.py:
  - empty acceptance_criteria  -> 422 (Pydantic)
  - placeholder phrase         -> 400/422 with 'acceptance_criteria' in body

* chore(make): add foundation-check drift gate (mirrors lifecycle-check)

* test(foundation): Phase 1 smoke gate — skeleton-task path returns incomplete_input

Phase 1 closes here: identity catalogs are single-sourced; the silent
acceptance_criteria fallback is gone; gateway delegate returns
incomplete_input with populated field_hints when criteria are missing.

The 2026-05-10 smoke run that produced skeleton tasks no longer can.
Phases 2-4 (tracing, journaling, communications, agent_loop, housekeeping)
get their own plans.

* feat(foundation/policy): journaling scope catalog (5 panel-UI scopes)

* feat(foundation/policy/journaling): role read tiers + protected journals

* refactor(content_actions): derive _VALID_NOTE_SCOPES from foundation.journaling

* refactor(services/journal): derive _SCOPE_TO_TYPE from foundation.journaling

* refactor(enforcement/journal_perms): import read-tier rules from foundation

PROTECTED_JOURNALS + ROLE_READ_TIERS now sourced from foundation.policy.journaling.
The local helpers (_check_protected_access, _check_cell_pm_access,
_check_cell_member_access) are collapsed into a single tier-driven check via
_decide_protected / _decide_by_tier. Pre-Phase-2 GLOBAL_READERS that lumped
CEO/auditor/PO/HoM/main_pm together is split into ReadTier.ALL (ceo+auditor —
includes protected) vs ReadTier.ALL_CELLS (others — excludes protected).
Observable behavior preserved.

* feat(foundation/policy): tracing Requirement enum + check_requirements

19 requirements (16 from pre-Phase-2 tracing_gate + 3 pre-gateway parity:
JOURNAL_NOTE_AT_CLAIM, JOURNAL_DECISION_AT_CLAIM, JOURNAL_DURING_WORK).
GateContext expanded with the new presence flags and journal_during_work_count.
Acceptance-criteria checker keeps the spec §9 item 1 reflect-note shortcut.

* feat(foundation/policy/tracing): VERB_REQUIREMENTS table + verb parity validator

Maps every gateway intent verb to its required-set. Includes the 6 inline
journal:decision callsites (submit_up, complete, unblock, escalate_up,
escalate_to_ceo, delegate) plus the 4 pre-gateway parity additions
(NOTE_AT_CLAIM, DECISION_AT_CLAIM, REFLECT on complete, DURING_WORK).
Validator asserts every spec verb is covered or explicitly waived, and
every Requirement enum value is used by at least one verb.

PLAN added to i_will_work_on / i_will_plan (mirrors spec.PRECONDITION_PLAN
in the tracing layer). SELF_VERIFIED added to i_am_done as a defense-in-depth
backstop (auto-set by the in_progress→verifying transition).

* refactor(gateway/i_am_done): tracing gates via foundation.policy.tracing

Adds JOURNAL_DURING_WORK_AT_LEAST_ONE check (pre-gateway parity P2 —
agents must write at least one decision/learning/struggle entry between
claim and submit). Adds journal.has_struggle_for_task helper.
Replaces the pre-Phase-2 tracing_gate.check_requirements call.

SELF_VERIFIED is filtered from the pre-flight required-set: the spec
composes (submit_verification, submit_qa) for i_am_done and the
auto-run submit_verification flips self_verified=True before submit_qa
runs. The flag therefore acts as a defense-in-depth backstop AFTER the
spec, not before — checking it pre-flight would block the auto-verify
path. SELF_VERIFIED stays in the foundation required-set and is
re-asserted by the spec action's own preconditions.

Test fixtures updated: 9 i_am_done success-path tests now mock
has_decision_for_task=True (or equivalent) so the new during-work
cadence gate is satisfied. NO_PR-token assertion broadened to also
accept the foundation token "pr_open".

* refactor(gateway/qa): pass/fail gates via foundation.policy.tracing

* refactor(gateway/doc): i_documented gates via foundation.policy.tracing

Doc-specific missing-key translations (docs_notes>=min, docs_files_non_empty)
added to the central _build_tracing_gap translator established in Task 9.

* refactor(gateway): unify 6 inline journal:decision checks via tracing.check_requirements

Pre-Phase-2 inline blocks at _impl.py lines ~2230/2394/2442/2574/2814/2895
each ran the same has_decision_for_task + Envelope.tracing_gap pattern. They
now call:

- _check_pm_decision_required(verb, ...) — for unblock, escalate_up,
  escalate_to_ceo, delegate. Each declares only JOURNAL_DECISION in
  VERB_REQUIREMENTS, so a single helper consuming
  tracing.requirements_for(verb) suffices.
- _check_complete_gates — for cell_pm_complete and main_pm_complete.
  Consumes VERB_REQUIREMENTS["complete"] = JOURNAL_DECISION + JOURNAL_REFLECT
  + NOTES_MIN_CHARS. The inline _subtasks_not_terminal_envelope is kept
  because its remediation enumerates the non-terminal subtask ids — strictly
  richer than the foundation hint.
- _check_submit_up_gates — for submit_up. Consumes
  VERB_REQUIREMENTS["submit_up"] minus SUBTASKS_TERMINAL (deferred to the
  inline envelope for the same reason as complete).

Also adds the journal:decision tracing gate to the delegate verb
(VERB_REQUIREMENTS["delegate"] = {JOURNAL_DECISION}) — pre-gateway PM.md
required journal:decision before each delegate, but the gateway path had
not yet enforced it. Threaded into _delegate_extra_guards so the verb
body's return count stays under the lint cap.

_build_tracing_gap gains hint translations for journal:decision, notes>=min,
and subtasks_terminal. The body is refactored to a static dispatch table
+ acceptance-criteria batch handler so the branch count stays under the
lint cap.

PM-verb success-path tests updated to provide notes >= 20 chars (the new
NOTES_MIN_CHARS gate); has_reflect_for_task mocks added to a few tests
where they're now load-bearing (AsyncMock truthiness covers most).
6575 tests passing, mypy + ruff clean.

* feat(gateway/claim): require journal:note_at_claim and journal:decision_at_claim

Pre-gateway parity P1, P3: developers wrote a note (scope='note') on
every claim; PMs wrote a decision (scope='decision') on plan. Restored
via foundation.policy.tracing requirements wired through a new
_post_claim_journal_gate helper that runs AFTER the composed
(claim, set_plan, start) sequence completes.

Failed checks return tracing_gap with a remediate hint that tells the
agent to journal then retry. The claim itself stays — the agent
journals and re-issues the verb (idempotent re-entry shortcuts back
to OK once the entry is present).

Adds journal.has_note_for_task helper paralleling
has_decision/reflect/learning/struggle. The PLAN requirement is
filtered out of the post-claim check because spec.PRECONDITION_PLAN
already enforced it before the runner ran — re-asserting at the
tracing layer would emit a misleading hint.

Two new tests verify the gate fires for missing note/decision; existing
success-path tests already mock the journal service via AsyncMock
(returning truthy) so no regressions.

* test(foundation): Phase 2 smoke gate + tracing_gate.py deleted

Phase 2 closes here:
- foundation/policy/journaling.py owns the 5-scope catalog + read tiers
- foundation/policy/tracing.py owns Requirement enum + VERB_REQUIREMENTS
- 6 inline journal:decision checks replaced with unified helpers
- pre-gateway parity restored: NOTE_AT_CLAIM, DECISION_AT_CLAIM,
  DURING_WORK, REFLECT-on-complete
- services/gateway/tracing_gate.py deleted
- enforcement/journal_perms.py read-tier rules canonicalized

Smoke gate 2 enforces: no inline has_decision_for_task remains; every
intent verb has a tracing decision; tracing_gate module is gone.

* fix(foundation/task_completeness): align hint strings with actual enum values

_HINT_NATURE listed 5 values (technical | bugfix | feature | refactor | docs)
but TaskNature only has 2 (TECHNICAL / NON_TECHNICAL). _HINT_ESTIMATED_COMPLEXITY
listed "critical" which Complexity doesn't have. _HINT_TEAM omitted FULLSTACK
(real, used) and didn't note that MARKETING is legacy seed-data. _HINT_TASK_TYPE
was already correct.

Hints now reflect the actual enums in roboco/models/base.py and
roboco/foundation/identity.py — agents reading the gateway's incomplete_input
remediate envelopes will no longer be told to send values the enums reject.

Tests using nature="feature" (DelegateRequest's nature is `str`, not the
enum, so it accepted the fake value silently) updated to nature="technical"
so they exercise a real enum value end-to-end.

* fix(orchestrator): remove dead "critical" complexity branches

Complexity enum has only LOW / MEDIUM / HIGH — no CRITICAL value.
The three "critical" branches in dispatch logic at lines ~3032 / 3331 /
5049 were dead code (the comparison can never be true). Removed.

Surfaced during Phase 2 closeout when the foundation hint string was
audited against the actual enum.

* feat(foundation/policy/communications): Priority + NOTIFY_SENDER_ROLES + ACK_REQUIRED_BY_TYPE

* feat(foundation/policy/communications): CHANNELS catalog (channel topology)

* refactor(agents_config): derive CHANNEL_ACCESS from foundation.communications

* refactor(seeds): derive DEFAULT_CHANNELS / CHANNEL_MEMBERSHIPS from foundation

* refactor(content_actions): derive notify allowlist + priorities from foundation

Replaces _NOTIFY_ALLOWED_ROLES + _VALID_NOTIFY_PRIORITIES literals with
derivations from foundation.communications.NOTIFY_SENDER_ROLES + Priority.

Behavior change: pre-Phase-3 the literal frozenset {cell_pm, main_pm,
product_owner, head_marketing} excluded CEO. Foundation includes CEO
(per spec 5.5). The contradiction with agents_config.NOTIFICATION_PERMISSIONS
(which already granted CEO can_send=True) is now resolved.

* refactor(notification_delivery): requires_ack from foundation.ACK_REQUIRED_BY_TYPE

* refactor(enforcement,agents_config): delete dead notification policy

- enforcement/notification_perms.py deleted (dead at call-graph; only
  the enforcement/__init__.py re-export kept it reachable, and that
  re-export is gone too).
- agents_config.NOTIFICATION_PERMISSIONS dict deleted; agents_config
  .can_send_notifications now derives from
  foundation.policy.communications.NOTIFY_SENDER_ROLES (auditor
  correctly excluded — silent observer per spec §5.5).
- services/permissions.py: _can_role_send_notifications and
  can_agent_send_notifications now derive from NOTIFY_SENDER_ROLES;
  _get_notification_scope encodes the scope rule (cell/all/list)
  locally as a function-of-role and returns list[AgentRole] instead
  of list[slug]; can_notify list-scope branch updated to match.
- enforcement/__init__.py: removed the notification_perms re-export
  and the NotificationPermissionError, get_notification_scope,
  validate_notification_permission names from __all__.

Closes the spec §3 contradiction: gateway content_actions
._NOTIFY_ALLOWED_ROLES (Task 5) and the legacy
agents_config.NOTIFICATION_PERMISSIONS no longer disagree about
whether auditor may call notify(). Both now derive from
foundation.NOTIFY_SENDER_ROLES.

* fix(content_actions): runtime auditor guard in say/dm (defense in depth)

Closes the spec §5.5 gap where the auditor's silent role was enforced
ONLY by manifest exclusion. The manifest pre-filters the tool surface
exposed to the auditor agent, but if anything bypassed it, the auditor
could speak. The new runtime guard in ContentActions.say/dm refuses
with Envelope.not_authorized when the caller's role is "auditor",
regardless of how the call arrived.

* fix(a2a): pass Priority tristate end-to-end (was reduced to boolean)

Pre-Phase-3 path:
  request priority: str -> services/a2a.py reduces to urgent: bool
  -> services/notification.py maps bool back to NotificationPriority
This made Priority.HIGH unreachable through the A2A path.

After this fix the full tristate (NORMAL/HIGH/URGENT) survives end-to-end:

  * services/a2a.py:create_a2a_notification parses metadata["priority"]
    (preferred) or falls back to legacy metadata["urgent"] / config.urgent
    (URGENT-only). Unknown values fall back to NORMAL.
  * services/notification.py:send_a2a_notification now takes
    a2a_context["priority"] (NotificationPriority); a defensive bool/str
    coerce keeps legacy callers from crashing.
  * runtime/orchestrator.py:_build_a2a_prompt reads priority off the
    notification row (the source of truth) instead of a non-existent
    metadata.urgent and renders three tiers: URGENT bold, HIGH softer,
    NORMAL no prefix.

Cosmetic [URGENT] body/subject prefix stays urgent-only; HIGH gets no
prefix but is recorded as HIGH at the NotificationTable.priority column.

Tests:
  * 9 new tests in tests/integration/test_a2a_priority_tristate.py
    pinning the round-trip for HIGH/NORMAL/URGENT through both layers
    plus legacy-bool backcompat.
  * Updated tests/unit/services/test_notification.py::test_send_a2a_notification
    to the new priority= contract.

Closes the spec section 3 contradiction flagged in the audit.

* feat(foundation/policy): agent_loop BudgetPolicy + VERB_RETRY_LIMITS

* refactor(agent_sdk): import budget thresholds from foundation

* refactor(orchestrator): import _PM_RESPAWN_MAX_UNPRODUCTIVE from foundation

* fix(post-tool-budget-hook): exit 1 on loop-halt (was exit 0 / non-blocking)

Pre-Phase-3 the hook printed [Loop] and exit 0'd — agents could ignore it
and keep retrying. The 2026-05-10 smoke run showed i_am_done retried 5+
times within the global 150-tool budget, never hitting a real wall.

Now the hook reads the SDK response's loop_action field (sourced from
foundation.BudgetPolicy.loop_action; default "halt") and exits 1 to
deny the wrapping tool call when the rolling-window loop detector fires
AND loop_action is "halt". Operators can soften via env
ROBOCO_AGENT_LOOP_ACTION=warn for debugging.

Changes:
- BudgetStatus pydantic model: add loop_action: Literal["warn", "halt"]
  (default "halt") so the SDK response carries the policy.
- agent_sdk/server.py: read ROBOCO_AGENT_LOOP_ACTION env override on top
  of foundation default and surface it in _budget_snapshot().
- post-tool-budget-hook.sh: parse .loop_action, exit 1 to stderr when
  loop+halt; falls back to legacy warn-only print if the field is
  missing (older SDK / partial deploy).

* feat(agent_sdk): per-verb retry circuit breaker via foundation.VERB_RETRY_LIMITS

Pre-Phase-3 the gateway had no per-verb retry cap. The 2026-05-10 smoke
showed i_am_done retried 5+ times in 2 minutes within the global 150-tool
budget — the agent never hit a real wall.

Now the SDK tracks (verb, task_id) -> deque[timestamp] over a 60s sliding
window. When the count for a verb exceeds foundation.retry_limit_for(verb),
the next attempt receives Envelope.circuit_open with a remediate hint
pointing to i_am_blocked / i_am_idle as graceful exits.

Verbs in foundation.UNLIMITED_RETRY_VERBS (give_me_work, triage,
evidence, etc.) bypass the breaker. Only rejection envelopes
(tracing_gap, invalid_state, not_authorized, incomplete_input) feed
the counter — successful calls do not count.

Wire-up:
- Envelope.circuit_open classmethod + as_dict pass-through
- _SessionState.verb_attempts: defaultdict[(verb, task_id), deque[float]]
- Helpers _record_verb_attempt / _verb_attempt_count / _check_verb_circuit
- POST /verb/attempted: hook posts after a rejected gateway call;
  response carries breaker state + (when open) the wire-format
  Envelope.circuit_open dict the agent should surface to itself
- GET /verb/circuit_status: read-only state probe
- _state.reset() (also POST /budget/reset) wipes the tracker on spawn

* test(foundation): Phase 3 smoke gate + foundation-check extended

Phase 3 closes here:
- foundation/policy/communications.py owns Priority, NOTIFY_SENDER_ROLES,
  ACK_REQUIRED_BY_TYPE, ChannelSpec, CHANNELS, parse_priority
- foundation/policy/agent_loop.py owns BudgetPolicy, VERB_RETRY_LIMITS,
  UNLIMITED_RETRY_VERBS, retry_limit_for
- 6 channel topology fork sites collapsed to one source (CHANNELS)
- Notification sender contradiction closed (CEO included; auditor excluded)
- A2A urgency tristate restored (HIGH reachable end-to-end); A2A
  service now consumes parse_priority instead of inlining branches
  (also drops create_a2a_notification CC from C/13 to A/<10)
- Auditor silent role enforced at runtime in say/dm
- enforcement/notification_perms.py deleted (was dead code)
- 7 hand-set requires_ack callsites consolidated to ACK_REQUIRED_BY_TYPE
- post-tool-budget-hook.sh exits 1 on loop-halt
- Per-verb retry circuit breaker live in agent_sdk (60s sliding window)

make foundation-check now validates communications + tracing + journaling
+ identity drift in one command. make quality green.

* refactor(lifecycle): copy spec.py to foundation/policy/lifecycle.py + shim

Phase 4 Task 1 — relocates the canonical lifecycle spec next to its policy
siblings (task_completeness, tracing, journaling, communications, agent_loop).

The original roboco/lifecycle/spec.py is now an explicit re-export shim;
consumers continue to work unchanged. Subsequent Phase 4 tasks (2-7) migrate
the imports in batches, then Task 8 deletes the shim.

No behavior change — pure code move.

* refactor(services): import lifecycle from foundation (Phase 4 batch)

* refactor(agents,enforcement): import lifecycle from foundation (Phase 4 batch)

* refactor(tests): import lifecycle from foundation (Phase 4 batch)

* refactor(foundation): absorb lifecycle _validate + _generators

Phase 4 Tasks 9 + 10. Moves the lifecycle spec's internal validators
to roboco/foundation/_validate_lifecycle.py and its RAG/prompt artifact
emitter to roboco/foundation/_generators.py.

The lifecycle validators live in a sibling module (not merged with
foundation/_validate.py) because the lifecycle spec imports from
foundation at module load — placing the lifecycle checks alongside the
identity checks would create an import cycle between
roboco.foundation and roboco.foundation.policy.lifecycle (the latter
calls the validators at the bottom of its own definition). The
_validate_lifecycle module defers its policy.lifecycle imports to
function bodies so it loads cleanly when the spec hasn't finished
initialising yet; the per-file PLC0415 exemption in pyproject.toml
documents the reason.

Test files relocated:
- tests/lifecycle/test_spec.py        -> tests/foundation/test_lifecycle_spec.py
- tests/lifecycle/test_generators.py  -> tests/foundation/test_lifecycle_generators.py

scripts/build_lifecycle_artifacts.py now imports the generators from
roboco.foundation; the on-disk artifacts (docs/rag/lifecycle,
panel/lib/lifecycle.json, agents/prompts/_generated/lifecycle-*.md)
regenerate byte-identically.

After this commit, roboco/lifecycle/ contains only the spec.py and
__init__.py re-export shims — Task 8 deletes those.

No behavior change. 6638 tests pass; make quality green.

* refactor(lifecycle): delete legacy roboco/lifecycle/ package

Phase 4 Task 8. All consumers migrated to roboco.foundation.policy.lifecycle
in Tasks 2-7; the internal validators + generators moved to foundation in
Tasks 9-10. The legacy package contained only re-export shims.

Also trims tests/foundation/test_role_reexport.py — the two assertions that
checked the lifecycle.spec shim's object-identity are gone with the shim.
The two models.base shim assertions (AgentRole / Team) are still
meaningful and stay.

Inline docstrings / comments in enforcement/task_lifecycle.py,
services/gateway/role_config.py, services/gateway/content_actions.py,
tests/integration/test_task_service_lifecycle_misc.py and
foundation/policy/lifecycle.py that referenced the now-deleted
roboco.lifecycle.spec module are updated to point at
roboco.foundation.policy.lifecycle.

After this commit, roboco.lifecycle is gone. Lifecycle policy lives only
at roboco.foundation.policy.lifecycle. Adding new lifecycle rules edits
exactly that one file.

* refactor(api): consolidate route-guard role-sets via foundation

Replace hand-written role-name string frozensets in roboco/api/deps.py
(_PM_OR_ABOVE_ROLES, _DEVELOPER_OR_ABOVE_ROLES, _GLOBAL_CELL_ACCESS_ROLES)
and roboco/api/routes/v2/_role_dep.py (require_dev/qa/doc/cell_pm/main_pm/
board/auditor) with foundation-derived expressions over PM_ROLES,
BOARD_ROLES, DEV_ROLES, and Role enum members.

Behavior is preserved: Role is a StrEnum, so the lowercase X-Agent-Role
header still compares equal to its matching member. HEAD_MARKETING stays
excluded from every -or-above set (marketing spokesperson, not approver);
the carve-out is now expressed as (BOARD_ROLES - {Role.HEAD_MARKETING})
instead of an opaque literal.

Adds tests/foundation/test_route_guard_consolidation.py (6 tests) pinning
both the foundation-derived membership and the import contract.

* test(foundation): Phase 4 smoke gate + housekeeping closeout

Phase 4 closes the foundation canonicalization effort (Phases 1-4 spanning
2026-05-10 -> 2026-05-11):

Phase 1 - identity + task_completeness (skeleton-task bug killed)
Phase 2 - tracing + journaling (pre-gateway cadence restored)
Phase 3 - communications + agent_loop (channel/notification/A2A/circuit-breaker)
Phase 4 - housekeeping (lifecycle moved to foundation; consumers migrated)

All cross-cutting policy now lives in roboco/foundation/. Adding a policy
edits exactly one file. The legacy roboco.lifecycle package is gone.
Smoke gates 1-4 enforce: no skeleton tasks, no inline journal:decision
checks, channel topology canonical, A2A tristate preserved, auditor silent
at runtime, lifecycle module path canonical.

make quality + make foundation-check both green.

* fix(mcp/agent_sdk): wire per-verb circuit breaker into response handler

Phase 3 Task 14 added the SDK infrastructure (tracker, endpoints,
Envelope.circuit_open, retry_limit_for) but nothing was actually
recording rejections — the breaker never tripped. This commit wires
the gateway-response path so every rejection envelope (tracing_gap /
invalid_state / not_authorized / incomplete_input) hits
POST /verb/attempted, and if the breaker is open, the envelope is
substituted with the circuit_open response before the agent sees it.

Best-effort: SDK-unreachable / malformed-response failures fall open
(agent sees the original rejection), so the breaker never breaks the
gateway path.

* fix(notification_delivery): retype CEO approval-flow notifications APPROVAL

notify_assignee_of_ceo_rejection and notify_ceo_of_escalation were both
typed NotificationType.TASK_ASSIGNMENT, which the Phase 3 foundation
table (ACK_REQUIRED_BY_TYPE in roboco/foundation/policy/communications.py)
maps to requires_ack=False. Both are approval-flow notifications and
should mandate acknowledgment.

Retyped both to NotificationType.APPROVAL so the table lookup yields
requires_ack=True via ACK_REQUIRED_BY_TYPE[NotificationType.APPROVAL].

* test(foundation): move lifecycle parity + smoke-replay tests under tests/foundation/

Phase 4 Task 8 deleted roboco/lifecycle/ but tests/lifecycle/ still held
two files importing roboco.foundation.policy.lifecycle. Mirror the layout
of test_lifecycle_spec.py and test_lifecycle_generators.py (moved in
Phase 4 Tasks 9+10) by relocating them under tests/foundation/ with the
test_lifecycle_* prefix, then delete the now-empty tests/lifecycle/
package.

  tests/lifecycle/test_consumer_parity.py
    -> tests/foundation/test_lifecycle_consumer_parity.py
  tests/lifecycle/test_smoke_replay.py
    -> tests/foundation/test_lifecycle_smoke_replay.py

* build(make): consolidate ci-lifecycle-check into foundation-check

ci-lifecycle-check was a thin wrapper that regenerated lifecycle artifacts
via scripts/build_lifecycle_artifacts.py and gated on git diff. After
Phase 4 it sat alongside foundation-check covering the same drift-gate
intent. Merge the lifecycle-artifact regen + git-diff step into
foundation-check so a single 'make foundation-check' is the canonical
drift gate.

Keep ci-lifecycle-check as a phony alias forwarding to foundation-check
for any external script or CI lane still using the old target name.
Drop the redundant ci-lifecycle-check call from 'make quality'.

* ++

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-05-11 02:15:47 +02:00
62bda0c497 Gateway/full (#9)
* chore(gateway): scaffold gateway package and test layout

* feat(config): add gateway feature flags, coordination thresholds, commit-validator settings

* feat(gateway): add standardized response envelope with ok/error variants

* feat(gateway): add remediation hint catalog for tracing-gap and invalid-state errors

* feat(gateway): add per-role flow/do tool catalog with developer, qa, doc, pm, board configs

* feat(db): add gateway columns — active_claimant_id, heartbeat, pre_block snapshot, acceptance_criteria_status, qa_evidence_inspected

* feat(db): create gateway_triggers table for dispatcher decision logging

* feat(db): align canonical skill set; substitute qa_review -> code_review across agent seeds

* fix(db/008): make skill alignment in-place + idempotent; preserve column and existing custom skills

* feat(gateway): add claimant_lock for single-active-agent invariant with heartbeat staleness

* feat(gateway): add trigger_filter with stale-cleanup, claimant-queue, and cooldown rules

* feat(gateway): add tracing_gate with plan, progress, journal, acceptance_criteria, qa requirements

* refactor(gateway): drop per-file ruff ignores; refactor tracing_gate with dispatch table + GateContext

* feat(gateway): add merge_chain to resolve PR target by branch hierarchy depth

* feat(gateway): add commit_validator with min-length, banned-words, and conventional-shape hints

* feat(gateway): add evidence_builder for verb-response evidence and capped context_briefing

* feat(gateway): add Choreographer skeleton with per-phase verb signatures and DI protocols

* feat(runtime): add spawn_manifest builder for per-role pre-loaded tool registration

Introduces SpawnInputs dataclass + build_for_role(inputs) + write_manifest()
in roboco/runtime/spawn_manifest.py; reads role_config for allowed verbs/tools,
emits JSON manifest that SDK shim reads at container startup to eliminate ToolSearch.

* feat(runtime): wire gateway pre-spawn check (trigger_filter + claimant_lock) into orchestrator behind ROBOCO_GATEWAY_ENABLED flag

- Add GatewayTriggerTable SQLAlchemy ORM model to roboco/db/tables.py
  (matches existing table from migration 007_gateway_triggers_table)
- Add module-level gateway_pre_spawn_check() + helpers to orchestrator.py
  (gated: returns ("spawn", "gateway disabled") immediately when flag is False)
- Wire gateway check into _safe_spawn() — the single dispatcher choke-point
  for all agent spawns; QUEUE or DROP outcome logs and returns None (no spawn)
- ROBOCO_GATEWAY_ENABLED defaults to False; legacy behaviour is unchanged

* feat(agent_sdk): load tool-manifest.json at startup behind ROBOCO_GATEWAY_ENABLED flag (no agent-visible change yet)

Adds load_tool_manifest() to the SDK server that reads env at call-time
so gateway-enabled agents can obtain their pre-registered tool list at
startup; returns None when the flag is off, leaving the legacy briefing
path completely unchanged.

* fix(optimal_brain): skip indexing when source ID is None to eliminate roboco://journals/None spam

- Add `build_doc_source(kind, id_)` module-level helper in indexes/base.py that
  returns None when id_ is None instead of producing a "roboco://journals/None" URI
- Update abstract `build_source_uri` return type to `str | None` so subclasses
  can legitimately signal a missing ID
- Short-circuit `ingest()` and `_prepare_docs_for_batch()` in BaseIndexPlugin
  when `build_source_uri` returns None (debug log, no push to vector store)
- Fix JournalsIndexPlugin.build_source_uri: `kwargs.get("entry_id")` returns the
  kwarg value even when it is None, so fall back to doc_id before calling
  build_doc_source
- Fix ConversationsIndexPlugin.build_source_uri: return None when session_id is
  None rather than producing "roboco://conversations/None-unknown"
- Add 9 unit tests with a piragi-free conftest that stubs sys.modules

* fix(agent_sdk): inject X-Agent-ID header on notification-poller requests

Both `_check_pending_a2a` and `_auto_ack_a2a_notifications` in
`roboco/mcp/a2a_server.py` were calling the main API without identity
headers, causing orchestrator `Missing X-Agent-ID header` warnings on
`GET /api/v1/notifications/pending-a2a` and the ack-a2a POST.

Add module-level `AGENT_ROLE` constant (mirrors the existing `AGENT_ID`
pattern) and pass `{"X-Agent-ID": AGENT_ID, "X-Agent-Role": AGENT_ROLE}`
on both requests.

* fix(git): use ROBOCO_PUBLIC_BASE_URL for commit-trailer Links instead of hardcoded localhost

* fix(test_runner): call uv run pytest/ruff directly; add make to orchestrator Dockerfile as backstop

FileNotFoundError was propagating as a raw 500 when a project had `make test`
configured but make was not installed in the orchestrator container.

Two fixes:
1. Catch FileNotFoundError in _run_command and re-raise as ValidationError (400)
   with a clear message telling the operator to reconfigure the project command
   (e.g. replace 'make test' with 'uv run pytest').
2. Add `make` to the orchestrator Dockerfile runner-stage apt-get so projects
   that legitimately use make targets continue to work without reconfiguration.

* fix(api/git): resolve project by slug or UUID in git_log endpoint

Add _resolve_project_slug() helper to git routes that tries UUID
lookup first and falls back to slug, matching the pattern already
used in project routes. Apply to all four read-only git endpoints:
status, log, branches, diff.

* fix(a2a): auto-create conversation when conversation_id absent; reject empty IDs in URL builder

* fix(agent_sdk): default subagent model to parent agent's model from spawn manifest, not hardcoded haiku

Inject CLAUDE_CODE_SUBAGENT_MODEL env var into every agent container at
spawn time.  Claude Code ≥2.1.x reads this variable to override the
default Task (Agent) subagent model, which otherwise hard-codes
claude-haiku-4-5-20251001.  When the parent runs on a non-Anthropic
provider (e.g. Ollama Cloud / minimax-m2.7:cloud) that Anthropic model
is unreachable, so subagent dispatch fails.

The value follows the same provider-aware translation already used for
the --model CLI flag: Anthropic short names go through MODEL_MAP, and
non-Anthropic identifiers are passed verbatim.  To avoid calling the
class by name inside a @staticmethod, the shared translation logic is
extracted to the module-level _resolve_agent_cli_model() helper;
_resolve_cli_model() now delegates to it.

Verified: CLAUDE_CODE_SUBAGENT_MODEL is present and honoured in the
Claude Code 2.1.123 binary (grep confirmed the env-var lookup pattern
`if(process.env.CLAUDE_CODE_SUBAGENT_MODEL) return KK(…)`).

* chore(makefile): add quality and quality-fast targets composing every PR gate

* chore(quality): add import-linter dependency and gateway boundary contract

* test(property): scaffold tracing-completeness assertion (filled in Phase 4)

* Format test file to pass ruff check

* fix(gateway): drop Protocol scaffolding from choreographer skeleton; del-statements on unused stub args; clear vulture whitelist

* linting

* feat(gateway): Phase 1 dev cutover — ChoreographerDeps + give_me_work

Add ChoreographerDeps frozen dataclass (7 deps: task, work_session, git,
a2a, journal, audit, evidence_repo), refactor Choreographer.__init__ to
accept the bundle, implement give_me_work + _briefing_for via
evidence_builder.build_context_briefing, and add property accessors for
all deps. All Phase 2-4 stubs gain del-statements and still raise
NotImplementedError. 3 tests added and passing; mypy/ruff/vulture clean.

* feat(gateway): implement i_will_work_on handling pending, claimed, and needs_revision recovery

* feat(gateway): implement i_have_committed with plan-required precondition

Replaces the NotImplementedError stub with the real implementation: looks up
the agent's active task, enforces plan presence before recording, calls
task.add_progress, and returns a structured Envelope. Adds 3 unit tests
(records progress, no active task → invalid_state, no plan → tracing_gap).

* feat(gateway): implement i_am_done with smart catch-up and skill resolution

* feat(gateway): implement i_am_blocked (struggle + escalate) and i_am_idle (with unread soft-block)

* feat(gateway): add ContentActions for commit, note, say, dm, evidence with auto-inject and validation

* feat(api/v2): add /api/v2/flow/dev/* endpoints delegating to Choreographer

Six intent-verb endpoints (give_me_work, i_will_work_on, i_have_committed,
i_am_done, i_am_blocked, i_am_idle) under /api/v2/flow/dev/, each a thin
handler that delegates to Choreographer. Includes Pydantic request schemas,
EvidenceRepo Phase 1 stub (all methods return []), get_choreographer FastAPI
dep wired with all 7 service deps, and 8 unit tests (all passing).

* feat(api/v2): add /api/v2/do/* endpoints for commit, note, say, dm, evidence

* feat(mcp): add roboco-flow MCP server for intent verbs (Phase 1: dev verbs implemented)

* feat(mcp): add roboco-do MCP server for smart-wrapped content tools

* feat(runtime): mount per-agent tool-manifest.json on developer-container spawn; gateway flag enabled for devs only

* docs(prompts): rewrite developer role prompt for gateway-only verbs (~15 lines vs 49)

* chore(mcp): confirm dev manifest excludes legacy task/journal/notify/a2a tools (Phase 1 cutover; servers retired in Phase 4)

* feat(gateway): implement claim_review with inline evidence (kills #15) and qa_evidence_inspected tracking

* feat(gateway): implement pass_review with qa_notes/learning/evidence tracing gates

* feat(gateway): implement fail_review with issue list, tracing gates, and dev A2A handoff

* feat(api/v2): add /api/v2/flow/qa/* endpoints (claim_review, pass, fail, give_me_work, i_am_idle)

* feat(mcp): add QA verbs (claim_review, pass, fail) to roboco-flow MCP server

* docs(prompts): rewrite QA role prompt for gateway verbs; explicitly warn against grep-the-commit anti-pattern

* feat(runtime): enable gateway flag for QA-role spawns (Phase 2 cutover)

* feat(gateway): implement claim_doc_task and i_documented with file-list and notes-min-chars gates

* feat(gateway): implement triage (cell PM) and triage_all (main PM) with priority order

* feat(gateway): implement unblock with pre_block_state restoration (kills #23)

* feat(gateway): implement cell_pm_complete with auto-merge to parent branch (kills #22 for cell scope)

* feat(gateway): implement main_pm_complete (open master PR + escalate to CEO)

* feat(gateway): add complete() dispatcher routing to cell_pm_complete or main_pm_complete by role

* feat(gateway): implement escalate_up routing by role.escalation_target

* feat(api/v2): add /api/v2/flow/{documenter,cell_pm,main_pm}/* endpoints

* feat(mcp): add Doc + PM verbs to roboco-flow MCP server (claim_doc_task, i_documented, triage, triage_all, unblock, complete, escalate_up)

* docs(prompts): rewrite Doc, Cell PM, and Main PM role prompts for gateway verbs

* feat(runtime): enable gateway flag for Doc, Cell PM, and Main PM roles (Phase 3 cutover)

* test(integration): full pending->awaiting_ceo_approval test through dev/QA/doc/cell-PM/main-PM gateway path

* chore(tests): rename unused args to _args in flow_server tests

Cleared RUF059 lint blocker for Phase 3 closeout. The destructured args was only consumed in URL-asserting tests; one variant only checks kwargs["json"], so its args is now _args.

* chore: untrack docs/superpowers/ + add to .gitignore

Plans + spec were inadvertently swept into commits 5d41a4b and de0c5b5 by subagent 'git add -A' calls. Removed from index and gitignored going forward; files remain on disk for ongoing reference. They still exist in history of those two commits — invoke a follow-up filter-repo if a full purge is desired.

* feat(gateway): implement Board escalate_to_ceo with role allow-list

Allows main_pm, product_owner, and head_marketing to escalate tasks to
CEO. Enforces awaiting_pm_review state and journal:decision tracing gate.
Closes Phase 4 Task 1.

* feat(gateway): implement board_triage prioritizing strategic root tasks

Adds Choreographer.board_triage and TaskService.list_strategic_for_board.
PO and Head Marketing get curated lists of strategic-nature root tasks
in awaiting_pm_review. Closes Phase 4 Task 2.

* feat(gateway): implement auditor_triage surfacing long-running blocked-task anomalies

Adds Choreographer.auditor_triage and TaskService.list_long_running_blocked.
The Auditor surfaces tasks blocked >30min as anomalies for reflect-note
observation. Closes Phase 4 Task 3.

* chore(tests): add return + arg type annotations to gateway tests

All gateway test functions now have -> None and parameter annotations. Cleared 63 mypy [no-untyped-def] errors that pre-existed since Phase 1. Mypy now clean across tests/unit/gateway/.

* feat(api/v2): add /api/v2/flow/{board,auditor}/* endpoints

Board: triage, escalate_to_ceo, i_am_idle.
Auditor: triage, i_am_idle (read-only role).
Adds EscalateToCeoRequest schema with reason min_length validation.
Closes Phase 4 Task 4.

* feat(mcp): add Board + Auditor verbs to roboco-flow MCP server

Adds escalate_to_ceo MCP tool used by Board (PO + Head Marketing) and Main PM. Updates the implemented set in _validate_role_compatibility. Auditor uses the existing triage tool with role-routing in URL.

Closes Phase 4 Task 5.

* docs(prompts): rewrite Board (PO, Head-Marketing, Auditor) prompts for gateway verbs

All 3 board identity files + roles/board.md now use the slim, gateway-aware shape (no ToolSearch directive, no state-tool table). Auditor is explicit about its read-only scope. Closes Phase 4 Task 6.

* feat(runtime): enable gateway manifest for ALL roles (Phase 4 cutover)

Adds product_owner, head_marketing, auditor to GATEWAY_ENABLED_ROLES. Every spawned agent now gets a gateway manifest mounted at /app/tool-manifest.json. The legacy briefing path is dead. Closes Phase 4 Task 8.

* test(property): implement tracing-completeness assertion across smoke-test batch

Replaces Phase 0 stub. Asserts the 6 tracing-contract requirements on every
completed task: audit_log agent_id non-null per state-transition row,
DEVELOPER:TASK_REFLECTION journal entry, QA:LEARNING journal entry,
CELL_PM/MAIN_PM:DECISION_LOG journal entry, acceptance_criteria_status
covering every criterion with a referencing_artifact_id, and
qa_evidence_inspected = true.

Uses an in-memory ephemeral Postgres test DB (`roboco_test_<pid>_<rand>`)
provisioned per pytest session, not SQLite — the production schema relies
on Postgres-only types (UUID, ARRAY) the SQLite dialect cannot compile.
Tests requesting db_session/smoke_test_batch are auto-skipped when no
Postgres is reachable on localhost:5432; ROBOCO_TEST_DB_HOST/PORT/USER
override the endpoint.

Schema is built via Base.metadata.create_all + manual ALTER for the
acceptance_criteria_status / qa_evidence_inspected columns, NOT via
`alembic upgrade head`. This sidesteps two pre-existing layer-drift items
that block any fresh migration run today:

  1. Migration 001 declares the agentrole Postgres enum with lowercase
     values (qa, developer, ...) but the SQLAlchemy ORM binds
     Enum(AgentRole) to the StrEnum's uppercase NAMES — production DBs
     mask this by being bootstrapped via create_all and stamped at 001.
  2. Migration 008 runs UPDATE agents SET skills WHERE id over an
     agents.skills column that no migration in this chain ever creates.

Documented in conftest.py so a future migrations cleanup can find them.
Also notes that acceptance_criteria_status/qa_evidence_inspected are in
the DB schema (per migration 006) but are NOT mapped on the ORM TaskTable
nor on the Pydantic Task model — services that read them via
`task.qa_evidence_inspected` rely on those values being set on raw rows.
The property test uses raw SQL to read the columns directly, matching the
DB-level contract.

Closes Phase 4 Task 11.

Side change: pyproject.toml — adds asyncpg.* to the existing
[[tool.mypy.overrides]] ignore_missing_imports list (asyncpg ships no
py.typed marker), matching the convention used for redis, anthropic,
piragi, etc.

Test count: 1; backend: Postgres (localhost test DB).

* style(mcp/flow_server): single-line _post call after format pass

* fix(db): map 7 gateway columns from migration 006 to TaskTable + Task model

active_claimant_id, last_heartbeat_at, pre_block_state, pre_block_assignee, pre_block_metadata, acceptance_criteria_status, qa_evidence_inspected: present in DB since migration 006 but absent from the ORM mapping. Gateway code (tracing_gate, choreographer, claimant_lock) reads these via task.<attr>; without the mapping, runtime would AttributeError. Closes PHASE4-BUG-A.

* fix(db): repair alembic chain — neutralize 008, add 009 enum reconcile, ORM uses values_callable

Three coordinated changes that close PHASE4-BUG-B:

1. roboco/db/tables.py — introduce _str_enum() helper that wraps Enum() with values_callable=lambda obj: [m.value for m in obj]. Apply to all 23 StrEnum-typed mapped columns. ORM now serializes by .value (lowercase) to match alembic 001's declared enum values; default Enum() was using .name (uppercase) which never matched.

2. alembic/versions/008_align_skills.py — replace with documented no-op. The original migration referenced agents.skills, a column that has never existed in any migration (the agents table has capabilities, not skills). The substitution intent (qa_review -> code_review) was already satisfied statically in roboco/agents_config.py.

3. alembic/versions/009_enum_reconcile.py — new migration that:
   - Adds missing enum values: agentrole.system, team.fullstack, taskstatus.quarantined.
   - Detects uppercase drift from a Base.metadata.create_all bootstrap and rebuilds agentrole/team/taskstatus enums with lowercase members + USING lower(col::text)::enum on every column referenced. No-op if already lowercase.

Tests stay green: 281 passed.

* feat(services): backfill 36 gateway-shaped methods for Choreographer

The gateway Choreographer was wired to call methods that the underlying
services did not expose. This adds them as thin wrappers + queries (most
alias canonical methods; a handful are gateway-specific variants).

TaskService — 26 methods: aliases (submit_verification, submit_qa,
list_blocked_for_team, list_blocked_all_teams,
list_awaiting_pm_review_for_team, list_assigned_for_agent), agent
queries (agent_for, qa_agent_for_team, documenter_for_team,
cell_pm_for_team, get_active_task_for_agent, list_paused_for_agent),
triage queries (list_awaiting_main_pm_all, all_subtasks_terminal),
state setters (set_plan, mark_evidence_inspected, mark_agent_idle),
QA/Doc claim variants (qa_claim, doc_claim, qa_pass, qa_fail),
PM completion (cell_pm_complete with merge_commit), unblock with
state restore (unblock_with_restore), and escalation
(escalate, escalate_up_to_role). Also adds GatewayAgentView
dataclass that unifies DB and config-derived agent attributes.

JournalService — 4 methods: existence checks (has_decision_for_task,
has_learning_for_task, has_reflect_for_task) + write_struggle.

GitService — 4 methods: branch-keyed entry points (create_pr,
pr_merge, pr_target, diff) plus push_branch helper. Each derives
project + workspace from the task that owns the branch / PR.

WorkSessionService — 2 methods: files_changed + has_unpushed_commits.
PR existence is the proxy for pushed (no per-commit push column).

Choreographer: switched git.push(branch_name) call to push_branch()
to dispatch to the new gateway-shaped helper.

* test(services): unit tests for 36 gateway-backfill methods

Adds happy-path + edge tests for every method added in the prior
backfill commit. Total 61 new tests across:
  - tests/unit/services/test_task.py (36)
  - tests/unit/services/test_journal.py (8)
  - tests/unit/services/test_git.py (10)
  - tests/unit/services/test_work_session.py (7)

Each test mocks at the session boundary (no DB) and stubs adjacent
service methods via a dynamic _bind helper to avoid mypy
[method-assign] noise without resorting to type:ignore comments.

* test(gateway): switch dev catch-up assertion to push_branch

The Choreographer's catch-up sequence was renamed from git.push(branch)
to git.push_branch(branch) when GitService got a gateway-shaped helper
in the prior commit. This updates the existing assertion to match.

* feat(mcp): add roboco-git-readonly server with status/log/diff/branches

Slim FastMCP server exposing the four read-only git tools every role
needs (status, log, diff, branch_list) by forwarding to /api/v1/git/*
on the orchestrator. Replaces the read-only half of the legacy
roboco-git server; write operations now go through gateway verbs in
roboco-flow / roboco-do.

The endpoint shapes mirror the panel-facing API (project_slug,
include_remote, staged/file_path) so the same backend handlers serve
both human and agent traffic.

* refactor(mcp): delete legacy task/journal/notify/a2a/message/project servers

Phase 4 cutover: agents now reach every state-changing surface through
the gateway (roboco-flow intent verbs + roboco-do content tools), with
roboco-git-readonly + roboco-optimal + roboco-docs covering reads. The
seven legacy MCP servers + their handler trees are dead code from the
agent side, so they're removed:

  roboco/mcp/task_server.py            (1020 LOC)
  roboco/mcp/journal_server.py         (512 LOC)
  roboco/mcp/notify_server.py          (440 LOC)
  roboco/mcp/a2a_server.py             (790 LOC)
  roboco/mcp/message_server.py         (682 LOC)
  roboco/mcp/project_server.py         (667 LOC)
  roboco/mcp/tasks/  (handlers+utils)  (~4300 LOC)
  roboco/mcp/test/   (in-container runner; replaced by gateway evidence
                      + manual smoke)
  roboco/mcp/git/    (full server; read-only half migrates to the new
                      slim roboco-git-readonly module, write half is
                      owned by gateway verbs)

Orchestrator updates:
  - _generate_mcp_config registers only roboco-flow, roboco-do,
    roboco-git-readonly, roboco-optimal, and (for docs roles) roboco-docs.
    No more per-role legacy fan-out.
  - base_allow flips to mcp__roboco-flow__*, mcp__roboco-do__*,
    mcp__roboco-optimal__*, mcp__roboco-git-readonly__*. Role-specific
    allow lists are reduced to file IO scoping, since gateway verbs
    enforce role policy server-side.
  - TRACEABILITY_TRIGGER_TOOLS rewritten in terms of the gateway servers
    (mcp__roboco-flow__* / mcp__roboco-do__*) instead of the now-deleted
    per-tool list.

Test fix: tests/unit/services/test_a2a.py imported _handle_send_chat_message
from the deleted a2a_server. The four MCP-layer URL-builder tests (empty
conversation_id guard) are dropped — the equivalent boundary now lives
in /api/v2/do/* which has its own integration coverage. The two
service-layer nil-UUID guard tests are kept; they exercise A2AService
directly and remain meaningful (the panel still uses the v1 chat surface,
where a buggy caller could pass the nil UUID).

Net: ~9600 LOC removed from roboco/mcp/. quality-fast green:
338 tests pass, mypy clean, ruff clean. No /api/v1/* router changes —
those endpoints stay live for the panel UI which still uses every
lifecycle action; agents have no prompts that name them so the path is
dead code from the agent side.

* docs(claude.md): replace legacy MCP listing with gateway/verb-surface section

Phase 4 cutover: agents go through roboco-flow + roboco-do (gateway), not the deleted task/journal/notify/a2a/message/project servers. Document the verb surface per role + the Envelope response shape so future Claude Code sessions land in the correct mental model. Closes Phase 4 Task 13.

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-05-02 03:11:49 +02:00