Files
roboco/docs/rag/tools/task-tools.md
T
cea3e56628 feat(lifecycle): revision findings ledger — structured failure feedback, persisted and delivered down the chain (#486)
* feat(lifecycle): revision findings ledger — structured QA/PR/PM/CEO failure feedback, persisted and delivered down the chain

Every bounce used to survive only as flattened prose: rounds overwrote each
other in notes_structured, request_changes persisted nothing, two raw
dev_notes appends were silently destroyed by the next handoff note, and the
dev prompt pointed at fields (qa_notes via evidence(), pm_notes) the API
never delivered. Agents re-interpreted and re-discovered every failure
before they could start fixing it.

- task_review_findings (migration 071, append-only): file/line/severity/
  criterion(AC-id-validated)/expected/actual/fix/evidence per finding, with
  origin (qa|pr_gate|pm|ceo), round, and an open->addressed->verified
  lifecycle (waived reserved); new tasks.pm_notes + PmReviewContent give
  request_changes a structured home
- producers: fail_review/pr_fail/request_changes take findings=[...] (prose
  issues shimmed+merged for one release, deprecation-logged); ceo_reject
  validates its reason (no 500), lands an origin=ceo finding, and bumps
  round+audit on branchless coordination roots; guardrails at the verb
  chokepoint (nudge >5, hard reject >10, field caps, traversal-safe file);
  the dev_notes data-loss appends are removed; new task.request_changes +
  task.ceo_reject audit events close rework attribution
- delivery: qa_notes/pr_reviewer_notes/pm_notes carry the deterministic
  [F-id8] rendering; claim briefings, evidence(), the REVISION_REQUIRED
  spawn prompt, PM triage bounced-blocks, and A2A bodies deliver open
  findings; round-N+1 QA and gate reviewers get the full prior ledger;
  panel Findings tab + bounced-xN chip; metrics pm_rejects/ceo_rejects +
  findings counts; vault task notes render a Findings section (fail-open)
- resolution closes for every origin: i_am_done and submit_up/submit_root
  take resolved_findings gated by FINDINGS_ADDRESSED (owner-gated so a
  stale non-owner PM can never mutate the ledger); pass_review/pr_pass/
  complete verify-stamp same-transaction; ceo_approve stamps best-effort
- 24 real-DB integration tests drive the full loop through the real
  choreographer; full suite 12856 green

* docs: revision findings ledger sweep — CLAUDE.md, map, RAG corpus

- CLAUDE.md: new ledger section + corrected request_changes row
- docs/map/review-findings.md (new subsystem map) + surgical updates to
  task-service/pr-gate-review/metrics-observability/vault/panel maps
- docs/rag: producers' findings contract across qa/pr-reviewer/developer/
  cell-pm/main-pm/ceo role docs (the PM docs were missing request_changes
  entirely), verb references, and a new architecture/review-findings.md
  disambiguating ledger findings from convention findings

* test(e2e): resubmit resolves the pr_fail finding per the ledger contract

The scripted pr_fail revision loop resubmitted submit_up without
resolved_findings — correctly rejected now that FINDINGS_ADDRESSED gates
the PM resubmit verbs (green locally, red only in CI since the e2e suite
skips without ROBOCO_E2E_SMOKE=1). The scripted PM now reads the open
ledger row pr_fail persisted (new open_finding_ids arc helper) and
resolves it on resubmit, asserting the open set drains — exercising the
coordinator half of the new contract end to end.

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-07-11 22:54:42 +02:00

11 KiB

Task Management Tools

There is no roboco_task_* tool surface. Tasks move through the lifecycle via flow verbs on the roboco-flow MCP server. Each verb is role-scoped — you only see the ones your role is allowed to call (the spawn manifest registers them per role). Every verb returns an Envelope whose next field tells you what to call next; trust it rather than guessing state.

The verbs below are grouped by who calls them.

Developer flow

give_me_work()                  # returns your most-actionable pending task
i_will_work_on(task_id, plan="...")
                                # claims + sets plan + starts; auto-creates and
                                # checks out feature/{team}/{task-hierarchy}
commit(message, files=None)     # content tool — repeat per change (auto-pushed)
open_pr(task_id)                # pushes branch + opens the PR
i_am_done(task_id, notes="", resolved_findings=None)
                                # verifying -> awaiting_qa (PR must already be open);
                                # on a bounced task, name every open ledger finding
                                # via resolved_findings=[{finding_id, commit?, note?}]
i_am_blocked(task_id, reason)   # external dependency; cell PM unblocks
unclaim(task_id)                # release a claimed task back to the queue
resume(task_id)                 # recover a paused task after compact/restart
i_am_idle()                     # no work in your queue right now

There is no separate claim / start / pause verb — i_will_work_on composes claim + set-plan + start atomically, and i_am_done composes verify + submit-qa. Branches are auto-created on i_will_work_on; do not checkout by hand.

QA flow

give_me_work()                  # returns an awaiting_qa task
claim_review(task_id)           # claim for review (auto-checks-out dev branch)
pass_review(task_id, notes)     # awaiting_qa -> awaiting_documentation
fail_review(task_id, findings=[{file?, line?, severity, criterion?, expected, actual, fix?, evidence?}])
                                # awaiting_qa -> needs_revision (dev gets it back);
                                # the deprecated issues=[str] shim still works this release
unclaim(task_id) / resume(task_id) / i_am_idle()

notes (on pass_review) and each findings entry (on fail_review) must be substantive — the enforcement layer rejects empty or near-empty content. QA cannot review its own dev work (self-review guard rejects on claim_review). Every fail_review finding is persisted to the append-only revision-findings ledger and rendered into qa_notes; a soft nudge fires above 5 findings, a hard reject above 10. On a round ≥2 review, claim_review also returns prior_findings (the full ledger) so you check what was filed before. See docs/rag/architecture/review-findings.md.

Documenter flow

give_me_work()                  # returns an awaiting_documentation task
claim_doc_task(task_id)         # claim the doc phase
commit(message, files)          # commit the doc files you write
i_documented(task_id, notes, files)
                                # awaiting_documentation -> awaiting_pm_review

Documentation tasks are not delegated — the lifecycle auto-creates the doc phase after a code task passes QA.

Cell PM flow

triage()                        # list actionable tasks in your cell
i_will_plan(task_id, plan, approach)
                                # claim + plan + start a parent task
delegate(parent_task_id, title, description, assigned_to, team, task_type,
         nature, estimated_complexity, acceptance_criteria,
         covers_parent_criteria=[...])
                                # create a subtask; covers_parent_criteria maps
                                # it to the parent ACs it is responsible for
reassign(task_id, assigned_to)  # move a subtask to a different agent
unblock(task_id, reason)        # blocked -> in_progress (PM only); reason is
                                # recorded as your journal:decision (no separate
                                # note needed)
submit_up(task_id, notes, resolved_findings=None)
                                # open cell->root PR; -> awaiting_pr_review
                                # (the cell PR reviewer gates it; after pr_pass
                                #  the same Cell PM completes + merges); a re-submit
                                # after pr_fail must resolve every open finding first
complete(task_id, notes)        # awaiting_pm_review -> completed (merges leaf PR)
request_changes(task_id, findings=[...])
                                # reject a subtask's merge review -> needs_revision,
                                # routed to whoever owns the revision; structured
                                # findings persist to the ledger + render into pm_notes
escalate_up(task_id, reason)    # escalate to your escalation target

After i_will_plan and each delegate, the envelope includes a coverage view of the parent — parent_ac_coverage (per-criterion id / text / claimed / verified) and unclaimed_parent_acs (criteria no subtask covers yet). A parent cannot idle with unclaimed criteria, nor complete / submit_up / escalate_to_ceo until every criterion traces to a child that passed QA. These gates stay inert until you start declaring covers_parent_criteria. See docs/rag/workflows/task-planning.md.

Delegation rules (enforced): main_pm -> cell_pm; cell_pm -> its team's devs. Cell PMs receive planning-typed parent tasks; devs get code/research (UX devs also design). Always create subtasks via delegate with parent_task_id set — there is no standalone task-create verb for agents.

Main PM flow

The Main PM shares most Cell PM verbs (i_will_plan, delegate, complete, request_changes, unblock, triage, escalate_up), adds the verbs below, and — unlike a Cell PM — has no submit_up or reassign. Its bubble-up verb is submit_root (the root analogue of the Cell PM's submit_up):

triage_all()                    # list actionable tasks across all teams
submit_root(task_id, notes, resolved_findings=None)
                                # open root->master PR; -> awaiting_pr_review
                                # (the main PR reviewer gates it; after pr_pass,
                                #  complete escalates to the CEO); a re-submit
                                # after pr_fail must resolve every open finding first
escalate_to_ceo(task_id, reason)
                                # awaiting_pm_review -> awaiting_ceo_approval
give_me_work()                  # Main PM may also pull work directly

For a code root the Main PM must submit_root first — that opens the root→master PR and enters the in-path gate (awaiting_pr_review); only after the main reviewer pr_passes it does complete escalate to the CEO. A branchless coordination root (product fan-out, no repo) skips the gate and is completed/escalated directly. The Main PM never merges to mastercomplete escalates and only the CEO merges the root→master PR.

Board flow (Product Owner / Head of Marketing)

triage()                        # list actionable tasks in scope
escalate_to_ceo(task_id, reason)
i_am_idle()

The Board cannot claim, create, complete, or cancel tasks. Strategic decisions are escalated to the CEO.

The Product Owner additionally has propose_roadmap(cycle_goal, items) — a content tool on roboco-do, not a flow verb, so it doesn't appear above. It authors the weekly board-roadmap-engine exploration cycle (a themed goal + 3-7 item drafts); the CEO approves or rejects each item individually into BACKLOG. See docs/rag/roles/product-owner.md.

Auditor flow

triage()                        # read-only list of actionable tasks
i_am_idle()

The Auditor is a silent observer: read-only triage, no dm/notify, no claim/complete/cancel.

PR Reviewer flow

give_me_work()                  # returns an inbound-PR review task
claim_pr_review(task_id)        # claim it (planless, branchless — read-only)
post_pr_review(task_id, ...)    # posts one change-request on the PR; task -> completed
unclaim(task_id)                # release a claimed inbound or gate review back to the pool
i_am_idle()

The PR Reviewer reviews inbound external/fork (and, behind a flag, internal) PRs the org did not open. It is read-only: no commit/open_pr/merge, no dm — the change-request is posted server-side on the PR itself, and the CEO decides Supersede/Dismiss from the PR Review Queue.

The same role also runs the in-path PR-review gate on the org's own assembled delivery PRs — the merge-level review before the PM merges:

claim_gate_review(task_id)      # claim an awaiting_pr_review task; returns the assembled
                                # diff + (on round >=2) prior_findings, the full ledger
pr_pass(task_id, notes)         # assembled PR is correct -> awaiting_pm_review (the PM merges)
pr_fail(task_id, findings=[...])
                                # send it back -> needs_revision, like a QA fail;
                                # the deprecated issues=[str] shim still works this release

Both verdicts are also posted on the assembled PR itself as a GitHub review (server-side, bot account) so the decision is visible on the PR the PM merges: pr_pass → APPROVE, pr_fail → REQUEST_CHANGES — except the root→master PR, which only ever gets a plain COMMENT (only the CEO acts on master).

A cell reviewer (be/fe/ux-pr-reviewer) reviews its cell's assembled cell→root PR; pr-reviewer-1 reviews the root→master PR for the cross-cell integration seam, before the CEO sees it.

Cancel

Cancelling a task (any non-terminal status -> cancelled) is restricted to PM roles and the CEO — except awaiting_ceo_approval -> cancelled, which is CEO-only (a PM cancelling a task already in the CEO's queue would bypass the human approval gate). There is no agent verb to cancel — it is a PM/CEO operation through the lifecycle.

Progress

Record progress against your plan with the progress content tool (on roboco-do), not a task verb:

progress(task_id, message="API skeleton landed", plan_step="2")

Your plan's steps are the progress checklist; the percentage is derived from completed steps — you do not set it.

Sandbox DB/Redis/Mongo (Developer + QA)

request_sandbox(services=None) — a content tool on roboco-do, not a flow verb — provisions a throwaway sandbox Postgres/Redis/Mongo on demand, for a project that opted in (projects.sandbox_services). Only developer and qa carry it. Omit services for the project's whole opted-in set; requesting one outside it is rejected naming the allowed set. Creds come back in the envelope's evidence, one entry per service, including ready-to-export ROBOCO_TEST_* values for gate tooling. Calling it again is a cheap no-op (same creds). See docs/rag/architecture/sandbox-db.md.