* docs: sync prompts/RAG/CLAUDE + bump to 0.11.0 for the run-hardening wave
Documentation + version sweep for everything shipped since 889f3689 (the 0.11.0
wave: MegaTask + #249-#253 run-hardening). Closes the doc drift behind the live
incidents — agents had no branch-behind-master guidance, so a Main PM invented
a bogus "rebase subtask".
Agent guidance (the headline gap):
- main_pm / cell_pm / developer prompts: a task branch is made current at CLAIM;
there is NO rebase/pull/merge verb at the agent layer. Never create a "rebase
subtask" or improvise git surgery; escalate a behind-base branch
(developer: i_am_blocked; PM: escalate_up). "A rebase subtask is always a mistake."
- board prompt: Board has no unblock verb; a blocked task assigned to it is a
mis-assignment -> escalate_to_ceo immediately, never sit on it (respawn loop).
- developer prompt: the shared clone is git-reset on a fresh claim; push/open_pr
target the task branch by name regardless of the current checkout.
- RAG (git-errors, blocked-tools, pr-creation): branch-behind-base, "src refspec
does not match any", and non-fast-forward recovery -> escalate, don't improvise.
CLAUDE.md: 9 shipped behaviors synced (session-limit parking, one-active-work-
session + migration 047, push/PR-by-name + origin ref recovery, fresh-claim
workspace reset, Board never owns a coordination root, verb-runner per-action
INVALID_STATE re-check, note fire-and-forget RAG indexing,
ROBOCO_GATEWAY_HEALTH_ENABLED flag, learnings not broadcast to human roles).
Version 0.10.0 -> 0.11.0: pyproject, roboco/__init__, config.app_version +
agent-image-tag example, panel/package.json, uv.lock, README/deploy examples;
CHANGELOG [Unreleased] cut to [0.11.0] - 2026-06-24.
* docs(site): document session-limit parking + the branch-behind-base operator flow
User-facing docs site updates for the 0.11.0 wave (the run-hardening behaviors
that are operator-visible):
- models/resilience.md: the Claude session-limit (5-hour usage window) parks
and auto-revives like an overload, not just per-request 429s / 5xx overloads.
- troubleshooting/common-issues.md: same session-limit note on the parked-
provider entries; plus a new "task stuck on a branch behind its base" entry —
agents have no rebase verb so they escalate it; the operator rebases from the
panel Git tab (auto-rebase-at-spawn is the roadmap cure).
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
51 KiB
Changelog
All notable changes to RoboCo are documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
[Unreleased]
[0.11.0] - 2026-06-24
Added
- MegaTask — describe several tasks in one intake chat and ship them as one sequenced batch. When the CEO wants several pieces of work at once — even across projects that don't share a codebase (e.g. a SaaS app, its open-source core engine, and a framework adapter) — the intake modal now offers a third scope, MegaTask, beside Single cell and Board-led. You pick the repos it spans; the intake agent reads them all and proposes the whole batch in one hand-off (the new
propose_batchtool), one draft per task, each carrying its own project plus a collision surface (which files it touches, whether it adds a migration, whether it edits a widely-shared component). A deterministic analyzer (SequencingService) turns those surfaces into conflict-free waves — file-overlap and migration-adding tasks are serialized, a shared-surface edit runs after what it overlaps, independent tasks run in parallel — and the Board reviews the batch once. On confirm RoboCo creates a branchless umbrella task (the Main PM's coordination + board-review + CEO-approve unit) over N root-subtasks, each a real coordination root with its own project, branch, and PR, wired with the analyzer's dependencies so the existing dependency-gate dispatches the waves in order. The umbrella assembles no PR of its own, is exempt from the branch gate, and completes only when every root-subtask is terminal (then it escalates to the CEO). On the Board route the root-subtasks are held until the umbrella is approved, then released. Surfaced as a core capability — no feature flag — branded "MegaTask" across the panel, prompts, and docs; internal names stay technical (batch_id,SequencingService). Addstasks.batch_id+ the three collision-surface columns (migration 046),confirm_live_batch+POST /prompter/live/{session}/confirm-batch, multi-project intake spawn (project_ids), thepropose_batchtool on both intake runtimes (Claude SDK driver + grok CLI server), and the panel's MegaTask scope + Review-MegaTask card.
Fixed
-
A Main PM blocking its own coordination root no longer hands the whole root to the Board (a respawn catch-22). Root cause: the generic escalation chain points
main-pm → product-owner, andi_am_blocked/ escalate REASSIGNS the task to that chain target. The board-advisory guard that refuses such a hand-off only covered descendant cell tasks (it requiredparent_task_id), so a top-level Main-PM coordination root slipped through and the entire root was reassigned to the Product Owner and marked blocked. A Board role has nounblockverb at all (only notify / note / triage / i_am_idle) and the unblock gate is assignee-only, so it could neither resolve the blocker nor hand it off — it just spam-notified the CEO while the blocked-task dispatcher respawned it every tick (one live incident burned an estimated 6400+ tool calls on a single root). Fixed at both layers: the escalation / reassign / revival guard now also refuses a Board owner for amain_pmcoordination task (root or MegaTask root-subtask) and diverts it to the pool for a role-matched (Main-PM) re-claim — the upstream cure — via a single shared_board_cannot_ownpredicate; and, as a defense-in-depth backstop, the orchestrator's blocker dispatcher no longer treats a Board role as a blocker resolver (it returns no resolver, so a mis-owned blocked task is skipped rather than respawned onto a role that physically cannot act). -
A racing state change mid-verb no longer crashes a PM into a respawn loop. The gateway's verb runner guards the initial task/agent against
None, but its composed atomic actions reassign the working task from each step (i_will_planruns claim → set_plan → start). When a concurrent agent transitioned the row between the verb's precondition gate and execution — e.g. a racingi_am_blockedmoved a coordination root fromneeds_revisiontoblocked—claim()found no valid transition and returnedNone, then the next step dereferencedNone.idand crashed with the opaque'NoneType' object has no attribute 'id', surfaced to the agent as a cryptic "verb runner failed" so the PM respawn-looped on the wedged root. The runner now re-checks after each composed action and fails fast with an actionableINVALID_STATEthat tells the agent the row changed under it and to re-fetch and re-issue its verb (the savepoint rolls the partial sequence back). -
A completed task no longer wedges when its branch is missing from a re-provisioned clone. Push-by-name (the fix that decoupled the push from the workspace checkout) still requires the named task branch to exist as a local ref — but a developer's shared clone can be freshly re-provisioned (the per-task workspace-collision recovery re-clones it), leaving the task branch absent locally even though its commits are safely on
originand the clone is parked on a different task's branch.git push origin <branch>then died with the crypticsrc refspec <branch> does not match anyand the task blocked-looped ati_am_done. The push now recovers a missing local ref fromorigin/<branch>first (a clean no-op when the work is already on origin); if the branch exists on neither the clone nor origin the commits are genuinely gone from this clone, so it fails loud with a recoverable "unclaim the task and re-claim it to rebuild the branch, then replay your commits" instruction instead of the raw refspec error. -
The orchestrator's own recovery actions now actually run. Its background dispatcher made internal HTTP calls to its own API without an agent identity, so every self-
PATCHto a task — auto-blocking a task with missing prerequisites, auto-resuming a PM's paused parent, auto-recovering a stale-blocked parent, annotating an SLA breach — was rejected with401 Missing X-Agent-IDand silently dropped. The visible effect was paused/blocked parent tasks staying wedged and their dependent work stranded (with the dispatcher logging a "respawning assignee" loop). Header propagation was inconsistent across the orchestrator's separate HTTP-client call-sites — only the main dispatch loop sent the identity. The system identity is now hoisted into one shared constant and applied to every API-facing dispatcher client (the external provider-recovery probe is intentionally excluded); thesystemrole holds the permission required for the audited status-override path those routes use. -
A developer's completed work no longer silently fails to reach GitHub ("No commits between"). A developer's single git clone is shared across all of their tasks, so by the time a task's PR is opened the clone has usually moved on to a later task's branch. The push at the QA-submission /
open_prboundary, and the PR's head branch, were both taken from the clone's current checkout — so the push was rejected (the workspace was parked on another task's branch) and the locally-committed work never reachedorigin, leaving the task branch empty andopen_prfailing with GitHub's "No commits between" 422. The work was on disk and correct, just never pushed. Both the push and the PR head now operate on the task's recorded branch by name, independent of the checkout (push(branch=…)targets the named ref; the PR head is the task'sbranch_name). Work committed on any of a shared clone's task branches now pushes and opens its PR correctly. -
Hitting the Claude session limit now parks the workforce instead of crash-looping it. When the org's Claude usage ("5-hour") limit is reached, each agent container exits with a 429 rejection; the orchestrator was treating that like any crash and immediately respawning the agent straight back into the limit, over and over, across the whole fleet. It already parks the provider on a persistent server overload (529/500/503) and revives the parked work once it recovers — but that detection only matched the overload signatures, not the session-limit 429. The same park-and-resume break now also recognizes the session limit: the provider is parked, dispatch goes quiet, and the background probe loop brings the agents back automatically when the window resets — no churn, no wasted respawns.
-
A failed PR review no longer looks green. On a task's detail page, the "PR Reviewer Notes" card was painted a fixed teal/green background regardless of the review verdict, so a
Failedreview — red badge and all — sat inside a green card and could read as passing at a glance. The card background now mirrors the verdict the way the QA Notes card already does: red on a failed review, green on approved/passed, amber on changes-requested, and neutral before a verdict is in. -
The CEO and other human roles no longer get spammed with agent "learnings." Whenever an agent recorded a learning, RoboCo broadcast it as a knowledge-share notification — and the recipient query swept in the human roles too (the CEO, plus the human-driven prompter and secretary). Agent knowledge-sharing is a signal for agents; in a human's inbox it is just noise. Those roles are now excluded from learning broadcasts.
-
A gateway verb on a vanished task/agent fails cleanly instead of crashing cryptically. The verb runner's atomic steps dereference
task.id/agent.idwith no guard, so a verb invoked when the task or agent could not be resolved (e.g. a task forced into an unexpected state out-of-band) crashed with an opaque'NoneType' object has no attribute 'id'. The runner now fails fast with an actionableINVALID_STATEerror that tells the agent to re-fetch and re-issue its claim verb. -
An agent could be permanently wedged in a respawn loop by duplicate work sessions on one task. A task is owned by one agent at a time, so it must have at most one active git work session — but nothing enforced that: when a task was re-claimed by a different agent (after a pool release, reaper unclaim, or escalation redirect) the prior holder's active session was left open.
WorkSessionService.get_active_for_taskthen ran a one-row query across the duplicates and raisedMultipleResultsFound; the caught failure surfaced as the cryptic'NoneType' object has no attribute 'id'that crashed the claim/plan/start flow, so the task could never advance — the orchestrator re-spawned its PM every ~30s forever and the task's dependents stayed blocked. (This was the real root cause behind the verb-runnerINVALID_STATEguard above, which only made the crash legible.) Fixed at three layers: the active-session lookups now return the most-recent session instead of raising; claiming a task supersedes any other agent's stale active session (the single-active-per-task invariant); and a partial unique index — migration 047, which first de-duplicates existing rows, keeping the most recent — enforces it at the database level so it can never recur. -
A dev claiming a new task no longer gets stuck on
BRANCH_MISMATCH. Each developer has one persistent clone shared across all their tasks, so a finished or abandoned prior task could leave the clone dirty and sitting on a sibling task's branch. The claim's git work (creating/checking out the new task's branch) runs as a side-effect after the claim's DB transition commits — so when the checkout failed on that dirty tree, the task was already marked assigned while the workspace stayed on the wrong branch, and the dev's next commit was rejected withBRANCH_MISMATCH(stalling, then blocking, the task). The claim now does agit reset --hardto clean the tree before the checkouts. It runs only on a fresh claim (resume short-circuits earlier), so the discarded changes are abandoned cruft from a finished task — never committed work, and never the gitignored.venv. -
The
notetool no longer times out under load. Writing a journal entry / note synchronously waited on RAG indexing, which embeds via Ollama — and Ollama is CPU-bound, so under concurrent load that embed slowed enough to time thenotegateway tool out entirely (despite a "non-blocking" comment on the code). The entry is already persisted before indexing, so indexing is pure best-effort enrichment: it now runs fire-and-forget on the event loop, and the note/journal write returns immediately. -
A feature flag stopped showing its raw internal key. In Settings → Feature Flags, the "Gateway-health recovery" toggle displayed its raw key
gateway_health_enabledas its description (the only flag missing a human blurb). Added the description, and changed the fallback so a future flag without one renders nothing rather than leaking a snake_case key.
[0.10.0] - 2026-06-23
Added
- Delivery observability dashboards — cycle-time, bottlenecks, rework rate, and per-agent/per-cell scorecards. A new "Delivery" tab on the Metrics page surfaces how work flows, built on data RoboCo already captures: per-stage cycle time reconstructed from the
audit_logtransition journey, a bottleneck view (which lifecycle stage holds the most cumulative time + how many tasks are parked there now), a rework view (how often work bounces toneeds_revision, by team and by agent, with the rejection attributed to the QA / PR-reviewer who made it, plus the rework's token cost), and fused per-agent / per-cell scorecards. Backed by new read-onlyMetricsServicemethods and/dashboard/metrics/{cycle-time,bottlenecks,rework,scorecard}endpoints. To make rework correct and O(1), each task now carries arevision_countincremented at the single transition chokepoint (migration 045, with a compositeaudit_log(target_id, event_type, timestamp)index for the reconstruction queries), and QA/PR-review bounces emit rejector-attributedtask.qa_fail/task.pr_failaudit events. No feature flag — it reads the always-on metrics surface. - Gateway-health recovery — a broken-but-alive agent is recovered instead of protected forever. The verb-driven heartbeat cannot distinguish a healthy agent quiet during a long edit/test cycle from one whose MCP gateway is broken (e.g. a corrupted
/app/.venvso every gateway tool import raises) while its container stays up — and the reaper's live-skip would shield that broken agent indefinitely. The reaper now probes the gateway out-of-band (docker exec: does the gateway venv import its deps?) and, once it has been broken longer thanROBOCO_GATEWAY_HEALTH_GRACE_SECONDS(so a transient probe miss is tolerated), kills + evicts the container so it falls through to release + respawn; a healthy or inconclusive probe spares it. Gated byROBOCO_GATEWAY_HEALTH_ENABLED(default-on reliability fix, in the panel Feature Flags). Builds on the shipped bash-guard/appblock and reaper Docker-liveness fallback — together the third leg the live incident exposed. - Edit a task's sequence from the task details page. A task's
sequence(its order within siblings — lower runs first) was display-only with no way to change it from the UI. The details page's Dependencies tab now carries an inline sequence editor alongside the parent / dependency editors, andPATCH /tasks/{id}accepts asequencefield (owner or privileged role), so an operator can re-order sibling work directly.
Fixed
- Metrics "hours" fields serialized as JSON strings, crashing the panel.
EXTRACT(epoch …)returnsnumericon PostgreSQL 14+, which asyncpg surfaces as aDecimal— and aDecimalserializes to a quoted JSON string. Every SQL-averaged hours field —avg_cycle_hourson the new Delivery scorecards, plus the pre-existingavg_completion_hours/avg_blocked_hours/longest_blocked_hours— was therefore a string, so the panel'svalue.toFixed(…)threwtoFixed is not a functionand blanked the tab. A single_as_hourscoercion now rounds each to a realfloat, so every hours field is a JSON number. (Token/cost fields were alreadyfloat()-cast and unaffected.) - The Main PM could not advance past its first coordination task — the developer single-task concurrency guards were deadlocking the coordinator. A PM plans and delegates many root tasks in parallel; the real work then runs in the delegated cells, not in the PM's own hands. But the claim-time guards that correctly keep a developer to one task at a time —
already_active(you have another claimed / in-progress task) andpaused(you have a paused task, resume it first — which fires afteri_am_idleauto-pauses the PM's own umbrella) — were applied to the PM as well, so once it held one root it could never plan a second: it thrashed between its claimed roots and respawned every few minutes, burning tokens for zero progress. These two guards are now skipped for the coordinator PM roles (main_pm/cell_pm): a PM may hold any number of roots in parallel, gated only by a genuine upstream sequence dependency (unmet_dependency), which still parks the task topendinguntil its dependency reaches a terminal state. As defense-in-depth thepausedguard now also excludes the target task itself, so a PM re-entering its own paused umbrella can never self-block. - Task notes were invisible in the panel — the API response dropped them. The
task_to_responseserializer (used by the task list and detail endpoints the panel reads) setdev_notes/qa_notes/quick_contextbut omittedpr_reviewer_notes,doc_notes, andnotes_structured, andTaskResponsedidn't even declarenotes_structured— so the PR-reviewer's notes, the documenter's notes, and the structured PR-review verdict were always blank in the UI no matter what the agents wrote to the DB (the structured-content write-path and obligation gates work; the data simply wasn't being serialized). The builder now returns all note sections plus the structured source of truth. (dev_notes/qa_noteson an in-flight task are still legitimately empty until the developer submits / QA reviews.)
[0.9.0] - 2026-06-23
Added
- Architectural Conventions Standard — a per-project, repo-canonical architecture map that gates where code may live. Beyond the
make-style checks (syntax, types, tests), each project can carry a.roboco/conventions.ymldeclaring which definition kinds belong in which modules, a toggleable rule set, custom regex rules, and waivers — so an agent can no longer land a Pydantic model inside a router or a lint suppression (a misplaced helper — any top-level function — warns rather than blocks). A tree-sitter validator CLI (Python + TypeScript) classifies every changed definition and emits findings; ablock-level finding refuses a developer'si_am_doneand the in-path PR gate'spr_passwith the offendingfile:lineand a fix hint, and findings surface in QA's review evidence. The auto-derived defaults exclude test and documentation trees, count an explicitdb.commit()in a route as legitimate (not a fat-route violation), and exempt a small allowlist of structurally-unavoidable framework suppressions (ruffTC001–TC003, pydanticprop-decorator). The committed file and repo scan are read from a dedicated project-level read clone the service ensures on demand — so the standard resolves even for a project created before it existed, with no manual workspace configuration. The file is auto-scaffolded on first clone, editable from a per-project Conventions tab in the panel, and a false positive is cleared by a waiver committed in the branch and reviewed in the PR. Gated byROBOCO_CONVENTIONS_ENABLED(default off) and fully inert when off. - Agent runtime toolchain matching — agents build each target project under the Python that project actually requires. The agent image bakes one interpreter, but the projects RoboCo builds don't all share it, so a self-gate could pass against the wrong runtime. The workspace now resolves each target's Python from its
requires-python/.python-version, provisions the clone withuv sync --extra dev --python <version>(fetching the interpreter on demand), and records a.git/.roboco-toolchainmarker. A guard refuses a developer'si_am_done, QA'spass_review, and the PR gate'spr_passwhen the suite cannot be collected under the provisioned interpreter, so "verifying by reading source" can't masquerade as a passing gate. Gated byROBOCO_TOOLCHAIN_MATCH_ENABLED(default off). - Provider overload circuit-break — a persistent model-API overload parks the provider instead of crash-retrying into it. A sustained 529/500/503 (the SDK already retries transient ones) now trips the same park-and-probe break as a rate limit: the spawn gate queues further work for that provider and a background loop revives it when the overload lifts, instead of respawning the agent straight back into the failure and burning tokens. Gated by
ROBOCO_OVERLOAD_BREAK_ENABLED(default on). - Structured content standard with obligated note sections. Every agent-authored handoff (developer, QA, documenter, PR-reviewer, auditor, PM resumption) is now a validated structured model persisted as the source of truth, with the legacy text column derived from it through a single chokepoint. An anti-soup guard rejects filler and all-token-noise free-text across the flow and content verbs, structured PR-review findings render a generated GitHub comment, and each role's note section is obligated at its lifecycle transition the way journals already were.
- User-facing documentation site. A MkDocs Material site (source under
docs/) is now built and deployed to GitHub Pages, publishing the organizational blueprint, role descriptions, task lifecycle, and how-to guides at the project's github.io site; the agent-facing RAG corpus underdocs/rag/stays excluded from the published site. Documenter output is also committed into the project repository (not only the RAG knowledge store) so it ships through the open PR.
Changed
- RoboCo adopts its own architectural standard. The repo now ships a canonical
.roboco/conventions.yml, and the inline request/response models that lived in thesystemand*_liveroute modules were relocated toroboco/api/schemas/so the codebase passes its own placement gate (no_models_in_routes/modular_cohesionare now clean and enforced atblock). - RoboCo's own
requires-pythonfloor is raised to>=3.13. The codebase importstomllib(3.11+) and runs on 3.13; the previous>=3.10floor made the toolchain resolver provision the self-hosted build at 3.10, where the suite cannot even be collected. Agent gate containers now also receive the test-database connection, so an agent'smake qualityruns the real, DB-backed suite instead of a coverage-collapsing unit-only subset.
Fixed
- Documentation now actually lands in the project repo. A documenter's output reached a host-mounted, RAG-indexed knowledge store and (more recently) was committed onto the task branch — but in the documenter's own workspace clone, and nothing ever pushed that commit, so the PM merged the already-open PR without the docs and the deliverable vanished on merge. The documenter's
i_documentednow pushes the task branch before handing off (mirroring the developer's pre-QA push), so the doc commit rides the open PR into the repository; a push failure holds the task inawaiting_documentationfor a retry instead of silently dropping the docs. - The conventions standard now resolves for projects created before it existed. It previously read the committed
.roboco/conventions.ymland the repo scan fromproject.workspace_path— a field only a manual API call ever set — so an older project (or one whose workspace was cleared) showed an empty "missing" map no matter what was pushed. The service now ensures a dedicated, default-branch read clone on demand and reads from it, persisting the resolved path + HEAD (the backfill). The panel tab, the spawn-time ambient block, and the per-task constraints all resolve the committed standard with no manual setup. - The conventions ambient prompt block no longer truncates mid-line. It now lists only modules that actually constrain a kind, and when the list would exceed its budget it trims at a line boundary with a
+N morepointer instead of cutting a module in half. - The conventions read clone now stays current on a private repo. Its refresh reused the orchestrator's token-less best-effort fetch, but the clone's remote URL is credential-stripped — so on a private repo the refresh fetch failed silently and the clone stayed frozen at clone-time, never seeing commits merged afterwards (the panel showed "auto-derived defaults" even after the standard was merged to the default branch). The refresh now performs a token-authenticated fetch + hard-reset, mirroring the clone.
- Self-heal fix tasks dispatch autonomously instead of being stranded. A self-heal task was opened
confirmed_by_human=falseand held out of dispatch until an "Approve & Start" — but that button only renders for board-reviewed Intake tasks, never for a self-heal task (team=main_pm, no board review), so there was no way to start it and it sat inpendingforever. Self-heal now opens the fix task confirmed + assigned to the Main PM, so the dispatcher picks it up immediately. The fix still ships through the normal gates (dev → QA → PR review → the CEO's merge); the loop never starts, merges, or deploys. - Self-heal no longer reads the wrong branch and fails silently. The CI-signal fetch filtered runs by
project.default_branch or "main"— the only"main"fallback in the codebase (everywhere else falls back to"master") — so a project whose default branch ismaster(like RoboCo) with an unsetdefault_branchmatched zero runs and the signal silently went dark: no fix task, no notification. The fallback now matches the rest of the codebase, and an armed self-heal that reads no CI signal (no/expired token, wrong branch, or a GitHub error) now logs a loud warning instead of an invisible no-op. - The toolchain gate no longer passes silently on an unverifiable workspace. A
brokeninterpreter still blocks; anunknownstatus — the smoke could not confirm the suite is collectable — now emits a warning when the gate proceeds, instead of slipping through unseen. - The crypto tests are hermetic. The Fernet round-trip tests supply their own key instead of depending on
ROBOCO_ENCRYPTION_KEYin the environment, so they pass in any gate container without the production secret being injected. ollama-initis best-effort and gates startup on the models being present, so a slow or unreachable model registry can no longer down a fully-cached deployment.- A PM can recover its own coordination task from
needs_revision, and lifecycle-transition notes are kept off the human-facingquick_context/dev_notescolumns. - Panel: a copyable task-id chip with a stable, non-shifting task header, clickable Branch / PR links with a branch-copy button, and clearer agent status badges.
- Panel: the per-project Conventions editor lays out in a responsive two-column grid (Module boundaries | Rules, then Waivers | Custom rules) with Recent violations full-width, inside a wider modal on large viewports — instead of one long single column. Each row's two cards share an equal height, and the Module-boundaries list scrolls internally so it matches the Rules card instead of running long. It collapses to a single column on mobile and is capped so it stays sane up to a 27" display.
[0.8.0] - 2026-06-20
Added
- In-path PR-review gate — every assembled PR is reviewed before the PM merges. A new
awaiting_pr_reviewstatus sits between the work and the PM merge: the cell PM'ssubmit_upopens the cell→root PR and the Main PM'ssubmit_rootopens the root→master PR, and each entersawaiting_pr_review, where a reviewerpr_passes it on to PM review orpr_fails it back for revision — the merge-level reject the PM previously lacked (motivated by a front-end/back-end seam bug that slipped straight through to master). Three new team-scoped cell PR-reviewers (backend, frontend, UX/UI) join the existing main reviewer, taking the company to 25 agents, each with its own first-class image and spawn manifest; leaf developer tasks and branchless coordination roots skip the gate. Ships migration 040 (theawaiting_pr_reviewenum value) plus the panel surfacing: a legible PR-review status badge, a dedicated "PR Review" kanban tab, and a PR-review column on the management board. - Panel test gate. The Next.js panel gains a baseline vitest suite over its lib and stores, a
pnpm teststep enforced in the CI panel job, and amake panel-gatetarget, so panel changes are quality-gated the way the Python side already is.
Changed
get_team_metricsreuses the sharedACTIVE_STATUSESconstant instead of re-listing the active task statuses inline, keeping the definition in one place.
Fixed
- The self-healing CI signal is now deterministic. The regression watch defaulted to the latest completed run across all of the repo's workflows, so on a multi-workflow repo an unrelated green run — or a green run on an older commit — could mask a red CI run and the loop fired only intermittently. The signal is now scoped to the
ci.ymlworkflow by default, pulls a window of recent completed runs and resolves the conclusion against the branch's current HEAD (a green re-run supersedes the failure; a stale green run can't hide it), and retries transient GitHub errors instead of reading one network blip as all-green. - Self-heal fix tasks are assigned to the Main PM agent, not just the
main_pmteam. A team-only task fell to slow unassigned-team routing after the CEO approved it; it is now assigned to the Main PM agent up front so the orchestrator dispatches it straight away once approved. The confirmed-by-human hold that keeps the task inert until CEO approval is unchanged.
Security
- bash-guard denies git verbs hidden in command substitutions — a
$(...)- or backtick-wrapped git command could previously slip past the guard. - Transcript retention matches the encoded workspaces root at a path boundary, so a sibling directory sharing a name prefix is no longer mistaken for the workspaces root during pruning.
- The v1 role guard binds to a verified agent token before trusting a role claim, so the role a request asserts is checked against its signed token rather than taken at face value.
- pydantic-settings upgraded to 2.14.2 to pull in the fix for GHSA-4xgf-cpjx-pc3j.
[0.7.0] - 2026-06-19
Added
- Grok agents on xAI's official
grokCLI, on a SuperGrok subscription. A newroboco/llm/providers/seam (anAgentProviderlifecycle ABC + aProviderRegistrykeyed byModelProvider) lets the orchestrator drive agent backends other than Claude Code, and the first is Grok — running xAI's officialgrokCLI authenticated by a SuperGrok subscription rather than a metered API key, so a Grok workforce can't stall mid-task on out-of-credits. It reaches parity with the Claude path by construction: the same MCP gateway + tool-manifest wiring, per-role tool removal and git-operation deny rules, a prompt-injection guard on the task prompt, headless tool auto-approval, and per-agent token/cost capture from the grok session store. It covers both one-shot delivery roles and the interactive Intake (Prompter) and Secretary chats (per-turngrok -pwith session resume, streamed turn-by-turn). The change is purely additive — onlyGROKroutes through the registry; Anthropic / Ollama Cloud / self-hosted spawns are untouched — and ships migration 038 (thegrokenum) + 039 (the seeded provider row), first-classroboco-agent-grok/-prompter/-secretaryimages wired into all three compose files and the release workflow, and a Settings provider card. - SuperGrok token auto-refresh. The grok access token has a fixed ~6h server-set TTL and the CLI cannot refresh it headlessly — on an expired token it hangs forever at an interactive login prompt — so the orchestrator now mints a fresh token from the offline-access refresh token (xAI's OIDC
refresh_tokengrant) before expiry and rewrites the sharedauth.jsonin place, keeping every Grok agent's credential live with no recurring manualgrok login. As a backstop the agent entrypoint refuses to start (exit 78) on a missing or expired token instead of hanging. - Self-healing CI loop (default-off). RoboCo can now watch its own repository's CI and, on a detected regression, open a fix task that is held out of dispatch until the CEO approves it — then dispatch it through the normal delivery flow, so the company repairs its own breakages. It is dormant by default and armed from two Feature-Flags panel toggles; the CI signal is scoped to a single named workflow, and task origination is bounded by rolling and per-cycle caps so it can't flood the backlog.
- Company Scorecard. A company scorecard on the panel's Business Goals tab.
Fixed
- The PR-reviewer is no longer wedge-killed before it can post a review.
pr_review_claimnow seeds the claim heartbeat like every other claim path; without it a Grok reviewer was treated as a silent (NULL-heartbeat) wedged container and killed before it could callpost_pr_review, churning the task back to pending in a respawn loop. - Grok one-shot runs are observable, and their usage is captured. The entrypoint streams agent activity to the container log live (
--output-format streaming-json) instead of buffering it to a file until the run ends, and per-agent token/cost is read from the grok session store's actual cumulative-total field (it was silently reading$0). - Path-injection hardening of the Grok usage directory. The agent id is validated and reduced to a single safe path component before it is used to build the per-agent usage path, on both the write/mount and finalize-read sides.
[0.6.0] - 2026-06-17
Added
- Inbound PR review — the org reviews, and can take over, pull requests it didn't open. A new read-only
pr_reviewerrole (a 22nd agent, its own first-class image and spawn manifest, migration 037) discovers inbound PRs, reviews the diff adversarially, and posts a single complete change-request as a real GitHub review on the PR itself — no agent-to-agent chatter. It covers external / fork PRs, gated by a configurable author allowlist, and — behind a second flag — internal org-repo PRs opened outside the agent task-flow (the org's own in-flight integration PRs are skipped, since a live task already owns their branch and they pass QA + PM review). Re-review is driven by the PR's head commit (an unchanged PR is skipped, new commits open a fresh review), and polling is repo-aware so a monorepo is no longer reviewed several times over. External-PR review is enabled by default in the shipped compose (with human-confirm on); internal-PR review is off by default. Both are flippable from the panel. - CEO decision queue + supersede for reviewed PRs. Completed reviews surface in a PR-review queue in the panel — in-flight reviews are shown too, linking to the PR, so it never goes dark. From there the CEO can dismiss a review, or supersede the PR: the system cuts a roboco-owned branch off the contributor's commits and opens a Main-PM coordination task to finish and harden the work to our standards on that branch, open our own PR, and — once that replacement actually merges — close and link the contributor's PR. We never push to a contributor's fork.
- Feature-flags panel. A Settings → Feature Flags card toggles env-gated subsystems (external / internal PR review, web research, the strategy engine, pitch provisioning, RAG auto-update, transcript pruning) from the panel instead of hand-editing environment variables. A toggle persists in the existing settings store and takes effect on the next backend restart; an unset flag falls back to its environment / config default, and secrets (API keys, tokens) are never surfaced to the client.
- Required-cells decomposition gate. When a coordination task names the cells that must deliver it, the Main PM can no longer go idle having silently dropped one —
i_am_idleis rejected until every named cell has a subtask, and the Main-PM prompt now insists on honoring explicitly-named cells rather than quietly dropping them. Inert until the marker is set, so existing flows are unaffected.
Changed
- Run from pre-built registry images. A standalone registry compose runs the full stack — including every per-agent image, now extended to the Secretary and the PR-reviewer — from published images rather than a local build. Both deploy paths, the registry knobs, and measured idle / under-load resource usage are documented.
- Documentation, for humans and agents. The how-to guide is now a structured, multi-chapter walkthrough under
docs/how-to/with a new business-workflow chapter (charter → Cockpit → Secretary, and the research / strategy / PR-review toggles); the published reference docs (README, usage, deployment, CLAUDE) were refreshed against the current code, with a CI guard that keeps documentation prose single-line. Agents also get richer in-context guidance: new RAG role docs for the Prompter, Secretary, and PR-reviewer, the 0.4.0 company layer (goals / research / strategy / provisioning) documented for them, and a refreshed guardrails surface.
Fixed
- The PR-reviewer no longer respawns in a loop. Without a spawn manifest the reviewer had no flow verbs, so it could never claim its review and was respawned over and over (burning tokens); it now ships a role-scoped manifest and reliably claims its work. The supersede close-on-land path was also hardened — the contributor PR is retired only once our replacement PR has actually merged, not merely when the umbrella task completed.
- A CEO-rejected coordination root no longer deadlocks. A product-linked coordination task the CEO sends back to needs-revision is re-dispatched to its owning PM (and the readiness gate now accepts a PM on a coordination root in that state), instead of sitting unowned forever because the developer dispatcher skipped it.
- Panel UI standardization + usability pass. A panel-wide pass plus targeted fixes: the Settings grid layout, the Journals and Kanban scroll regions, the agent-list item, the Projects table, kanban cards whose text overflowed, the PR-review queue's empty state, and the Secretary chat composer buttons.
Internal
- DB-backed and httpx-mocked test coverage for the inbound-PR read and lifecycle paths (ingest / dedup / classify / claim / complete / supersede); a cyclomatic-complexity refactor of the git PR-creation and lifecycle-validation code to clear the xenon gate; and the cell-PM / main-PM role docs corrected to the real
delegatesignature and cross-linked to each other.
[0.5.0] - 2026-06-16
Added
- Acceptance-criteria & decomposition guardrails. Every task's acceptance criteria now carry stable per-criterion ids, and each decomposed subtask records which parent criteria it is responsible for (
covers_parent_criteria). Two gates build on that linkage: a PM can no longer go idle leaving a parent criterion with no subtask responsible for it (the decomposition floor), and a parent can no longer complete / submit up / escalate to the CEO unless every one of its criteria traces to a child that passed QA on it (the roll-up gate). PMs see live coverage in their briefings (parent_ac_coverage,unclaimed_parent_acs) after eachdelegate. Safe-by-construction: every gate stays inert until a PM starts declaring coverage, so existing decompositions are never blocked. (Migration 036.) - Per-dev sequenced code queues. A cell PM now delegates each developer its full queue of code subtasks up front instead of one task at a time. Both cell developers build in parallel, and each works its own queue one task at a time, in order — enforced by a per-lane dispatch barrier, with leaf PRs still merged in sequence into the shared cell branch. The old "two code subtasks per parent" ceiling is removed; the 12-subtask hard cap and a same-title duplicate guard remain.
- Unified Business page. The Company Goals, Secretary, and Pitches pages are consolidated into one tabbed Business page (Goals / Secretary / Pitches), modeled on the Knowledge Base page with deep-linkable
?tab=URLs. A single sidebar entry replaces four.
Changed
- Company Goals, Secretary, and Pitches brought to the panel's standards. Skeleton loading and offline/error states, structured fields instead of raw JSON dumps, required-note confirmation dialogs for pitch and directive decisions, and markdown rendering in the Secretary chat.
Removed
- The standalone Cockpit page. Its data duplicated the Dashboard and Metrics; its one unique element — the strategy-engine "needs your attention" signals — was relocated to the Dashboard, served by a new lightweight
GET /api/cockpit/signalsendpoint. The/cockpit,/company-goals,/secretary, and/pitchespanel routes are all retired (404); the Goals, Secretary, and Pitches views now live under/business?tab=….
Fixed
- Agent MCP/SDK servers no longer stall on spawn. They launch with
uv run --no-sync, so a workspace clone whose lockfile has drifted from the baked image no longer triggers a multi-minute dependency re-sync that left the gateway tools stuck "pending" and the developer respawning in a loop. open_prno longer fails on a missing base branch.create_prauto-creates and pushes the PR's base branch off the default branch when it is not yet on the remote, instead of returning a GitHub 422.- Admin status overrides restore task ownership. Forcing a blocked task back to pending / in_progress now restores its pre-block assignee, so an escalated code task no longer re-enters the pool still owned by a PM and is dispatched to that PM as if it were a developer.
- A developer can idle past its own queued work. With per-dev queues, a dev whose current leaf has moved to QA now idles cleanly while its later queue items wait their turn (the orchestrator respawns it when the lane clears), instead of looping on the idle guard or claiming the next leaf out of order.
- 26 verified panel UI bugs across the dashboard, kanban, task detail, and API layer: consistent priority labels and badge sizing, dark-mode coverage, kanban drag-and-drop that prompts for the required audit note, auto-scroll in the message and mentor-chat views, corrected WebSocket reconnect counting,
PATCH(notPUT) for partial task updates, working "Activate Task" and "Start Revision" actions for backlog and needs-revision tasks (no more dead-end menus), the previously-dead "New / Generate Report" buttons, a duplicate agent id, a "0h ago" timestamp, and more.
Internal
- Verb-table generation no longer emits tables for the driver-based roles (prompter, secretary), whose real tools live in their SDK drivers rather than the gateway verb surface; and the
_briefing_fortyped stub was aligned with its implementation so the composed choreographer type-checks under full mypy.
[0.4.0] - 2026-06-15
Added
- Business Goals — the company charter. A single CEO-owned charter (north star, prioritized objectives, constraints, operating policy) injected compactly into every agent's briefing so all work is goal-aware.
GET /api/company-goals(any agent) /PUT(CEO-only), with a panel editor. - Web research for the Board and PMs. Pluggable
web_search/web_fetchexposed through aroboco-searchMCP server backed by/api/research/*, with Tavily / Brave / Exa adapters and a graceful no-op when no provider is configured. The provider key stays server-side — agent containers never make the external request themselves — and a per-agent daily quota (Redis, fail-open) bounds cost. - Pitch → approve → provision. The Board proposes a product (a "pitch"); on CEO approval the system provisions a GitHub repo per target cell, registers a Project for each (and a Product when multi-cell), and seeds one Main-PM delivery task — reusing the existing Product / coordination-task machinery. Default-off: with no provisioning token configured, approval is refused and nothing is created.
- Autonomous strategy engine (dormant). An optional second engine that watches the company against its standing goals and surfaces drift, idle, and long-stranded blocked work to the CEO (notify-only — it never spends, builds, or auto-approves). Off by default; the delivery lifecycle is unchanged.
- The Secretary — the CEO's chief-of-staff. A live conversational agent (its own role, distinct from the Prompter) the CEO chats with in the panel. It acts only under the CEO's command: it reads company state and relays dictated messages directly, but high-impact actions — editing the charter, starting / cancelling / overriding tasks, approving a pitch, announcements — are queued and run only after the CEO's explicit confirmation (the gate list). Its authority is HMAC-scoped to the secretary role and routed through the existing enforcement, never a parallel permission model.
- The Cockpit. A read-only
/cockpitview answering "is the business winning, what's happening, what needs me" — the charter, delivery counts, 30-day spend vs the budget cap, pending pitches, and the strategy engine's signals. Honestly stampedbasis: proxy(a proxy until real launches).
All of these are additive and opt-in or default-off — an unconfigured deployment behaves exactly as before.
[0.3.0] - 2026-06-15
Added
- In-house RAG engine. Replaced the piragi/torch retrieval stack with an in-house pgvector engine (asyncpg), then added hybrid retrieval — pgvector cosine fused with Postgres full-text ranking — retiring HyDE, plus an embed-once / concurrent-search pass that cut multi-index query latency.
- Self-hosted LLM provider with dynamic model discovery, so agents can run against a local or self-hosted model endpoint.
- Quality gates at the source. Developers run a fast quality gate at
i_am_doneand the full fast gate (including complexity) at their desk; QA requires a per-acceptance-criterion verdict before passing; cells run two developers in parallel with split-before-claim sizing. - Board redraft loop — the Board can send a drafted task back to intake for an in-context re-draft before it starts.
- Transcript retention — a background sweep prunes old agent transcripts, with a panel-tunable retention window.
tests/type-gated under mypy — the whole test suite now type-checks in CI.
Fixed
- PR-divergence respawn-loop meltdown. Capped the PM respawn loop-gate, added CEO god-mode status override, a PR-conflict auto-resolver (rebase → close-superseded / re-merge / escalate), and sequence-ordered sibling merge; the dispatcher can now claim an ownerless
awaiting_pm_reviewtask without transitioning it. - Git robustness. Fall back to a permitted merge method when the repo refuses the requested one, and retarget a PR's base to the default branch when the resolved base is missing on the remote.
- RAG outage. Migrated the live
chunks_*tables to the in-house schema (offline-renderable migration), closed engine audit gaps, decoded jsonb metadata returned as a string by asyncpg, and kept the embedding model resident to stop ingest timeouts. - Panel. Fixed task lifecycle (updates, merge, reassignment, copy), responsive grids + mobile overflow, the status dropdown duplicating the current status, the orchestrator-status reachability signal, and surfaced the CEO "Approve & Start" gate so it can't be missed.
- Usage attribution. Agent transcripts are attributed by an orchestrator-assigned session id, fixing zeroed token/cost capture for review-role agents.
- Composed the prompter role layer for the intake agent; aligned auditor channel permissions; made the app route-registration test robust to FastAPI 0.137; cleared an xenon complexity failure and fixable test warnings.
Security
- Documented that WebSocket authentication is REST-only and
/ws/systemis unauthenticated.
[0.2.0] - 2026-06-11
Added
- Provider rate-limit handling. End-to-end backpressure for LLM-provider 429s: a Redis-backed
RateLimitStateTracker, a spawn gate that queues (never drops) work while a provider is rate-limited, agent parking viai_am_blocked(reason="rate_limited"), and a background probe-and-resume loop that auto-revives parked agents when the limit lifts — escalating to the CEO after repeated failed probes. Surfaced live in the panel via a rate-limit banner. - Token usage & cost analytics. Per-agent-session token capture read from the Claude Code transcript (
/usage/sync), persisted to spawn-session rows and daily rollups, with provider-aware pricing (Anthropic models priced; local/Ollama models intentionally $0). Visible on the usage dashboard. /ws/systemoperator WebSocket stream with awebsocket_bridgethat forwards system events from the event bus to panel clients in real time — the rate-limit lifecycle and live token/cost usage (USAGE_UPDATE/USAGE_SNAPSHOT), so the dashboard's "Token Usage & Cost" panel updates over the socket and falls back to HTTP polling when it drops.
Fixed
- Agent workspaces now install the project's
devextra (uv sync --extra dev) so spawned agents have the fullmake qualitytoolchain (ruff/mypy/xenon) and can gate their own work — closing the gap that let lint/type/complexity debt merge unchecked. - Token-usage capture: the dashboard previously recorded zeros because nothing populated the per-session counters.
- Panel rate-limit endpoint shape (
/api/system/rate-limitsreturns the{ entries: [...] }envelope the dashboard expects) and the doubled/ws/ws/systemWebSocket path. - Control-panel logo and all
/publicassets returning 500 — the panel image copied them without chowning to the non-root runtime user. - Provider-aware pricing (Opus corrected to $5/$25 per 1M; non-Anthropic models no longer warn or mis-price).
[0.1.0] - 2026-06-09
Added
- Initial public release of RoboCo — an open-source AI agent "company": a virtual organization of 20 AI agents and 1 human CEO that plans, builds, reviews, documents, and ships software.
- Organizational hierarchy: on-demand Intake, Board (Product Owner, Head of Marketing, Auditor), Main PM, and Backend / Frontend / UX-UI cells.
- Task Assistant (the intake Prompter): a live, codebase-aware chat that interviews the CEO and drafts a well-formed, board-ready task — objective, per-cell breakdown, and acceptance criteria — then launches it into the lifecycle (Board review, or straight to the Main PM).
- Agent gateway (
roboco-flow,roboco-do) backed by the server-side Choreographer; intent-verb tool surface per role. - Task lifecycle state machine with role-based transitions and git workflow (PR-before-QA, CEO approval for major work).
- A2A protocol, journals, channels/notifications, kanban, and RAG (piragi + pgvector) knowledge base.
- Next.js control panel (
panel/) behind a single nginx entry point. - Multi-agent workspace management with per-project encrypted git tokens.
[0.5.0]: https://github.com/rennf93/roboco/compare/v0.4.0...v0.5.0 [0.4.0]: https://github.com/rennf93/roboco/compare/v0.3.0...v0.4.0 [0.3.0]: https://github.com/rennf93/roboco/compare/v0.2.0...v0.3.0 [0.2.0]: https://github.com/rennf93/roboco/compare/v0.1.0...v0.2.0 [0.1.0]: https://github.com/rennf93/roboco/releases/tag/v0.1.0