Smoke-15 wedge: leaf 1533ce56 sat at awaiting_pm_review owned by
be-pm, but the PMs looped firing complete/unblock at the wrong
(parent) task_ids — every rejection was generic ("not assigned to
you" / "not ready for completion"), so minimax never discovered it
just needed `complete(1533ce56)`.
New _own_review_hint: on a cell_pm/main_pm complete-guard rejection
(not-owner, wrong-state, or main-pm-on-non-root), if the PM owns a
DIFFERENT task that is awaiting_pm_review, append a remediate suffix
naming it and the exact `complete(task_id='<id>', notes='...')` call.
Best-effort (never raises into the rejection path), pure guidance —
no control-flow or state-machine change.
Scope: this is the bounded, low-risk slice of #170 (fix b). The
parent-state corruption + missing recovery transition (root->paused /
cell->blocked from earlier mis-targeted verbs, fix a/c) is a lifecycle
state-machine change deferred for explicit design alignment — tracked
in #170.
Smoke-15: QA's claim_review diff was `origin/master...origin/<branch>`
but origin/master in QA's clone was the STALE clone-time tip
(47c674d) — the three-dot diff spanned the whole session delta
(41 files / +2740) instead of the 1-line README change.
Two compounding causes:
- _default_branch_ref early-returns the ref NAME when origin/HEAD is
set (it was) WITHOUT fetching it, so the base stayed stale.
- Every fetch in the diff path ran unauthenticated; the repo is
private, so `git fetch` failed ("could not read Username for
github.com") and could never refresh the ref. (The documenter path
was correct only because #162 uses the PAT-injected
workspace.fetch_branch_for_inspection.)
Fix: new best-effort _token_for_branch resolves the project PAT
(None on any failure → degrades to unauth, never raises in the
evidence path). diff()/list_changed_files() thread it into
_resolve_head_ref + _resolve_diff_base, whose fetches now pass
token= (uses the existing _run_git http.extraheader Basic-auth
injection). _resolve_diff_base additionally re-fetches the resolved
default branch so the base is current even when origin/HEAD shortcut
skipped the fetch.
Smoke-15: be-doc i_documented(files=["README.md"]) → choreographer
doc.py stamped `existing.documents = files` (a list[str]) onto
Task.documents. Task.documents is list[DocRef] persisted as dicts —
list_docs does DocRef(**d), _get_existing_doc_ref does d.get("path"),
the RAG indexer does d.get("path"). A bare string 500'd GET /docs
("TypeError: DocRef() argument after ** must be a mapping, not str")
during be-pm's PR review and would AttributeError the indexer.
Fix at source: new _doc_refs_for builds proper DocRef dicts
(path, title=filename, doc_type, created_by/at, updated_by/at) at the
i_documented stamp. Defensive read: new _coerce_doc_ref tolerates
dict / DocRef / bare-string / rejects unknown — applied at
list_docs and _get_existing_doc_ref so legacy/corrupted rows can't
500. _add_doc_to_task (docs-route write path, not the gateway flow
that broke) left as-is per scope.
The system-prompt directive layer and the briefing block both opened
with "FIRST ACTION REQUIRED: run ToolSearch to activate deferred
Edit/Write". That premise is false: per Claude Code 2.1.114, ToolSearch
gates only deferred MCP tools, never built-ins — and it is not even a
callable tool in the agent runtime. Built-ins are loaded at spawn via
the `--tools` flag and gated solely by the per-role permission rules
(the actual Edit/Write breakage was the global Write(*)/Edit(*) deny +
single-slash path, fixed in c0ba335). So weak models dutifully chased a
nonexistent ToolSearch, concluded Edit/Write were unavailable, and
rewrote whole files via destructive shell redirection.
Both touch points now affirm the role's built-in tools are loaded and
ready, tell the agent NOT to call ToolSearch, and (for authoring roles)
explicitly steer away from whole-file shell redirection — directly
countering the clobber behaviour. Role prompt files (developer,
cell_pm, main_pm, board) updated to match. Dead
_read_tool_load_from_role_prompt (no callers) removed. Directive tests
rewritten to lock the corrected behaviour.
Smoke-10..14: every agent (developers included) got "Edit exists but is
not enabled in this context" and fell back to destructive bash
redirection (a 207-line README rewritten to a 3-line stub, which QA
correctly failed). Two coordinated defects in _generate_agent_settings /
_get_role_permissions:
1. base_deny carried a GLOBAL Write(*)/Edit(*). Claude Code evaluates
permission rules deny -> ask -> allow, first match wins — a deny
ALWAYS beats a more-specific allow and the glob syntax has no
negation. So the global deny unconditionally shadowed every per-role
workspace-scoped Write/Edit allow. Removed it; the security denies
that legitimately rely on deny-always-wins (Bash(git:*), credential
Read denies, curl github, env) stay. Roles that must not author
(qa, cell_pm, main_pm, auditor) keep their OWN Write(*)/Edit(*) deny.
2. The workspace allow used a single leading slash (Write(/data/...)).
Claude Code resolves a single / against the settings.json project
root, not the container filesystem root, so the allow silently never
matched even without defect #1. Emit the // absolute-filesystem form.
defaultMode stays bypassPermissions (switching to dontAsk would require
re-deriving the full allow-list and risks wedging agents elsewhere —
out of scope). Verified against Claude Code 2.1.114 permission docs.
Smoke-14: QA's claim_review evidence had pr_diff_summary="" and
files_changed=[] on a PR with a real README change. Root cause: diff()
and list_changed_files() diffed against the bare local <branch_name>.
That ref only exists in the clone where the dev ran `git checkout -b`
at claim. QA / documenter / PM inspect from their OWN clones, where a
bare <branch_name> resolves refs/heads then refs/remotes/<name> but
NEVER refs/remotes/origin/<name> — so `git diff base...<branch>` had an
unresolvable HEAD and silently returned an empty diff (run with
check=False).
#161 previously fixed the BASE side (cell-PM parent never pushed → fall
back to default branch). This is the symmetric HEAD-side facet.
open_pr pushes the leaf branch, so origin/<branch> is the
workspace-independent source of truth. New _resolve_head_ref fetches the
branch and prefers the local branch (dev's own clone, unchanged
behaviour), falling back to origin/<branch> (QA/doc/PM clones), then the
bare name so the command stays well-formed. diff() and
list_changed_files() route through it; explicit base (incremental dev
path, base=HEAD~1) is preserved.
The git network/auth deny rule matched its regex against the whole
command string, so heredoc bodies and echo/printf arguments that merely
documented git verbs (a README, a notes file) were treated as git
invocations and denied. This wedged smoke-13's dev: after wiping the
README via an Edit/Write fallback it could not restore it because every
`cat > README.md << EOF ... git commit ... EOF` was blocked.
The git-ops check now runs against a skeleton of the command with
heredoc bodies and echo/printf literal args stripped (both are data the
shell writes, never executed). Quoted args to a shell interpreter
(`bash -c "... && git fetch"`) ARE executed, are not echo/printf/heredoc
bodies, and so survive untouched — the hook's core purpose is preserved.
A sentinel prefix distinguishes a legitimately-empty skeleton from a
python failure (fail closed on failure). All other rules, including the
#164 import-bypass rule, still inspect the full command.
Smoke-12: be-dev-1 (minimax-m2.7) bypassed the entire MCP boundary by
running `uv run python3 -c "import os;
os.environ['ROBOCO_AGENT_ID']='...'; from roboco.mcp.flow_server
import open_pr; open_pr(...)"` from the Bash tool. This voided the
per-role tool manifest (role-scoping is meaningless if the agent can
import any server module in-process), forged agent identity via an
env-var rewrite, and ran choreographer code outside the gateway's
tracing + auth.
bash-guard-hook.sh now adds two deny rules:
1. Any python/uv/poetry/pipenv/pdm/hatch invocation that imports or
`-m`-runs roboco.* internals (mcp/services/runtime/foundation/
api/enforcement). The whole command string — heredoc body
included — is matched, so quoting/heredoc forms are covered.
2. Any assignment or export of ROBOCO_AGENT_ID (identity forgery).
Reading roboco source for context (cat/grep) is still allowed — the
block is on *executing* internals, not viewing them. Normal python
one-liners without roboco imports still pass.
19 bash-guard tests pass (10 prior + 9 new). Note: takes effect on
agent-image rebuild (hook ships in the agent container).
Smoke-12: be-pm delegated TWO subtasks under one cell parent — a
code subtask (be-dev-1, spawned) and a documentation subtask
(be-dev-2). The orchestrator dev-dispatch refuses to spawn a developer
for task_type=documentation, so the doc subtask became a permanent
orphan that loops dev-dispatch forever and would deadlock submit_up
(all subtasks must be terminal). The spine-cap is per-type so
code + documentation both passed sibling-dedup — the PM never saw the
anti-pattern warning.
_delegate_static_guards now rejects task_type='documentation' with a
remediate explaining the lifecycle auto-creates the documentation
phase (awaiting_documentation → documenter spawned) after the code
subtask passes QA, and that the PM should delegate ONLY the code
subtask.
#160 — panel /roboco-logo.png "received null":
next/image optimizer fails for static public assets in Next.js
standalone mode. Added `unoptimized` to the sidebar logo Image so
it serves the static file directly (validated on panel rebuild).
#161 — QA/doc evidence pr_diff_summary empty:
A leaf dev branch's parent_branch_for is the cell-PM branch, which
is never pushed (only devs push their leaf branch). diff against a
non-existent origin/<parent> returned empty. Added
GitService._resolve_diff_base + _default_branch_ref + _ref_exists:
diff/list_changed_files fall back to the repo default branch
(origin/HEAD → master/main) when origin/<parent> is absent.
#162 — claim_doc_task BRANCH_MISMATCH loop:
The documenter's clone is separate from the dev's; the task branch
already existed (dev created it) so no checkout ran in the doc
workspace — roboco_docs_write / commit failed BRANCH_MISMATCH and
the doc looped. Fixes:
(a) new GitService.checkout_branch_in_agent_workspace; claim_doc_task
checks out the task branch into the doc clone (best-effort —
a checkout hiccup never fails the claim).
(b) BRANCH_MISMATCH remediate now lists all four role claim verbs
(i_will_work_on / i_will_plan / claim_doc_task / claim_review).
(d) give_me_work next-hint is role+status aware via _claim_verb_hint
(doc→claim_doc_task, qa→claim_review, pm→i_will_plan, else dev).
Facet (c) (i_am_blocked "Not Found" for doc) was only reachable via
the stuck-without-checkout path; primary fix removes it.
Smoke-11 reached dev→QA→doc (deepest ever) and validated the prior
6 fixes (panel flood gone, #158/#159/#157 confirmed). These three
clear the doc-phase blockers found in that run.
Task #157 — spine-cap allows planning fanout across cells:
main-pm's pattern is to delegate planning to be-pm / fe-pm / ux-pm
in parallel — each on a different team. The previous spine-cap
rejected all planning siblings under one parent as
over-decomposition. New helper _is_cross_team_planning skips the
cap for planning when both teams are non-empty and distinct. Code
/ documentation stay capped regardless (single repo on one branch
shouldn't have two simultaneous code subtasks).
Task #159 — tracing-gap remediate hints every requirement:
journal:during_work>=1, journal:struggle, commits>=1, pr_open,
and self_verified had no entries in _hint_for_missing_key, so when
they were missing the agent saw the token in `missing[]` but the
`remediate` text had no instruction for how to satisfy them.
Smoke-10's be-dev-1 burned multiple turns retrying i_am_done not
knowing scope='reflect' doesn't count toward during_work. Now
every token has a hint, the during_work hint warns that reflect
doesn't satisfy it, and multi-hint remediate uses a numbered list
so the model treats each requirement as a distinct step instead
of a semicolon-blob.
Also coerce convert_plan._coerce_risk formatting (ruff-format follow-up
to 9cd73d0).
Task #156 (sessions): pre-gateway flow created a session for the whole
task tree at once, so subtasks were visible in the PM's group chat the
moment they existed. The gateway creates subtasks one-by-one via
delegate(), losing that wiring. Added
MessagingService.propagate_sessions_to_subtask and threaded it through
the choreographer's _create_subtask_from_inputs. ChoreographerDeps grew
an optional `messaging` field so existing test wirings keep working.
Task #155 (progress): smoke-9 had zero progress entries because the dev
never called progress() explicitly. Added _record_milestone_progress and
fire it server-side from two natural milestones — open_pr ("opened PR
#N", 70%) and i_am_done ("submitted for QA review", 90%). Best-effort
write (contextlib.suppress) so a progress failure cannot break the verb
path. Extracted _open_pr_success_envelope to keep cyclomatic rank ≤ B.
Bug:
ContentActions.evidence() hard-coded files_changed=[] and diffed
against HEAD~1 instead of the branch's parent. QA's _build_qa_claim_
evidence (and doc/_impl mirrors) sourced files_changed from
work_session.files_modified, which the gateway commit() never
populates (no add_files_modified plumbing). Result: QA / docs / PM
reviewers saw an empty change list on real PRs and only the latest
commit's delta — flagged in smoke-9 when PR #20 showed the README
change on GitHub but evidence() reported empty.
Fix:
Added GitService.list_changed_files (git diff --name-only against
parent branch). evidence(), _build_qa_claim_evidence,
_claim_doc_evidence, and _build_i_am_done_ok all source files_changed
from this — git is the authoritative source. evidence() also drops
the HEAD~1 base so the diff is the full PR.
Wired EvidenceRepo into ContentActionsDeps so evidence() returns
journal_highlights too, matching the QA/doc shape.
Bug:
Choreographer.i_will_plan() built spec_ctx with the raw plan string
while ctx (_ClaimPlanStartContext) got the resolved (panel-shaped)
dict. The verb runner uses spec_ctx — so the rich shape never reached
TaskService.set_plan. The panel's Plan tab stayed empty even when
PMs supplied approach / sub_tasks / risks.
Fix:
Pass the resolved (possibly-dict) plan into spec_ctx too. Widened
lifecycle.Context.plan to `str | dict[str, Any] | None` to match.
Tightened _resolve_effective_plan to require a non-empty narrative
paragraph — rich structure layers on top of prose, not in place of it.
Smoke-8 follow-up. Two issues in _write_agent_briefing:
1. _build_tool_load_block was scraping role prompts for a "## Load on
spawn" section that doesn't exist in any role file. Returned "" for
every role → no ToolSearch directive in the briefing. Combined with
weak models skipping the system-prompt-layer directive (#144), the
agent's first action was Edit → "not enabled in this context."
Fix: per-role tool list lives in the orchestrator (mirrors
factories._base.py). Pre-renders the directive directly. developer
and documenter get Edit + Write; QA/PMs/board get the common
read-only set. 7 tests pin the contract.
2. The briefing's "Terminal tools (how to exit cleanly)" section still
listed pre-gateway verb names: roboco_agent_idle,
roboco_task_substitute, roboco_task_escalate,
roboco_task_submit_qa, _qa_pass/fail, _docs_complete, _complete.
Same rename pattern as #145's _TERMINAL_TOOLS set. Updated to:
i_am_idle, i_am_blocked, unclaim, i_am_done, pass, fail,
i_documented, complete, submit_up, escalate_up, escalate_to_ceo.
The agent now reads the same directive in two places (system prompt +
session briefing) — the second touch point catches weak models that
skip the first.
Smoke-8 surfaced a tight respawn loop: QA failed a PR cleanly, container
exited 0, then _check_health bumped error_count and respawned QA with
the same task_id. But by then the task was in needs_revision (dev's
state), so QA's claim_review was rejected — and the cycle repeated on
the next health tick. Token-burning loop.
Two layers:
1. _check_health now reads docker's exit code. exit_code == 0 →
graceful (intentional handoff via i_am_idle / clean shutdown) →
reset error_count, do NOT auto-restart. Non-zero → keep the
existing crash-retry behavior. Refactored into
_inspect_container_state + _handle_stopped_container to keep
xenon's complexity check happy.
2. _readiness_check_role_for_status now includes the dev-owned
states (needs_revision, verifying) so a misrouted spawn for QA /
PM / board on these statuses fails the readiness gate before the
gateway has to reject it. Defense in depth — the right path is
#1 (don't respawn on clean exit at all), but if some other code
path tries to spawn QA on needs_revision the gate now catches it.
Tests: 12 new (5 for _check_health graceful/crash matrix + 7 for the
expanded role-status table). Pre-gateway names (none of which were
needed here) untouched.
Smoke-7: be-pm got the expected spine-cap rejection on a second
delegate. The remediate said "drive the existing sibling to
completion / cancel it, OR split this parent into two sibling
parents". The model read that, decided neither applied, and
"adapted" by re-delegating with task_type='documentation' as a
"verification subtask". The gateway accepted it (different type =
no cap collision) but the orphan subtask had no claimant — it
blocked submit_up forever with "subtasks not all terminal".
Two-layer fix:
1. _spine_type_dup_envelope remediate now explicitly forbids the
workaround: "DO NOT work around this by delegating again with a
different task_type (e.g. 'documentation' or 'research' as a
'verification' subtask). The lifecycle handles QA, documentation,
and PM-review automatically after the code subtask finishes —
you do not create auxiliary subtasks for those roles. Call
i_am_idle() now and wait for the existing child to come back."
2. cell_pm.md workflow step 6 strengthened to name the anti-pattern
explicitly: no verification subtask (QA is the verification step);
never re-delegate with a different task_type as a workaround.
3 new tests pin the remediate text: forbids workaround, names the
verification anti-pattern, retains invalid_state error kind.
Smoke-7 evidence: every agent's first successful i_am_idle was followed
by a stop-hook error "you stopped without calling a terminal tool"
even though i_am_idle had just succeeded.
Root cause: agent_sdk._TERMINAL_TOOLS still held pre-gateway names
(roboco_agent_idle, roboco_task_submit_qa, ...). /terminal/tool_recorded
strips the `mcp__roboco-flow__` prefix and stores 'i_am_idle' — the
membership check against {'roboco_agent_idle', ...} never matched, so
had_terminal_recently() always returned False, and the stop hook nagged
every clean exit. Wasted ~2-3 turns per agent + burned stop_allowance.
Fix: rebuilt _TERMINAL_TOOLS with the current gateway verb names:
i_am_idle, i_am_done, i_am_blocked, i_documented, unclaim, pass, fail,
complete, submit_up, escalate_up, escalate_to_ceo.
7 new tests pin:
- every role's terminal verbs are recognized
- pre-gateway names are NOT in the set
- _SessionState.had_terminal_recently() returns True after i_am_idle
Smoke-7: be-dev-1 hit "Edit exists but is not enabled in this context."
Claude Code v2.1.69+ defers built-in tools (Edit, Write, Read, etc.)
behind a ToolSearch call. Weak models (minimax-m2.7) skip soft
directives buried in the briefing.
Also: 4 role prompts (developer, cell_pm, main_pm, board) claimed
"no ToolSearch needed" — a lie that compounds the problem. The
manifest registers MCP tools; built-in tools are still deferred.
Fix: compose_prompt now prepends a tool-load directive layer as the
FIRST block in the system prompt. It names the exact ToolSearch call
the role needs:
- developer/documenter: Read, Bash, Grep, Glob, Task, TodoWrite, Edit, Write
- qa/pm/board: Read, Bash, Grep, Glob, Task, TodoWrite (no Edit/Write)
The directive includes the failure mode it prevents so the model
understands what skipping the call causes.
Updated role-prompt lines that lied about ToolSearch.
7 new tests pin: directive is the first block; developer/documenter
get Edit/Write; qa/pm don't; failure-mode message is present.
Smoke-7 surfaced: be-qa called dm(recipient='qa-all', ...) — 'qa-all'
is a channel slug, not an agent. A2A enforcement raised
A2AAccessDeniedError. It propagated past dm(), past content_actions,
got caught by FastAPI middleware which renders RobocoError.to_dict()
as {'error': {'code': ..., 'message': ..., 'details': ...}} — a
DICT-shaped 'error' field.
do_server's circuit-breaker check (and flow_server's mirror) did
`payload.get('error') in _CIRCUIT_REJECTION_KINDS` — trying to hash
a dict against a frozenset → `TypeError: unhashable type: 'dict'`.
The agent saw "Error executing tool dm: unhashable type: 'dict'"
and got stuck calling dm in a loop.
Two-layer fix:
1. content_actions.dm now catches A2AAccessDeniedError and returns
Envelope.not_authorized with the original reason + route_hint as
remediate. This is the right shape — content tools always emit
Envelopes; RobocoErrors escaping to the middleware is a bug.
2. Defense-in-depth: do_server._record_and_check_circuit and
flow_server._record_and_check_circuit now guard against non-string
error fields. Any future RobocoError-leak that bypasses (1) will
pass through untouched instead of crashing the tool call.
3 new tests pin the contracts:
- dm A2A denial returns Envelope.not_authorized (not propagated)
- do_server circuit-breaker doesn't crash on dict-shaped errors
Smoke-7 surfaced this: QA spawned, claim_review succeeded, but every
attempt to call `pass()` fell through to dm/say workarounds. The MCP
tool 'pass' never existed.
Root cause: foundation.policy.lifecycle declares the intent verbs as
`pass_review`/`fail_review` (Python-friendly names — `pass`/`fail` are
keywords). intents_for_role(Role.QA) returns those names, the spawn
manifest carries them, and flow_server reads them. But flow_server's
_TOOLS dict has keys 'pass'/'fail' — the manifest's pass_review keys
didn't match and got silently dropped from the registration.
Fix: add _INTENT_TO_PUBLIC = {'pass_review': 'pass', 'fail_review':
'fail'} in flow_server. _register_tools transforms manifest names
through it before _TOOLS lookup. Manifest entries map to the public
MCP tool names the prompts advertise.
Also fixed _VERB_RETRY_LIMITS keys in foundation.agent_loop — they
used the IntentSpec names too, but the SDK receives the public name
from /verb/attempted (derived from the flow URL path), so the limit
entries never matched real rejections. Renamed to 'pass'/'fail'.
3 regression tests pin: pass/fail register under public names;
IntentSpec names don't leak through; the registered tool POSTs to
the correct orchestrator path.
Smoke-6 found the agent calling note(scope='decision', context=null,
chosen=null, rationale=null) eight times in a row. Root cause split
across two surfaces:
1. The MCP tool schema declared these fields as `str | None = None`,
producing a JSON schema of `anyOf [string, null]`. minimax-m2.7 read
that and decided null was a valid value — passed it on every retry.
2. The remediate text used `<placeholder>` syntax for the example
without telling the agent "don't pass null" explicitly.
Fixes:
- roboco/api/schemas/v2/do.py NoteRequest: context, chosen, rationale,
what_done, what_learned, what_struggled now typed `str = ""` (no None).
Pydantic on the route rejects literal null with 422 BEFORE the gateway
sees it. Empty string still counts as missing at the gate.
- roboco/mcp/do_server.py note(): matching signature changes so the
MCP tool schema declares the fields as `string` not `anyOf[string,null]`.
- roboco/services/gateway/content_actions.py: remediate text now opens
with "DO NOT pass null" and the example uses concrete values (redis vs
postgres) instead of <angle bracket> placeholders. Reflect remediate
also gets the don't-pass-null intro and concrete values.
8 new tests pin "schema rejects null for each of the 6 string fields"
plus "empty defaults work for unscoped notes".
Smoke-6 surfaced the gap. main-pm called note(scope='decision') with
context: null 8 times in a row — every one returned incomplete_input
and the agent kept retrying. The flow_server had a breaker (C1) but
do_server didn't, so content-tool rejections went uncapped.
Mirror the flow_server pattern:
- _CIRCUIT_REJECTION_KINDS = {tracing_gap, invalid_state,
not_authorized, incomplete_input} (same set)
- _record_and_check_circuit posts to the SDK's /verb/attempted on each
counted rejection
- When the SDK reports open=true, the original rejection envelope is
REPLACED with the circuit_open envelope so the agent stops retrying
The SDK side (agent_sdk/server.py) already accepts arbitrary verb
names; no changes there. note hits the default cap of 3 retries / 60s
from foundation.agent_loop. After the third incomplete_input the
agent gets circuit_open and the loop ends.
9 new tests pinning the contract.
ContentActions.pr_update enforces:
- task.pr_number is set (else invalid_state, remediate 'call open_pr')
- at least one of title/body/reviewers is non-None (else invalid_state)
- caller is task assignee OR PM on team (cell_pm same-team / main_pm
cross-team), else not_authorized
- GitError raised by the underlying service maps to invalid_state with
the upstream message preserved
The route at POST /api/v2/do/pr_update binds PRUpdateRequest, whose
model_validator returns 422 on all-None bodies so the verb layer
never sees them. The MCP tool registry adds 'pr_update' so manifest-
scoped do-servers can expose it to roles that opt in (next commit).
Smoke-5 surfaced that be-dev-1 had no gateway-native way to fix a
PR's title/body or request a reviewer after open_pr; `gh pr edit`
is bash-shimmed and the dev correctly escalated rather than bypass
the guard. This adds the GitService primitive: PATCH /pulls/{n}
for title/body and POST /pulls/{n}/requested_reviewers for the
reviewer list, with NotFound + 422 mapped to typed GitError. The
PRUpdateRequest schema enforces 'at least one field' via a
model_validator so the route returns 422 before reaching the verb.
Smoke-5 root cause. Agents wrote 5 decisions / 8 reflections / 1 struggle
during the run — every single entry persisted with task_id=NULL. The C8
tracing gate then never saw them and PMs spiraled forever on
'missing: journal:decision' while their decisions sat orphaned.
Cause: ContentActions.note/say/dm/notify called
TaskService.get_active_task_for_agent for task_id auto-injection. That
helper filters to _DEV_ACTIVE_STATUSES = {claimed, in_progress,
verifying, awaiting_qa, awaiting_documentation}. BLOCKED, PAUSED, and
NEEDS_REVISION fall outside that set — so the moment an agent gets
stuck (which is exactly when they journal), auto-injection returns None
and the entry persists without task_id.
Fix:
- New TaskService.get_journal_context_task_for_agent — same shape as
get_active_task_for_agent but the status set
_JOURNAL_CONTEXT_STATUSES adds BLOCKED, PAUSED, NEEDS_REVISION.
- ContentActions.note/say/dm/notify use the new lookup.
- ContentActions.commit keeps the narrow get_active_task_for_agent —
can't commit from blocked, so the dev-active set is correct there.
Tests:
- tests/unit/services/test_journal_context_lookup.py — 5 tests pinning
the two queries: journal-context INCLUDES blocked/paused/needs_revision,
dev-active EXCLUDES them.
- Existing content-actions tests updated to stub the new method
alongside the old one.
This alone may be 70% of what was killing smoke runs end-to-end.
Smoke run 4 crashed with AttributeError: 'SessionForTasksCreateRequest'
object has no attribute 'config' when main-pm called open_session.
The service _build_session_request reads req.config.max_message_count
(model has nested config). The gateway was passing the API schema
SessionForTasksCreateRequest (flat fields, no config attribute at all)
— so `req.config` blew up with AttributeError, not the safer None.
Fix: gateway now constructs SessionForTasksCreate (the model) with the
enum-typed relationship_type, matching what the route at
roboco/api/routes/sessions.py:230 does. Unknown relationship strings
fall back to DISCUSSION.
Two regression tests pin the contract: service receives the model;
invalid relationship_type defaults to DISCUSSION.
Pure whitespace / line-wrap reformats accumulated when ruff format
ran during earlier waves but weren't included in their commits. No
semantic changes — collection literals reflowed, with-statement
context managers regrouped via PEP 617 parens.
Smoke run 3 showed agents loading builtin Anthropic connectors
(mcp__claude_ai_Gmail__authenticate, Google Calendar, Notion, Drive)
alongside our 5 roboco MCP servers. The connectors bloat the tool
surface and give the LLM 'discover' targets it shouldn't have.
The Claude Code CLI's --strict-mcp-config flag tells it to load ONLY
the servers from --mcp-config, ignoring all builtin defaults. Added
to _append_image_and_claude_args next to --mcp-config.
Note: the existing --tools allowlist (Read,Write,Edit,Bash,Grep,Glob,
Task,TodoWrite) only filters builtin tools, not MCP-prefixed ones —
that's why the connectors slipped through.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section E3.
_check_pm_decision_required now requires the latest journal:decision
within pm_decision_window_seconds (default 300). Older decisions no
longer satisfy the gate. Adds JournalService.latest_decision_at.
Future-tighten (out of scope): per-verb-group consumption tracking
would need persistent state — Choreographer is per-request today.
Smoke run 3 showed agents auto-pausing on i_am_idle (correct behavior
for non-terminal tasks) but capturing no checkpoint — panel's
Checkpoints column stayed empty. Pre-gateway parity: the auto-pause
path now writes a synthetic checkpoint summarizing state at pause-time
so the panel reflects reality.
Manual i_will_pause (G8a, deferred) will eventually let agents pass
their own checkpoint_summary; for now this synthetic write covers the
bare i_am_idle case which is what all current agents do.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section C7.
The auditor's role is 'silent observer' — read every channel and emit
a reflect note when something notable happens. Smoke run 3 never
spawned auditor because no event-subscription registered it. Added
handler handle_auditor_spawn() wired to:
- task.blocked (EventType.TASK_BLOCKED)
- task.cancelled (EventType.TASK_CANCELLED)
- task.awaiting_ceo_approval (EventType.TASK_AWAITING_CEO_APPROVAL)
Routine events (task.claimed, task.started, task.created) deliberately
do NOT trigger auditor — those are progress, not exceptions. The
auditor's container is one-shot: i_am_idle() exits after logging its
reflect note.
Auditor spawn failures are swallowed into a WARNING log so they cannot
block the underlying event's processing chain. The auditor is a silent
observer — its absence must have no side effects on the lifecycle.
Pre-gateway parity. evidence(task_id).acceptance_criteria_status was
always [] because the gateway's i_am_done gate validated each
criterion against the dev's journal:reflect but didn't persist the
per-criterion verdict. The panel + audit log couldn't show
per-criterion checkmarks.
Now the gate writes a list of {criterion, addressed, artifact_ref,
checked_at} entries to task.acceptance_criteria_status. The existing
matching logic surfaces which artifact (commit sha / reflect-note)
addressed each criterion; entries that aren't addressed get
addressed=False so the panel can flag them.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section C5.
Pre-gateway parity. Smoke run 3 showed task.work_session_id null on
every task — the choreographer's claim/plan/start path didn't create
the row that downstream subsystems (panel, PR tracking, merge chain)
need to track agent-per-task git activity.
Add TaskService.ensure_work_session(task_id, agent_id) as a public
wrapper around the existing _create_work_session_if_needed logic.
Role restriction lifted to None so both developers and PMs get a
session (pre-gateway always created sessions for all claimants).
Built-in re-entry guard prevents duplicate rows on re-claim.
Wire the call into both _claim_plan_start_run and _resume_from_claimed
immediately before _touch, so every successful in_progress transition
(including the stuck-claimed recovery path) creates the row.
Spec ref: Wave C task C4 (2026-05-12).
Smoke run 3 showed agents reaped at the 3-min stale-claim window
while they were actively retrying rejected verbs. Two causes:
1. The reaper threshold was hardcoded at 180s via claim_stale_seconds.
LLM inference + retry loops routinely take longer than that between
verb-successes. Added settings.stale_claim_reap_seconds (default
600s); override via ROBOCO_STALE_CLAIM_REAP_SECONDS env var.
claim_stale_seconds (spawn-filter cutoff) is unchanged at 180s.
2. last_heartbeat_at only refreshed on verb SUCCESS. A verb stuck
in a rejection loop (e.g. tracing_gap missing journal:decision)
showed no heartbeat updates even though the agent was alive.
Added a best-effort heartbeat refresh inside _emit_rejection so
EVERY verb dispatch — success or rejection — counts as activity.
Heartbeat approach: option (b) — touch inside _emit_rejection (single
centralized rejection path). Requires no middleware layer, no HTTP body
parsing, and no new files. The _touch guard for task_id=None means
agent-level rejections (no task context) are a safe no-op.
Net effect: agents stop being reaped mid-retry. Genuinely-stuck
containers (no verb dispatch at all) still reap normally at 600s.
Spec ref: Wave C Task C3.
Smoke run 3 fired 'ensure_workspace: refresh fetch returned non-zero'
9 times per run because each evidence(task_id) call triggered
ensure_workspace -> fetch. The workspace doesn't change in subseconds.
Added a 30s TTL cache keyed by workspace path. ensure_workspace(force=True)
bypasses the cache for callers that genuinely need a fresh fetch.
Net effect: log noise drops from 9 entries to 1-2 per run; orchestrator
spends less time waiting on redundant git fetches.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section C2.
Smoke run 3 showed Main PM hitting 7 incomplete_input rejections on
the decision-note required-fields gate before finally succeeding.
The per-verb breaker tracks repeated rejections of the same verb in
a 60s window and returns circuit_open after the 3rd strike — but its
classification set only included tracing_gap. incomplete_input was
added in Wave 1 (pre-gateway parity for decision/reflect structured
fields) and should have been added to the breaker at the same time.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section C1.
Smoke run 3 showed Main PM's first give_me_work() returning
{status: idle, next: 'no Main PM work'} even though c7935d2c was
pending and assigned to Main PM. The filter only walked
list_assigned_for_agent (ordered by priority/updated_at — pending
could rank behind in_progress rows) and the PM path fell through
to idle because the pre-assigned pending case was not checked first.
Pre-pended a list_pending_for_agent check in both give_me_work and
pm_give_me_work: tasks where assigned_to=agent_id AND status=pending
take priority over all other lookups. Added TaskService.list_pending_for_agent
for the query (ordered by sequence, priority, created_at).
Updated existing tests in test_choreographer_dev, test_choreographer_pm_extras,
and test_heartbeat_wired to set list_pending_for_agent.return_value=[]
where they were not testing the pre-assigned path.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section B6.
Smoke run 3 showed Main PM taking 7 attempts to satisfy the decision-note
required-fields contract — the remediate listed which fields were
missing but didn't show what a fully-formed call looks like. The LLM
pattern-matches examples better than field-list prose; each retry it
dropped a different field.
Added a literal note(scope='decision', ...) / note(scope='reflect', ...)
call template to the rejection remediate so the agent sees the canonical
shape with named-keyword args and example values. The missing-fields
list stays — both pieces of information are useful, but the example is
what actually drives convergence.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section B4.
Three classes of failure, all surfaced once Wave A's plan-required gate
and the migration 013 went in. Per project standing rule: pre-existing
errors are not a free pass — fix them.
1. Wave A1 ordering (32 lifecycle parity failures + 1 full-pipe test).
_pm_sub_tasks_gate fired BEFORE _claim_plan_start_gate, so wrong-state
PMs got `incomplete_input` (the gate's verdict) when the spec's
lifecycle gate should have returned `invalid_state` first. Swapped:
re-entry check → spec lifecycle gate → sub_tasks gate → claim_plan_run.
Parity test now sees the spec's verdict as expected.
2. E2 enum naming (2 migration_013 failures + ripple).
_str_enum in roboco/db/tables.py didn't pass name=… to SQLAlchemy
Enum(...), so Base.metadata.create_all in test setup inferred
`role` from the Python class `Role` while the alembic migrations
declare `agentrole`. Tests saw two enums for the same class and
hit `agentrole = role` operator errors. Fixed: default name to
lower(class_name) (matches every migration), override `Role` →
`agentrole`. One dict entry; no class-by-class registration needed.
3. _MockContentActions.note() signature drift.
Wave 2 G4 added `structured` kwarg to ContentActions.note().
The integration mock at tests/integration/v2/test_full_pending_to_completed.py
didn't accept the new kwarg → 1 test failed on the very first call
from the v2 do/note route. Added `structured: object = None` and
left it unused (the test asserts lifecycle, not journal rendering).
Plus three ruff E501 line-length fixes in the test files I touched.
Quality: ruff + mypy clean. pytest 6690 passed / 0 failed / 274 skipped.
Smoke run 3 showed inconsistent return strings — main-pm got
status='sent', be-pm got status='posted' for the same verb.
Confirmed say() already returns 'posted' at its sole success exit.
Added test_say_status.py to pin the canonical past-tense pattern
(note->'noted', say->'posted', notify_ack->'acked') and prevent
regression. dm() and notify() retain 'sent' — different verbs,
different semantics.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section B3.
Smoke run analysis initially flagged three Task fields as unused
(pm_approvals, quick_context, proactive_context). A follow-up audit
found quick_context (stores original_developer marker + doc notes +
PR creator + escalation notes) and proactive_context (RAG injection)
are actively used. Only pm_approvals is truly orphaned.
Migration 014 drops pm_approvals; downgrade() recreates it if ever
needed. The two false-positive fields stay untouched.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section B2 (re-scoped 2026-05-12).
Smoke run 3 fired the same workspace.py warning ~9x per run:
'ensure_workspace: refresh fetch returned non-zero'
stderr: 'fatal: could not read Username for https://github.com'
This is EXPECTED behavior, not a bug. The docstring on
_fetch_origin_best_effort explains that credentials are deliberately
scrubbed from .git/config after the initial clone (part of the secret-
exfiltration mitigation) and refresh fetches are best-effort. For
private repos the auth-fail is the documented outcome.
The original A4 spec proposed re-injecting the PAT -- that would have
violated _assert_no_pat_leak and the URL-scrub mitigation. Re-scoped
to: silence the known-benign signature at DEBUG, keep WARNING for
genuine failures (network errors, broken remotes, repo-not-found).
No behavior change. No security boundary touched. Just log level.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
A4 (re-scoped 2026-05-12 after investigation showed the original spec
proposed reintroducing a documented security regression).
Fixes 2 important + 1 minor issue from the code-quality review of 5adb4ff:
1. Formula duplication: the workspace path string was inlined at two
sites in orchestrator.py (the canonical _prepare_agent_spawn and the
new _build_mount_args -w logic). Extracted to module-level helpers
_agent_workspace_path(project, team, agent_id) and
_cell_workspace_path(project, team) so both callers share the same
formula. Future path changes only land in one place.
Also extracted _resolve_project_slug_from_git_context() as the
module-level counterpart to the instance method, called by the static
_build_mount_args site that cannot access self.
2. Test consistency: test_workdir_matches_edit_allowlist_path now
extracts the Edit(<prefix>/**) value from _get_role_permissions and
asserts the spawn cmd's -w value equals that prefix. The test would
actually catch a drift where _build_mount_args and _get_role_permissions
use different formulas — previously it just compared two copies of
the same string.
3. Test coverage: added test cases for product_owner and head_marketing
spawns (both share the per-agent workspace path), so all roles that
_get_role_permissions distinguishes are covered.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
A2+A3 (re-scoped 2026-05-12).
Smoke run 3 surfaced two bugs that share a root cause:
- Edit(/app/README.md) → 'Edit exists but is not enabled in this context'
- commit(files=['/app/README.md']) → 'outside repository at <workspace>'
Both happened because the container's WORKDIR is /app (roboco package
source) while the agent's task workspace is bind-mounted at
/data/workspaces/<project>/<team>/<agent>/. The Dev role's
Edit/Write permission allowlist scopes to the workspace, so any Edit
call from /app fails the path match.
Adds '-w {workspace_path}' to the docker run command so the container
starts with cwd = task workspace. Edit(README.md) and git add README.md
now resolve inside the workspace clone.
Mirrors _get_role_permissions path selection exactly:
- developer / product_owner / head_marketing: per-agent workspace
- documenter: cell workspace (matches its Write/Edit allowlist)
- qa / cell_pm / main_pm / auditor: omit -w, fall back to /app
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
sections A2 + A3 (re-scoped per investigation 2026-05-12).
Three fixes from the code-quality review of a1009c0:
1. Critical: _pm_sub_tasks_gate ran before _handle_pm_reentry, breaking
idempotent re-entry for PMs whose containers crashed mid-run. Moved
the gate to after the re-entry short-circuit so initial-claim is the
only path that hits the gate.
2. Critical: gate had no direct unit test (the HTTP-layer test mocked
the choreographer). Added tests/unit/gateway/test_i_will_plan_sub_tasks_gate.py
with six tests: empty sub_tasks → incomplete_input, missing rich_plan
→ incomplete_input, filled sub_tasks → gate passes, developer with
empty sub_tasks → gate passes (devs don't decompose), sub_tasks filled
but approach empty → incomplete_input, in_progress re-entry short-
circuits before gate even with no sub_tasks.
3. Important: approach was only enforced at the HTTP Pydantic boundary.
Direct service-layer callers (MCP, test fixtures, orchestrator-
internal) could persist a plan with no approach. Gate now also checks
approach >= 20 chars (_PM_APPROACH_MIN_LEN constant) and includes it
in the rejection's missing list when absent.
Plus: stale docstring on i_will_plan corrected; _handle_pm_reentry
docstring rewritten to lead with the domain reason (re-entry contracts)
not the PLR0911 linter justification; unused pytest import removed from
test_i_will_plan_rich_required.py; pre-existing PLR2004/PLC0415 issues
in that file fixed.
i_will_plan now requires approach (min_length=20) at the schema and
non-empty sub_tasks at the gateway when the caller is a PM role. Restores
pre-gateway parity for _validate_claimed_start — agents could not
transition claimed -> in_progress without filling the rich plan.
Smoke run 3 (2026-05-11) showed PMs calling i_will_plan with just
plan='paragraph' and the gateway accepting it; Plan tab stayed empty
because no agent filled approach/sub_tasks/risks/open_questions.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section A1.
Wave 1 wired notify_list/get/ack into ContentActions but pointed them at
`self.notifications` (which is NotificationService — sender side, with
send_blocker_notification / send_qa_ready_notification / etc.). The
read methods (list_for_agent, get_for_recipient_and_mark_read,
acknowledge) live on `NotificationDeliveryService` instead.
Smoke run 2026-05-11 surfaced this immediately:
AttributeError: 'NotificationService' object has no attribute 'list_for_agent'
Fixes:
- roboco/api/deps.py — import NotificationDeliveryService and wire it
in as a new ContentActionsDeps field `notification_delivery`.
- roboco/services/gateway/content_actions.py — add notification_delivery
to ContentActionsDeps (Optional with `None` default for back-compat
with any tests that don't supply it). Point notify_list, notify_get,
notify_ack at self._deps.notification_delivery.
- tests/unit/gateway/test_content_actions.py — _make_deps adds a default
AsyncMock for notification_delivery so existing tests continue to pass.
Quality: ruff + mypy clean. 505 unit tests pass.
Pre-gateway parity (G8 part b of the 2026-05-11 design). The pre-gateway
TaskBlockInput at 254cc93:roboco/mcp/schemas/__init__.py required
blocker_type (external|internal|question|dependency) and what_needed
so PMs could triage their inbox by class. Current i_am_blocked dropped
both fields — every blocked task looked the same to the PM.
Now i_am_blocked accepts both as optional kwargs:
- Back-compat: callers that omit them still work (blocker_type defaults
to None → rendered as flat reason in the struggle entry).
- New: when supplied, the struggle journal entry body is structured
markdown (## Blocker Type / ## What Needed sections) so the panel's
journal view renders named blocks instead of one flat sentence.
Validator on blocker_type enforces the enum at the Pydantic boundary
with a clear "must be one of: ..." error if the agent invents a value
(same pattern as the Wave 3 G7 validators).
G8 part a — typed `pause(checkpoint_summary, remaining_work)` — defers.
That gap needs a new IntentSpec in foundation/policy/lifecycle.py
(currently pause is an ActionSpec only; agents auto-pause via i_am_idle)
plus checkpoint wiring through TaskService.add_checkpoint. Material
work, deferred until after the user has deployed and verified G7 +
G8b lands cleanly.
Wired:
- roboco/api/schemas/v2/flow.py — IAmBlockedRequest gains optional
blocker_type + what_needed; @field_validator enforces the enum
- roboco/api/routes/v2/flow_dev.py — passes the new fields through
- roboco/services/gateway/choreographer/_impl.py — i_am_blocked signature
+ structured struggle-entry rendering
- roboco/mcp/flow_server.py — typed wrapper with the kwargs
- agents/prompts/roles/developer.md — updated verb table
- tests/unit/mcp_servers/test_flow_server.py — updated to expect the
new optional kwargs as None when omitted
Quality: ruff + mypy clean. 505 tests pass.
Pre-gateway parity for G7 of the 2026-05-11 design. The pre-gateway
TaskCreateInput at 254cc93:roboco/mcp/schemas/__init__.py:210-235 had
@field_validator hooks that caught the most common LLM-vs-schema
confusions with helpful "did you mean X?" hints. Those validators were
lost in the gateway refactor.
Three validators added to DelegateRequest:
- estimated_complexity: rejects ints (some agents send 1/2/3 thinking
it's a priority), enforces enum {low|medium|high|critical}. Hint
steers them to drop priority (which isn't a delegate parameter).
- nature: rejects invented values like the 2026-05-11 'standard'
regression. Enum is {technical|non_technical}. Hint explicitly cites
the regression so the LLM knows why this is enforced.
- task_type: rejects invented task_type values. Enum is {code,
documentation, research, planning, design, administrative}.
Fail-fast at the Pydantic boundary returns a 422 with the structured
hint inline, so the agent loops a single retry instead of leaking a
TaskCompletenessError up the stack.
Existing tests in tests/unit/api/routes/v2/test_flow_*.py used
"nature": "feature" — a value that the gateway's TaskNature enum
never accepted, so it would have been rejected at completeness check
anyway. Updated both to "technical".
Spec ref: docs/superpowers/specs/2026-05-11-pre-gateway-parity-design.md