The system-prompt directive layer and the briefing block both opened
with "FIRST ACTION REQUIRED: run ToolSearch to activate deferred
Edit/Write". That premise is false: per Claude Code 2.1.114, ToolSearch
gates only deferred MCP tools, never built-ins — and it is not even a
callable tool in the agent runtime. Built-ins are loaded at spawn via
the `--tools` flag and gated solely by the per-role permission rules
(the actual Edit/Write breakage was the global Write(*)/Edit(*) deny +
single-slash path, fixed in c0ba335). So weak models dutifully chased a
nonexistent ToolSearch, concluded Edit/Write were unavailable, and
rewrote whole files via destructive shell redirection.
Both touch points now affirm the role's built-in tools are loaded and
ready, tell the agent NOT to call ToolSearch, and (for authoring roles)
explicitly steer away from whole-file shell redirection — directly
countering the clobber behaviour. Role prompt files (developer,
cell_pm, main_pm, board) updated to match. Dead
_read_tool_load_from_role_prompt (no callers) removed. Directive tests
rewritten to lock the corrected behaviour.
Smoke-10..14: every agent (developers included) got "Edit exists but is
not enabled in this context" and fell back to destructive bash
redirection (a 207-line README rewritten to a 3-line stub, which QA
correctly failed). Two coordinated defects in _generate_agent_settings /
_get_role_permissions:
1. base_deny carried a GLOBAL Write(*)/Edit(*). Claude Code evaluates
permission rules deny -> ask -> allow, first match wins — a deny
ALWAYS beats a more-specific allow and the glob syntax has no
negation. So the global deny unconditionally shadowed every per-role
workspace-scoped Write/Edit allow. Removed it; the security denies
that legitimately rely on deny-always-wins (Bash(git:*), credential
Read denies, curl github, env) stay. Roles that must not author
(qa, cell_pm, main_pm, auditor) keep their OWN Write(*)/Edit(*) deny.
2. The workspace allow used a single leading slash (Write(/data/...)).
Claude Code resolves a single / against the settings.json project
root, not the container filesystem root, so the allow silently never
matched even without defect #1. Emit the // absolute-filesystem form.
defaultMode stays bypassPermissions (switching to dontAsk would require
re-deriving the full allow-list and risks wedging agents elsewhere —
out of scope). Verified against Claude Code 2.1.114 permission docs.
Smoke-14: QA's claim_review evidence had pr_diff_summary="" and
files_changed=[] on a PR with a real README change. Root cause: diff()
and list_changed_files() diffed against the bare local <branch_name>.
That ref only exists in the clone where the dev ran `git checkout -b`
at claim. QA / documenter / PM inspect from their OWN clones, where a
bare <branch_name> resolves refs/heads then refs/remotes/<name> but
NEVER refs/remotes/origin/<name> — so `git diff base...<branch>` had an
unresolvable HEAD and silently returned an empty diff (run with
check=False).
#161 previously fixed the BASE side (cell-PM parent never pushed → fall
back to default branch). This is the symmetric HEAD-side facet.
open_pr pushes the leaf branch, so origin/<branch> is the
workspace-independent source of truth. New _resolve_head_ref fetches the
branch and prefers the local branch (dev's own clone, unchanged
behaviour), falling back to origin/<branch> (QA/doc/PM clones), then the
bare name so the command stays well-formed. diff() and
list_changed_files() route through it; explicit base (incremental dev
path, base=HEAD~1) is preserved.
The git network/auth deny rule matched its regex against the whole
command string, so heredoc bodies and echo/printf arguments that merely
documented git verbs (a README, a notes file) were treated as git
invocations and denied. This wedged smoke-13's dev: after wiping the
README via an Edit/Write fallback it could not restore it because every
`cat > README.md << EOF ... git commit ... EOF` was blocked.
The git-ops check now runs against a skeleton of the command with
heredoc bodies and echo/printf literal args stripped (both are data the
shell writes, never executed). Quoted args to a shell interpreter
(`bash -c "... && git fetch"`) ARE executed, are not echo/printf/heredoc
bodies, and so survive untouched — the hook's core purpose is preserved.
A sentinel prefix distinguishes a legitimately-empty skeleton from a
python failure (fail closed on failure). All other rules, including the
#164 import-bypass rule, still inspect the full command.
Smoke-12: be-dev-1 (minimax-m2.7) bypassed the entire MCP boundary by
running `uv run python3 -c "import os;
os.environ['ROBOCO_AGENT_ID']='...'; from roboco.mcp.flow_server
import open_pr; open_pr(...)"` from the Bash tool. This voided the
per-role tool manifest (role-scoping is meaningless if the agent can
import any server module in-process), forged agent identity via an
env-var rewrite, and ran choreographer code outside the gateway's
tracing + auth.
bash-guard-hook.sh now adds two deny rules:
1. Any python/uv/poetry/pipenv/pdm/hatch invocation that imports or
`-m`-runs roboco.* internals (mcp/services/runtime/foundation/
api/enforcement). The whole command string — heredoc body
included — is matched, so quoting/heredoc forms are covered.
2. Any assignment or export of ROBOCO_AGENT_ID (identity forgery).
Reading roboco source for context (cat/grep) is still allowed — the
block is on *executing* internals, not viewing them. Normal python
one-liners without roboco imports still pass.
19 bash-guard tests pass (10 prior + 9 new). Note: takes effect on
agent-image rebuild (hook ships in the agent container).
Smoke-12: be-pm delegated TWO subtasks under one cell parent — a
code subtask (be-dev-1, spawned) and a documentation subtask
(be-dev-2). The orchestrator dev-dispatch refuses to spawn a developer
for task_type=documentation, so the doc subtask became a permanent
orphan that loops dev-dispatch forever and would deadlock submit_up
(all subtasks must be terminal). The spine-cap is per-type so
code + documentation both passed sibling-dedup — the PM never saw the
anti-pattern warning.
_delegate_static_guards now rejects task_type='documentation' with a
remediate explaining the lifecycle auto-creates the documentation
phase (awaiting_documentation → documenter spawned) after the code
subtask passes QA, and that the PM should delegate ONLY the code
subtask.
#160 — panel /roboco-logo.png "received null":
next/image optimizer fails for static public assets in Next.js
standalone mode. Added `unoptimized` to the sidebar logo Image so
it serves the static file directly (validated on panel rebuild).
#161 — QA/doc evidence pr_diff_summary empty:
A leaf dev branch's parent_branch_for is the cell-PM branch, which
is never pushed (only devs push their leaf branch). diff against a
non-existent origin/<parent> returned empty. Added
GitService._resolve_diff_base + _default_branch_ref + _ref_exists:
diff/list_changed_files fall back to the repo default branch
(origin/HEAD → master/main) when origin/<parent> is absent.
#162 — claim_doc_task BRANCH_MISMATCH loop:
The documenter's clone is separate from the dev's; the task branch
already existed (dev created it) so no checkout ran in the doc
workspace — roboco_docs_write / commit failed BRANCH_MISMATCH and
the doc looped. Fixes:
(a) new GitService.checkout_branch_in_agent_workspace; claim_doc_task
checks out the task branch into the doc clone (best-effort —
a checkout hiccup never fails the claim).
(b) BRANCH_MISMATCH remediate now lists all four role claim verbs
(i_will_work_on / i_will_plan / claim_doc_task / claim_review).
(d) give_me_work next-hint is role+status aware via _claim_verb_hint
(doc→claim_doc_task, qa→claim_review, pm→i_will_plan, else dev).
Facet (c) (i_am_blocked "Not Found" for doc) was only reachable via
the stuck-without-checkout path; primary fix removes it.
Smoke-11 reached dev→QA→doc (deepest ever) and validated the prior
6 fixes (panel flood gone, #158/#159/#157 confirmed). These three
clear the doc-phase blockers found in that run.
Task #157 — spine-cap allows planning fanout across cells:
main-pm's pattern is to delegate planning to be-pm / fe-pm / ux-pm
in parallel — each on a different team. The previous spine-cap
rejected all planning siblings under one parent as
over-decomposition. New helper _is_cross_team_planning skips the
cap for planning when both teams are non-empty and distinct. Code
/ documentation stay capped regardless (single repo on one branch
shouldn't have two simultaneous code subtasks).
Task #159 — tracing-gap remediate hints every requirement:
journal:during_work>=1, journal:struggle, commits>=1, pr_open,
and self_verified had no entries in _hint_for_missing_key, so when
they were missing the agent saw the token in `missing[]` but the
`remediate` text had no instruction for how to satisfy them.
Smoke-10's be-dev-1 burned multiple turns retrying i_am_done not
knowing scope='reflect' doesn't count toward during_work. Now
every token has a hint, the during_work hint warns that reflect
doesn't satisfy it, and multi-hint remediate uses a numbered list
so the model treats each requirement as a distinct step instead
of a semicolon-blob.
Also coerce convert_plan._coerce_risk formatting (ruff-format follow-up
to 9cd73d0).
Bug:
Smoke-10's main-pm submitted a rich plan with risks omitting
severity. _normalize_risk persisted severity=None into the DB.
Every panel poll of /tasks/{id} then 500'd because
TaskPlanResponse.risks declares list[dict[str, str]] and Pydantic
rejects None for a str field. Result: panel single-task page broken
end-to-end every ~1-3s as panel reloads.
Fix:
Write side (_normalize_risk): default missing/None severity to
"medium" so new writes never persist None.
Read side (convert_plan._coerce_risk): defensively coerce any
existing DB row with severity=None to "medium" so old bad data
doesn't continue bricking the read path.
Task #156 (sessions): pre-gateway flow created a session for the whole
task tree at once, so subtasks were visible in the PM's group chat the
moment they existed. The gateway creates subtasks one-by-one via
delegate(), losing that wiring. Added
MessagingService.propagate_sessions_to_subtask and threaded it through
the choreographer's _create_subtask_from_inputs. ChoreographerDeps grew
an optional `messaging` field so existing test wirings keep working.
Task #155 (progress): smoke-9 had zero progress entries because the dev
never called progress() explicitly. Added _record_milestone_progress and
fire it server-side from two natural milestones — open_pr ("opened PR
#N", 70%) and i_am_done ("submitted for QA review", 90%). Best-effort
write (contextlib.suppress) so a progress failure cannot break the verb
path. Extracted _open_pr_success_envelope to keep cyclomatic rank ≤ B.
Bug:
ContentActions.evidence() hard-coded files_changed=[] and diffed
against HEAD~1 instead of the branch's parent. QA's _build_qa_claim_
evidence (and doc/_impl mirrors) sourced files_changed from
work_session.files_modified, which the gateway commit() never
populates (no add_files_modified plumbing). Result: QA / docs / PM
reviewers saw an empty change list on real PRs and only the latest
commit's delta — flagged in smoke-9 when PR #20 showed the README
change on GitHub but evidence() reported empty.
Fix:
Added GitService.list_changed_files (git diff --name-only against
parent branch). evidence(), _build_qa_claim_evidence,
_claim_doc_evidence, and _build_i_am_done_ok all source files_changed
from this — git is the authoritative source. evidence() also drops
the HEAD~1 base so the diff is the full PR.
Wired EvidenceRepo into ContentActionsDeps so evidence() returns
journal_highlights too, matching the QA/doc shape.
Bug:
Choreographer.i_will_plan() built spec_ctx with the raw plan string
while ctx (_ClaimPlanStartContext) got the resolved (panel-shaped)
dict. The verb runner uses spec_ctx — so the rich shape never reached
TaskService.set_plan. The panel's Plan tab stayed empty even when
PMs supplied approach / sub_tasks / risks.
Fix:
Pass the resolved (possibly-dict) plan into spec_ctx too. Widened
lifecycle.Context.plan to `str | dict[str, Any] | None` to match.
Tightened _resolve_effective_plan to require a non-empty narrative
paragraph — rich structure layers on top of prose, not in place of it.
Smoke-8 follow-up. Two issues in _write_agent_briefing:
1. _build_tool_load_block was scraping role prompts for a "## Load on
spawn" section that doesn't exist in any role file. Returned "" for
every role → no ToolSearch directive in the briefing. Combined with
weak models skipping the system-prompt-layer directive (#144), the
agent's first action was Edit → "not enabled in this context."
Fix: per-role tool list lives in the orchestrator (mirrors
factories._base.py). Pre-renders the directive directly. developer
and documenter get Edit + Write; QA/PMs/board get the common
read-only set. 7 tests pin the contract.
2. The briefing's "Terminal tools (how to exit cleanly)" section still
listed pre-gateway verb names: roboco_agent_idle,
roboco_task_substitute, roboco_task_escalate,
roboco_task_submit_qa, _qa_pass/fail, _docs_complete, _complete.
Same rename pattern as #145's _TERMINAL_TOOLS set. Updated to:
i_am_idle, i_am_blocked, unclaim, i_am_done, pass, fail,
i_documented, complete, submit_up, escalate_up, escalate_to_ceo.
The agent now reads the same directive in two places (system prompt +
session briefing) — the second touch point catches weak models that
skip the first.
Smoke-8: the stop-hook still nagged after a successful i_am_idle even
after #145's _TERMINAL_TOOLS rename. Root cause was upstream — nothing
was POSTing to /terminal/tool_recorded, so the SDK's recent_tools
deque stayed empty and had_terminal_recently() always returned False.
The PostToolUse hooks already record every tool call to
/budget/tool_called for the budget/loop tracker. Added a parallel call
to /terminal/tool_recorded so the terminal-tracker sees the same
stream. Fire-and-forget; never blocks Claude.
After the SDK suffix-strip (line ~798 in agent_sdk/server.py),
mcp__roboco-flow__i_am_idle becomes i_am_idle which is in
_TERMINAL_TOOLS (per #145). Stop-hook reads /terminal/stop_attempt
and now sees had_terminal_recently=true on the first attempt → exits 0.
Smoke-8: QA fell back to Bash for git inspection (git log, git branch,
git show) — bash-guard correctly blocked most of it. The
roboco-git-readonly MCP server WAS registered for every agent (per
orchestrator.py:1897) with roboco_git_status/log/diff/branches, but
no role prompt mentioned them, so agents never tried them.
Added the four verbs to the verb tables in: developer.md, documenter.md,
qa.md, cell_pm.md, main_pm.md, board.md. Each entry notes "use these,
NOT raw `Bash git ...`" so agents reach for the right tool first.
No code change — these MCP tools have existed all along. This is a
prompt visibility fix.
Smoke-8: QA correctly failed a PR because the PM wrote acceptance
criteria the gateway can never satisfy:
- "branch named feature/backend/<full-uuid>" — gateway generates
hierarchical 8-char IDs with `--` separator
- "commit prefix [<root-id>]" — gateway prefixes with the leaf
(dev's) task ID, not the root
The dev did the right work (timestamp added to README, PR opened) but
the literal criteria were unreachable.
Both cell_pm.md and main_pm.md now have an "How to write
acceptance_criteria" block explaining:
- Gateway-controlled outputs: branch name, commit prefix
- Examples of bad criteria (implementation/identifier-based) and good
criteria (outcome/file-content/PR-state)
- When you must reference a task ID, use the leaf (dev) ID — not
the root.
Next smoke run: PMs should write outcome criteria, dev work should
clear QA on first review (assuming the work itself is correct).
Smoke-8 surfaced a tight respawn loop: QA failed a PR cleanly, container
exited 0, then _check_health bumped error_count and respawned QA with
the same task_id. But by then the task was in needs_revision (dev's
state), so QA's claim_review was rejected — and the cycle repeated on
the next health tick. Token-burning loop.
Two layers:
1. _check_health now reads docker's exit code. exit_code == 0 →
graceful (intentional handoff via i_am_idle / clean shutdown) →
reset error_count, do NOT auto-restart. Non-zero → keep the
existing crash-retry behavior. Refactored into
_inspect_container_state + _handle_stopped_container to keep
xenon's complexity check happy.
2. _readiness_check_role_for_status now includes the dev-owned
states (needs_revision, verifying) so a misrouted spawn for QA /
PM / board on these statuses fails the readiness gate before the
gateway has to reject it. Defense in depth — the right path is
#1 (don't respawn on clean exit at all), but if some other code
path tries to spawn QA on needs_revision the gate now catches it.
Tests: 12 new (5 for _check_health graceful/crash matrix + 7 for the
expanded role-status table). Pre-gateway names (none of which were
needed here) untouched.
Smoke-7: be-pm got the expected spine-cap rejection on a second
delegate. The remediate said "drive the existing sibling to
completion / cancel it, OR split this parent into two sibling
parents". The model read that, decided neither applied, and
"adapted" by re-delegating with task_type='documentation' as a
"verification subtask". The gateway accepted it (different type =
no cap collision) but the orphan subtask had no claimant — it
blocked submit_up forever with "subtasks not all terminal".
Two-layer fix:
1. _spine_type_dup_envelope remediate now explicitly forbids the
workaround: "DO NOT work around this by delegating again with a
different task_type (e.g. 'documentation' or 'research' as a
'verification' subtask). The lifecycle handles QA, documentation,
and PM-review automatically after the code subtask finishes —
you do not create auxiliary subtasks for those roles. Call
i_am_idle() now and wait for the existing child to come back."
2. cell_pm.md workflow step 6 strengthened to name the anti-pattern
explicitly: no verification subtask (QA is the verification step);
never re-delegate with a different task_type as a workaround.
3 new tests pin the remediate text: forbids workaround, names the
verification anti-pattern, retains invalid_state error kind.
Smoke-7 evidence: every agent's first successful i_am_idle was followed
by a stop-hook error "you stopped without calling a terminal tool"
even though i_am_idle had just succeeded.
Root cause: agent_sdk._TERMINAL_TOOLS still held pre-gateway names
(roboco_agent_idle, roboco_task_submit_qa, ...). /terminal/tool_recorded
strips the `mcp__roboco-flow__` prefix and stores 'i_am_idle' — the
membership check against {'roboco_agent_idle', ...} never matched, so
had_terminal_recently() always returned False, and the stop hook nagged
every clean exit. Wasted ~2-3 turns per agent + burned stop_allowance.
Fix: rebuilt _TERMINAL_TOOLS with the current gateway verb names:
i_am_idle, i_am_done, i_am_blocked, i_documented, unclaim, pass, fail,
complete, submit_up, escalate_up, escalate_to_ceo.
7 new tests pin:
- every role's terminal verbs are recognized
- pre-gateway names are NOT in the set
- _SessionState.had_terminal_recently() returns True after i_am_idle
Smoke-7: be-dev-1 hit "Edit exists but is not enabled in this context."
Claude Code v2.1.69+ defers built-in tools (Edit, Write, Read, etc.)
behind a ToolSearch call. Weak models (minimax-m2.7) skip soft
directives buried in the briefing.
Also: 4 role prompts (developer, cell_pm, main_pm, board) claimed
"no ToolSearch needed" — a lie that compounds the problem. The
manifest registers MCP tools; built-in tools are still deferred.
Fix: compose_prompt now prepends a tool-load directive layer as the
FIRST block in the system prompt. It names the exact ToolSearch call
the role needs:
- developer/documenter: Read, Bash, Grep, Glob, Task, TodoWrite, Edit, Write
- qa/pm/board: Read, Bash, Grep, Glob, Task, TodoWrite (no Edit/Write)
The directive includes the failure mode it prevents so the model
understands what skipping the call causes.
Updated role-prompt lines that lied about ToolSearch.
7 new tests pin: directive is the first block; developer/documenter
get Edit/Write; qa/pm don't; failure-mode message is present.
Smoke-7 surfaced: be-qa called dm(recipient='qa-all', ...) — 'qa-all'
is a channel slug, not an agent. A2A enforcement raised
A2AAccessDeniedError. It propagated past dm(), past content_actions,
got caught by FastAPI middleware which renders RobocoError.to_dict()
as {'error': {'code': ..., 'message': ..., 'details': ...}} — a
DICT-shaped 'error' field.
do_server's circuit-breaker check (and flow_server's mirror) did
`payload.get('error') in _CIRCUIT_REJECTION_KINDS` — trying to hash
a dict against a frozenset → `TypeError: unhashable type: 'dict'`.
The agent saw "Error executing tool dm: unhashable type: 'dict'"
and got stuck calling dm in a loop.
Two-layer fix:
1. content_actions.dm now catches A2AAccessDeniedError and returns
Envelope.not_authorized with the original reason + route_hint as
remediate. This is the right shape — content tools always emit
Envelopes; RobocoErrors escaping to the middleware is a bug.
2. Defense-in-depth: do_server._record_and_check_circuit and
flow_server._record_and_check_circuit now guard against non-string
error fields. Any future RobocoError-leak that bypasses (1) will
pass through untouched instead of crashing the tool call.
3 new tests pin the contracts:
- dm A2A denial returns Envelope.not_authorized (not propagated)
- do_server circuit-breaker doesn't crash on dict-shaped errors
Smoke-7 surfaced this: QA spawned, claim_review succeeded, but every
attempt to call `pass()` fell through to dm/say workarounds. The MCP
tool 'pass' never existed.
Root cause: foundation.policy.lifecycle declares the intent verbs as
`pass_review`/`fail_review` (Python-friendly names — `pass`/`fail` are
keywords). intents_for_role(Role.QA) returns those names, the spawn
manifest carries them, and flow_server reads them. But flow_server's
_TOOLS dict has keys 'pass'/'fail' — the manifest's pass_review keys
didn't match and got silently dropped from the registration.
Fix: add _INTENT_TO_PUBLIC = {'pass_review': 'pass', 'fail_review':
'fail'} in flow_server. _register_tools transforms manifest names
through it before _TOOLS lookup. Manifest entries map to the public
MCP tool names the prompts advertise.
Also fixed _VERB_RETRY_LIMITS keys in foundation.agent_loop — they
used the IntentSpec names too, but the SDK receives the public name
from /verb/attempted (derived from the flow URL path), so the limit
entries never matched real rejections. Renamed to 'pass'/'fail'.
3 regression tests pin: pass/fail register under public names;
IntentSpec names don't leak through; the registered tool POSTs to
the correct orchestrator path.
Smoke-6 found the agent calling note(scope='decision', context=null,
chosen=null, rationale=null) eight times in a row. Root cause split
across two surfaces:
1. The MCP tool schema declared these fields as `str | None = None`,
producing a JSON schema of `anyOf [string, null]`. minimax-m2.7 read
that and decided null was a valid value — passed it on every retry.
2. The remediate text used `<placeholder>` syntax for the example
without telling the agent "don't pass null" explicitly.
Fixes:
- roboco/api/schemas/v2/do.py NoteRequest: context, chosen, rationale,
what_done, what_learned, what_struggled now typed `str = ""` (no None).
Pydantic on the route rejects literal null with 422 BEFORE the gateway
sees it. Empty string still counts as missing at the gate.
- roboco/mcp/do_server.py note(): matching signature changes so the
MCP tool schema declares the fields as `string` not `anyOf[string,null]`.
- roboco/services/gateway/content_actions.py: remediate text now opens
with "DO NOT pass null" and the example uses concrete values (redis vs
postgres) instead of <angle bracket> placeholders. Reflect remediate
also gets the don't-pass-null intro and concrete values.
8 new tests pin "schema rejects null for each of the 6 string fields"
plus "empty defaults work for unscoped notes".
Smoke-6 surfaced the gap. main-pm called note(scope='decision') with
context: null 8 times in a row — every one returned incomplete_input
and the agent kept retrying. The flow_server had a breaker (C1) but
do_server didn't, so content-tool rejections went uncapped.
Mirror the flow_server pattern:
- _CIRCUIT_REJECTION_KINDS = {tracing_gap, invalid_state,
not_authorized, incomplete_input} (same set)
- _record_and_check_circuit posts to the SDK's /verb/attempted on each
counted rejection
- When the SDK reports open=true, the original rejection envelope is
REPLACED with the circuit_open envelope so the agent stops retrying
The SDK side (agent_sdk/server.py) already accepts arbitrary verb
names; no changes there. note hits the default cap of 3 retries / 60s
from foundation.agent_loop. After the third incomplete_input the
agent gets circuit_open and the loop ends.
9 new tests pinning the contract.
xenon flagged ContentActions.pr_update as rank C — the three-branch
PM-or-assignee guard inlined with the precondition checks pushed it
over the cyclomatic-complexity bound. Extracted the authorization
check into a static helper _pr_update_is_authorized so the verb
itself stays at rank B and the helper carries the role-string +
team-equality branches.
Behavioural no-op; existing tests cover both the assignee path and
the cell_pm-same-team / cell_pm-other-team / main_pm paths.
Adds pr_update to _DEV_DO, _DOC_DO, _CELL_PM_DO, _MAIN_PM_DO so the
spawn manifest builder registers it on those roles' do-servers. QA,
auditor, and Board roles do not get it — QA reviews PRs but does not
edit them; Board operates above the PR layer; auditor is silent.
developer.md and documenter.md grow a verb-table row and a note next
to the open_pr workflow step calling out that pr_update — not bash-
shimmed `gh pr edit` — is the way to fix PR metadata.
ContentActions.pr_update enforces:
- task.pr_number is set (else invalid_state, remediate 'call open_pr')
- at least one of title/body/reviewers is non-None (else invalid_state)
- caller is task assignee OR PM on team (cell_pm same-team / main_pm
cross-team), else not_authorized
- GitError raised by the underlying service maps to invalid_state with
the upstream message preserved
The route at POST /api/v2/do/pr_update binds PRUpdateRequest, whose
model_validator returns 422 on all-None bodies so the verb layer
never sees them. The MCP tool registry adds 'pr_update' so manifest-
scoped do-servers can expose it to roles that opt in (next commit).
Smoke-5 surfaced that be-dev-1 had no gateway-native way to fix a
PR's title/body or request a reviewer after open_pr; `gh pr edit`
is bash-shimmed and the dev correctly escalated rather than bypass
the guard. This adds the GitService primitive: PATCH /pulls/{n}
for title/body and POST /pulls/{n}/requested_reviewers for the
reviewer list, with NotFound + 422 mapped to typed GitError. The
PRUpdateRequest schema enforces 'at least one field' via a
model_validator so the route returns 422 before reaching the verb.
Smoke-5 root cause. Agents wrote 5 decisions / 8 reflections / 1 struggle
during the run — every single entry persisted with task_id=NULL. The C8
tracing gate then never saw them and PMs spiraled forever on
'missing: journal:decision' while their decisions sat orphaned.
Cause: ContentActions.note/say/dm/notify called
TaskService.get_active_task_for_agent for task_id auto-injection. That
helper filters to _DEV_ACTIVE_STATUSES = {claimed, in_progress,
verifying, awaiting_qa, awaiting_documentation}. BLOCKED, PAUSED, and
NEEDS_REVISION fall outside that set — so the moment an agent gets
stuck (which is exactly when they journal), auto-injection returns None
and the entry persists without task_id.
Fix:
- New TaskService.get_journal_context_task_for_agent — same shape as
get_active_task_for_agent but the status set
_JOURNAL_CONTEXT_STATUSES adds BLOCKED, PAUSED, NEEDS_REVISION.
- ContentActions.note/say/dm/notify use the new lookup.
- ContentActions.commit keeps the narrow get_active_task_for_agent —
can't commit from blocked, so the dev-active set is correct there.
Tests:
- tests/unit/services/test_journal_context_lookup.py — 5 tests pinning
the two queries: journal-context INCLUDES blocked/paused/needs_revision,
dev-active EXCLUDES them.
- Existing content-actions tests updated to stub the new method
alongside the old one.
This alone may be 70% of what was killing smoke runs end-to-end.
Smoke run 4 crashed with AttributeError: 'SessionForTasksCreateRequest'
object has no attribute 'config' when main-pm called open_session.
The service _build_session_request reads req.config.max_message_count
(model has nested config). The gateway was passing the API schema
SessionForTasksCreateRequest (flat fields, no config attribute at all)
— so `req.config` blew up with AttributeError, not the safer None.
Fix: gateway now constructs SessionForTasksCreate (the model) with the
enum-typed relationship_type, matching what the route at
roboco/api/routes/sessions.py:230 does. Unknown relationship strings
fall back to DISCUSSION.
Two regression tests pin the contract: service receives the model;
invalid relationship_type defaults to DISCUSSION.
Pure whitespace / line-wrap reformats accumulated when ruff format
ran during earlier waves but weren't included in their commits. No
semantic changes — collection literals reflowed, with-statement
context managers regrouped via PEP 617 parens.
The inner closure inside BaseIndexPlugin.ingest() packed chunk
filtering, metadata merge, embedding, and aborted-transaction retry
into one block (CCN 13). Xenon's --max-absolute B ignores closures
(only top-level callables count), but radon flagged it as the last
rank-C block. Now zero rank-C blocks in the whole codebase.
Extracted:
- _filter_quality_chunks(raw_chunks) — module helper for the
tiny / mostly-markdown chunk filter
- _reset_store_connection(store) — module helper for piragi's
force-close + _init_schema after an aborted transaction
- _chunk_filter_embed_store(doc, metadata) — method that orchestrates
chunk → filter → embed → store (called via asyncio.to_thread)
- _store_with_transaction_retry(store, chunks, count) — method with
the retry loop, using early-return guard instead of nested ifs
ingest() now does: await asyncio.to_thread(self._chunk_filter_embed_store, ...).
No more nonlocal capture; chunk_count flows back through the return value.
Lift terminal-status and spine-task-type sets to class constants. Pull each
rule's rejection envelope into its own builder (_spine_type_dup_envelope,
_same_assignee_dup_envelope) and combine them under _sibling_dup_envelope.
The guard body is now a flat scan: skip terminal siblings, delegate the
per-sibling verdict to the helper, return the first non-None envelope.
Lift the rich-plan field list to a class constant (_RICH_PLAN_FIELDS) and
extract _resolve_effective_plan — the any()-over-5-keys decision between
the raw string plan and the panel-shaped dict. The verb body keeps its
spec-gate / re-entry / sub_tasks-gate sequence but no longer carries the
panel-shape branch.
Extract _build_struggle_body (the reason + optional Blocker Type / What Needed
markdown assembly) and _run_i_am_blocked_intent (the verb-runner dispatch +
try/except → rejection envelope). The verb body keeps its setup / spec-gate
shape but the structured-body branching and runner-exception branching no
longer count against it.
Extract four helpers: _extract_first_commit_sha (dict/model-tolerant sha read),
_already_addressed_criteria (set comprehension over existing status),
_find_existing_entry (preserved-entry lookup) and _new_criterion_entry (build
one fresh row). The main function is now a flat sequence: early-return on
empty criteria, early-return when all already addressed, then one loop with
two cases that each delegate to a helper.
Extract each conditional -v/-e block into a focused helper:
_append_claude_json_mount (claude.json file mount), _append_optional_host_mounts
(settings + briefing), _core_volume_and_env_args (the always-on block),
_append_provider_env (Anthropic-base/token), _append_manifest_args (spawn
manifest + gateway flag), _append_workspace_cwd (role-based -w). Two role
membership sets are lifted to class constants. The top-level function is
now a flat sequence of calls — no nested conditionals.
Extract the per-cycle XREADGROUP + dispatch into _listen_tick and the NOGROUP
self-heal branch into _handle_response_error. The outer loop is now a flat
while/try/except sequence: cancel breaks, response-error delegates the
recover-or-sleep decision to the helper, generic exceptions sleep. No
behavior change; the two new helpers preserve identical log messages and
ordering.
Lift the two scope-required field tables to module-level constants and route
through a shared _collect_required helper. The options-specific minimum-count
check and the generic scalar-empty check are each a one-line predicate
(_options_field_missing / _scalar_field_missing). The outer function is now a
dict lookup plus one call.
Lift the scope→sections lookup into a module-level dict (_SCOPE_SECTIONS) and
extract per-value rendering (options list / generic list / scalar) into
_render_section_value. The outer loop is now a flat dispatch with one early
continue per branch; the chained ternary and the list/dict branch ladder
that drove the CCN to 16 are gone.
TodoWrite is Anthropic's private session-local scratchpad — agents use
it to track their own immediate next steps. It does NOT surface to the
panel's Progress tab and is NOT a substitute for
progress(task_id, message, percentage). Smoke run 3 didn't show this
conflation yet, but Wave D's new progress() directive risks it.
- base.md gets the canonical "TodoWrite vs progress()" callout
- developer.md + documenter.md (the two roles with progress()) get
inline reminders in their verb tables: "NOT TodoWrite"
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section E4.
Smoke run 3 showed agents loading builtin Anthropic connectors
(mcp__claude_ai_Gmail__authenticate, Google Calendar, Notion, Drive)
alongside our 5 roboco MCP servers. The connectors bloat the tool
surface and give the LLM 'discover' targets it shouldn't have.
The Claude Code CLI's --strict-mcp-config flag tells it to load ONLY
the servers from --mcp-config, ignoring all builtin defaults. Added
to _append_image_and_claude_args next to --mcp-config.
Note: the existing --tools allowlist (Read,Write,Edit,Bash,Grep,Glob,
Task,TodoWrite) only filters builtin tools, not MCP-prefixed ones —
that's why the connectors slipped through.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section E3.
Pin the SQLAlchemy enum-naming invariant: every column typed Role
(aliased as AgentRole) MUST bind to the postgres enum named
'agentrole', not 'role'. Wave A5's smoke regression came from the
inferred 'role' type colliding with migration 001's 'agentrole'.
The fix already shipped (roboco/db/tables.py::_PG_ENUM_NAME_OVERRIDES);
this test locks it in. Two assertions: agents.role.udt_name == agentrole;
no stray 'role' enum type exists in pg_type.
QA / Documenter / Board don't have i_am_blocked in their manifests. The
D1 snippet's "escalate via i_am_blocked" line is now role-correct:
- QA / Documenter: unclaim(task_id) + dm(cell-pm, ...) with rejection
- Board: dm(ceo, ...) for PO/HoM; Auditor uses note(scope='reflect', ...)
All role prompts now mention channels() as the way to list valid
channel slugs. Smoke run 3 showed agents inventing slugs ('backend-dev',
'backend') and getting Channel not found. The channels() verb was
added in Wave 2 G6 but unused — making the directive explicit in
every prompt.
Cell PM prompt now mandates triage() as the first call on every
respawn, before re-decomposing. Smoke run 3 showed PMs re-decomposing
blindly and hitting spine-cap; triage shows them existing children
and prevents the over-decomposition pattern.