Smoke-10..14: every agent (developers included) got "Edit exists but is
not enabled in this context" and fell back to destructive bash
redirection (a 207-line README rewritten to a 3-line stub, which QA
correctly failed). Two coordinated defects in _generate_agent_settings /
_get_role_permissions:
1. base_deny carried a GLOBAL Write(*)/Edit(*). Claude Code evaluates
permission rules deny -> ask -> allow, first match wins — a deny
ALWAYS beats a more-specific allow and the glob syntax has no
negation. So the global deny unconditionally shadowed every per-role
workspace-scoped Write/Edit allow. Removed it; the security denies
that legitimately rely on deny-always-wins (Bash(git:*), credential
Read denies, curl github, env) stay. Roles that must not author
(qa, cell_pm, main_pm, auditor) keep their OWN Write(*)/Edit(*) deny.
2. The workspace allow used a single leading slash (Write(/data/...)).
Claude Code resolves a single / against the settings.json project
root, not the container filesystem root, so the allow silently never
matched even without defect #1. Emit the // absolute-filesystem form.
defaultMode stays bypassPermissions (switching to dontAsk would require
re-deriving the full allow-list and risks wedging agents elsewhere —
out of scope). Verified against Claude Code 2.1.114 permission docs.
Smoke-14: QA's claim_review evidence had pr_diff_summary="" and
files_changed=[] on a PR with a real README change. Root cause: diff()
and list_changed_files() diffed against the bare local <branch_name>.
That ref only exists in the clone where the dev ran `git checkout -b`
at claim. QA / documenter / PM inspect from their OWN clones, where a
bare <branch_name> resolves refs/heads then refs/remotes/<name> but
NEVER refs/remotes/origin/<name> — so `git diff base...<branch>` had an
unresolvable HEAD and silently returned an empty diff (run with
check=False).
#161 previously fixed the BASE side (cell-PM parent never pushed → fall
back to default branch). This is the symmetric HEAD-side facet.
open_pr pushes the leaf branch, so origin/<branch> is the
workspace-independent source of truth. New _resolve_head_ref fetches the
branch and prefers the local branch (dev's own clone, unchanged
behaviour), falling back to origin/<branch> (QA/doc/PM clones), then the
bare name so the command stays well-formed. diff() and
list_changed_files() route through it; explicit base (incremental dev
path, base=HEAD~1) is preserved.
The git network/auth deny rule matched its regex against the whole
command string, so heredoc bodies and echo/printf arguments that merely
documented git verbs (a README, a notes file) were treated as git
invocations and denied. This wedged smoke-13's dev: after wiping the
README via an Edit/Write fallback it could not restore it because every
`cat > README.md << EOF ... git commit ... EOF` was blocked.
The git-ops check now runs against a skeleton of the command with
heredoc bodies and echo/printf literal args stripped (both are data the
shell writes, never executed). Quoted args to a shell interpreter
(`bash -c "... && git fetch"`) ARE executed, are not echo/printf/heredoc
bodies, and so survive untouched — the hook's core purpose is preserved.
A sentinel prefix distinguishes a legitimately-empty skeleton from a
python failure (fail closed on failure). All other rules, including the
#164 import-bypass rule, still inspect the full command.
Smoke-12: be-dev-1 (minimax-m2.7) bypassed the entire MCP boundary by
running `uv run python3 -c "import os;
os.environ['ROBOCO_AGENT_ID']='...'; from roboco.mcp.flow_server
import open_pr; open_pr(...)"` from the Bash tool. This voided the
per-role tool manifest (role-scoping is meaningless if the agent can
import any server module in-process), forged agent identity via an
env-var rewrite, and ran choreographer code outside the gateway's
tracing + auth.
bash-guard-hook.sh now adds two deny rules:
1. Any python/uv/poetry/pipenv/pdm/hatch invocation that imports or
`-m`-runs roboco.* internals (mcp/services/runtime/foundation/
api/enforcement). The whole command string — heredoc body
included — is matched, so quoting/heredoc forms are covered.
2. Any assignment or export of ROBOCO_AGENT_ID (identity forgery).
Reading roboco source for context (cat/grep) is still allowed — the
block is on *executing* internals, not viewing them. Normal python
one-liners without roboco imports still pass.
19 bash-guard tests pass (10 prior + 9 new). Note: takes effect on
agent-image rebuild (hook ships in the agent container).
Smoke-12: be-pm delegated TWO subtasks under one cell parent — a
code subtask (be-dev-1, spawned) and a documentation subtask
(be-dev-2). The orchestrator dev-dispatch refuses to spawn a developer
for task_type=documentation, so the doc subtask became a permanent
orphan that loops dev-dispatch forever and would deadlock submit_up
(all subtasks must be terminal). The spine-cap is per-type so
code + documentation both passed sibling-dedup — the PM never saw the
anti-pattern warning.
_delegate_static_guards now rejects task_type='documentation' with a
remediate explaining the lifecycle auto-creates the documentation
phase (awaiting_documentation → documenter spawned) after the code
subtask passes QA, and that the PM should delegate ONLY the code
subtask.
#160 — panel /roboco-logo.png "received null":
next/image optimizer fails for static public assets in Next.js
standalone mode. Added `unoptimized` to the sidebar logo Image so
it serves the static file directly (validated on panel rebuild).
#161 — QA/doc evidence pr_diff_summary empty:
A leaf dev branch's parent_branch_for is the cell-PM branch, which
is never pushed (only devs push their leaf branch). diff against a
non-existent origin/<parent> returned empty. Added
GitService._resolve_diff_base + _default_branch_ref + _ref_exists:
diff/list_changed_files fall back to the repo default branch
(origin/HEAD → master/main) when origin/<parent> is absent.
#162 — claim_doc_task BRANCH_MISMATCH loop:
The documenter's clone is separate from the dev's; the task branch
already existed (dev created it) so no checkout ran in the doc
workspace — roboco_docs_write / commit failed BRANCH_MISMATCH and
the doc looped. Fixes:
(a) new GitService.checkout_branch_in_agent_workspace; claim_doc_task
checks out the task branch into the doc clone (best-effort —
a checkout hiccup never fails the claim).
(b) BRANCH_MISMATCH remediate now lists all four role claim verbs
(i_will_work_on / i_will_plan / claim_doc_task / claim_review).
(d) give_me_work next-hint is role+status aware via _claim_verb_hint
(doc→claim_doc_task, qa→claim_review, pm→i_will_plan, else dev).
Facet (c) (i_am_blocked "Not Found" for doc) was only reachable via
the stuck-without-checkout path; primary fix removes it.
Smoke-11 reached dev→QA→doc (deepest ever) and validated the prior
6 fixes (panel flood gone, #158/#159/#157 confirmed). These three
clear the doc-phase blockers found in that run.
Task #157 — spine-cap allows planning fanout across cells:
main-pm's pattern is to delegate planning to be-pm / fe-pm / ux-pm
in parallel — each on a different team. The previous spine-cap
rejected all planning siblings under one parent as
over-decomposition. New helper _is_cross_team_planning skips the
cap for planning when both teams are non-empty and distinct. Code
/ documentation stay capped regardless (single repo on one branch
shouldn't have two simultaneous code subtasks).
Task #159 — tracing-gap remediate hints every requirement:
journal:during_work>=1, journal:struggle, commits>=1, pr_open,
and self_verified had no entries in _hint_for_missing_key, so when
they were missing the agent saw the token in `missing[]` but the
`remediate` text had no instruction for how to satisfy them.
Smoke-10's be-dev-1 burned multiple turns retrying i_am_done not
knowing scope='reflect' doesn't count toward during_work. Now
every token has a hint, the during_work hint warns that reflect
doesn't satisfy it, and multi-hint remediate uses a numbered list
so the model treats each requirement as a distinct step instead
of a semicolon-blob.
Also coerce convert_plan._coerce_risk formatting (ruff-format follow-up
to 9cd73d0).
Task #156 (sessions): pre-gateway flow created a session for the whole
task tree at once, so subtasks were visible in the PM's group chat the
moment they existed. The gateway creates subtasks one-by-one via
delegate(), losing that wiring. Added
MessagingService.propagate_sessions_to_subtask and threaded it through
the choreographer's _create_subtask_from_inputs. ChoreographerDeps grew
an optional `messaging` field so existing test wirings keep working.
Task #155 (progress): smoke-9 had zero progress entries because the dev
never called progress() explicitly. Added _record_milestone_progress and
fire it server-side from two natural milestones — open_pr ("opened PR
#N", 70%) and i_am_done ("submitted for QA review", 90%). Best-effort
write (contextlib.suppress) so a progress failure cannot break the verb
path. Extracted _open_pr_success_envelope to keep cyclomatic rank ≤ B.
Bug:
ContentActions.evidence() hard-coded files_changed=[] and diffed
against HEAD~1 instead of the branch's parent. QA's _build_qa_claim_
evidence (and doc/_impl mirrors) sourced files_changed from
work_session.files_modified, which the gateway commit() never
populates (no add_files_modified plumbing). Result: QA / docs / PM
reviewers saw an empty change list on real PRs and only the latest
commit's delta — flagged in smoke-9 when PR #20 showed the README
change on GitHub but evidence() reported empty.
Fix:
Added GitService.list_changed_files (git diff --name-only against
parent branch). evidence(), _build_qa_claim_evidence,
_claim_doc_evidence, and _build_i_am_done_ok all source files_changed
from this — git is the authoritative source. evidence() also drops
the HEAD~1 base so the diff is the full PR.
Wired EvidenceRepo into ContentActionsDeps so evidence() returns
journal_highlights too, matching the QA/doc shape.
Bug:
Choreographer.i_will_plan() built spec_ctx with the raw plan string
while ctx (_ClaimPlanStartContext) got the resolved (panel-shaped)
dict. The verb runner uses spec_ctx — so the rich shape never reached
TaskService.set_plan. The panel's Plan tab stayed empty even when
PMs supplied approach / sub_tasks / risks.
Fix:
Pass the resolved (possibly-dict) plan into spec_ctx too. Widened
lifecycle.Context.plan to `str | dict[str, Any] | None` to match.
Tightened _resolve_effective_plan to require a non-empty narrative
paragraph — rich structure layers on top of prose, not in place of it.
Smoke-8 follow-up. Two issues in _write_agent_briefing:
1. _build_tool_load_block was scraping role prompts for a "## Load on
spawn" section that doesn't exist in any role file. Returned "" for
every role → no ToolSearch directive in the briefing. Combined with
weak models skipping the system-prompt-layer directive (#144), the
agent's first action was Edit → "not enabled in this context."
Fix: per-role tool list lives in the orchestrator (mirrors
factories._base.py). Pre-renders the directive directly. developer
and documenter get Edit + Write; QA/PMs/board get the common
read-only set. 7 tests pin the contract.
2. The briefing's "Terminal tools (how to exit cleanly)" section still
listed pre-gateway verb names: roboco_agent_idle,
roboco_task_substitute, roboco_task_escalate,
roboco_task_submit_qa, _qa_pass/fail, _docs_complete, _complete.
Same rename pattern as #145's _TERMINAL_TOOLS set. Updated to:
i_am_idle, i_am_blocked, unclaim, i_am_done, pass, fail,
i_documented, complete, submit_up, escalate_up, escalate_to_ceo.
The agent now reads the same directive in two places (system prompt +
session briefing) — the second touch point catches weak models that
skip the first.
Smoke-8 surfaced a tight respawn loop: QA failed a PR cleanly, container
exited 0, then _check_health bumped error_count and respawned QA with
the same task_id. But by then the task was in needs_revision (dev's
state), so QA's claim_review was rejected — and the cycle repeated on
the next health tick. Token-burning loop.
Two layers:
1. _check_health now reads docker's exit code. exit_code == 0 →
graceful (intentional handoff via i_am_idle / clean shutdown) →
reset error_count, do NOT auto-restart. Non-zero → keep the
existing crash-retry behavior. Refactored into
_inspect_container_state + _handle_stopped_container to keep
xenon's complexity check happy.
2. _readiness_check_role_for_status now includes the dev-owned
states (needs_revision, verifying) so a misrouted spawn for QA /
PM / board on these statuses fails the readiness gate before the
gateway has to reject it. Defense in depth — the right path is
#1 (don't respawn on clean exit at all), but if some other code
path tries to spawn QA on needs_revision the gate now catches it.
Tests: 12 new (5 for _check_health graceful/crash matrix + 7 for the
expanded role-status table). Pre-gateway names (none of which were
needed here) untouched.
Smoke-7: be-pm got the expected spine-cap rejection on a second
delegate. The remediate said "drive the existing sibling to
completion / cancel it, OR split this parent into two sibling
parents". The model read that, decided neither applied, and
"adapted" by re-delegating with task_type='documentation' as a
"verification subtask". The gateway accepted it (different type =
no cap collision) but the orphan subtask had no claimant — it
blocked submit_up forever with "subtasks not all terminal".
Two-layer fix:
1. _spine_type_dup_envelope remediate now explicitly forbids the
workaround: "DO NOT work around this by delegating again with a
different task_type (e.g. 'documentation' or 'research' as a
'verification' subtask). The lifecycle handles QA, documentation,
and PM-review automatically after the code subtask finishes —
you do not create auxiliary subtasks for those roles. Call
i_am_idle() now and wait for the existing child to come back."
2. cell_pm.md workflow step 6 strengthened to name the anti-pattern
explicitly: no verification subtask (QA is the verification step);
never re-delegate with a different task_type as a workaround.
3 new tests pin the remediate text: forbids workaround, names the
verification anti-pattern, retains invalid_state error kind.
Smoke-7 evidence: every agent's first successful i_am_idle was followed
by a stop-hook error "you stopped without calling a terminal tool"
even though i_am_idle had just succeeded.
Root cause: agent_sdk._TERMINAL_TOOLS still held pre-gateway names
(roboco_agent_idle, roboco_task_submit_qa, ...). /terminal/tool_recorded
strips the `mcp__roboco-flow__` prefix and stores 'i_am_idle' — the
membership check against {'roboco_agent_idle', ...} never matched, so
had_terminal_recently() always returned False, and the stop hook nagged
every clean exit. Wasted ~2-3 turns per agent + burned stop_allowance.
Fix: rebuilt _TERMINAL_TOOLS with the current gateway verb names:
i_am_idle, i_am_done, i_am_blocked, i_documented, unclaim, pass, fail,
complete, submit_up, escalate_up, escalate_to_ceo.
7 new tests pin:
- every role's terminal verbs are recognized
- pre-gateway names are NOT in the set
- _SessionState.had_terminal_recently() returns True after i_am_idle
Smoke-7: be-dev-1 hit "Edit exists but is not enabled in this context."
Claude Code v2.1.69+ defers built-in tools (Edit, Write, Read, etc.)
behind a ToolSearch call. Weak models (minimax-m2.7) skip soft
directives buried in the briefing.
Also: 4 role prompts (developer, cell_pm, main_pm, board) claimed
"no ToolSearch needed" — a lie that compounds the problem. The
manifest registers MCP tools; built-in tools are still deferred.
Fix: compose_prompt now prepends a tool-load directive layer as the
FIRST block in the system prompt. It names the exact ToolSearch call
the role needs:
- developer/documenter: Read, Bash, Grep, Glob, Task, TodoWrite, Edit, Write
- qa/pm/board: Read, Bash, Grep, Glob, Task, TodoWrite (no Edit/Write)
The directive includes the failure mode it prevents so the model
understands what skipping the call causes.
Updated role-prompt lines that lied about ToolSearch.
7 new tests pin: directive is the first block; developer/documenter
get Edit/Write; qa/pm don't; failure-mode message is present.
Smoke-7 surfaced: be-qa called dm(recipient='qa-all', ...) — 'qa-all'
is a channel slug, not an agent. A2A enforcement raised
A2AAccessDeniedError. It propagated past dm(), past content_actions,
got caught by FastAPI middleware which renders RobocoError.to_dict()
as {'error': {'code': ..., 'message': ..., 'details': ...}} — a
DICT-shaped 'error' field.
do_server's circuit-breaker check (and flow_server's mirror) did
`payload.get('error') in _CIRCUIT_REJECTION_KINDS` — trying to hash
a dict against a frozenset → `TypeError: unhashable type: 'dict'`.
The agent saw "Error executing tool dm: unhashable type: 'dict'"
and got stuck calling dm in a loop.
Two-layer fix:
1. content_actions.dm now catches A2AAccessDeniedError and returns
Envelope.not_authorized with the original reason + route_hint as
remediate. This is the right shape — content tools always emit
Envelopes; RobocoErrors escaping to the middleware is a bug.
2. Defense-in-depth: do_server._record_and_check_circuit and
flow_server._record_and_check_circuit now guard against non-string
error fields. Any future RobocoError-leak that bypasses (1) will
pass through untouched instead of crashing the tool call.
3 new tests pin the contracts:
- dm A2A denial returns Envelope.not_authorized (not propagated)
- do_server circuit-breaker doesn't crash on dict-shaped errors
Smoke-7 surfaced this: QA spawned, claim_review succeeded, but every
attempt to call `pass()` fell through to dm/say workarounds. The MCP
tool 'pass' never existed.
Root cause: foundation.policy.lifecycle declares the intent verbs as
`pass_review`/`fail_review` (Python-friendly names — `pass`/`fail` are
keywords). intents_for_role(Role.QA) returns those names, the spawn
manifest carries them, and flow_server reads them. But flow_server's
_TOOLS dict has keys 'pass'/'fail' — the manifest's pass_review keys
didn't match and got silently dropped from the registration.
Fix: add _INTENT_TO_PUBLIC = {'pass_review': 'pass', 'fail_review':
'fail'} in flow_server. _register_tools transforms manifest names
through it before _TOOLS lookup. Manifest entries map to the public
MCP tool names the prompts advertise.
Also fixed _VERB_RETRY_LIMITS keys in foundation.agent_loop — they
used the IntentSpec names too, but the SDK receives the public name
from /verb/attempted (derived from the flow URL path), so the limit
entries never matched real rejections. Renamed to 'pass'/'fail'.
3 regression tests pin: pass/fail register under public names;
IntentSpec names don't leak through; the registered tool POSTs to
the correct orchestrator path.
Smoke-6 found the agent calling note(scope='decision', context=null,
chosen=null, rationale=null) eight times in a row. Root cause split
across two surfaces:
1. The MCP tool schema declared these fields as `str | None = None`,
producing a JSON schema of `anyOf [string, null]`. minimax-m2.7 read
that and decided null was a valid value — passed it on every retry.
2. The remediate text used `<placeholder>` syntax for the example
without telling the agent "don't pass null" explicitly.
Fixes:
- roboco/api/schemas/v2/do.py NoteRequest: context, chosen, rationale,
what_done, what_learned, what_struggled now typed `str = ""` (no None).
Pydantic on the route rejects literal null with 422 BEFORE the gateway
sees it. Empty string still counts as missing at the gate.
- roboco/mcp/do_server.py note(): matching signature changes so the
MCP tool schema declares the fields as `string` not `anyOf[string,null]`.
- roboco/services/gateway/content_actions.py: remediate text now opens
with "DO NOT pass null" and the example uses concrete values (redis vs
postgres) instead of <angle bracket> placeholders. Reflect remediate
also gets the don't-pass-null intro and concrete values.
8 new tests pin "schema rejects null for each of the 6 string fields"
plus "empty defaults work for unscoped notes".
Smoke-6 surfaced the gap. main-pm called note(scope='decision') with
context: null 8 times in a row — every one returned incomplete_input
and the agent kept retrying. The flow_server had a breaker (C1) but
do_server didn't, so content-tool rejections went uncapped.
Mirror the flow_server pattern:
- _CIRCUIT_REJECTION_KINDS = {tracing_gap, invalid_state,
not_authorized, incomplete_input} (same set)
- _record_and_check_circuit posts to the SDK's /verb/attempted on each
counted rejection
- When the SDK reports open=true, the original rejection envelope is
REPLACED with the circuit_open envelope so the agent stops retrying
The SDK side (agent_sdk/server.py) already accepts arbitrary verb
names; no changes there. note hits the default cap of 3 retries / 60s
from foundation.agent_loop. After the third incomplete_input the
agent gets circuit_open and the loop ends.
9 new tests pinning the contract.
ContentActions.pr_update enforces:
- task.pr_number is set (else invalid_state, remediate 'call open_pr')
- at least one of title/body/reviewers is non-None (else invalid_state)
- caller is task assignee OR PM on team (cell_pm same-team / main_pm
cross-team), else not_authorized
- GitError raised by the underlying service maps to invalid_state with
the upstream message preserved
The route at POST /api/v2/do/pr_update binds PRUpdateRequest, whose
model_validator returns 422 on all-None bodies so the verb layer
never sees them. The MCP tool registry adds 'pr_update' so manifest-
scoped do-servers can expose it to roles that opt in (next commit).
Smoke-5 surfaced that be-dev-1 had no gateway-native way to fix a
PR's title/body or request a reviewer after open_pr; `gh pr edit`
is bash-shimmed and the dev correctly escalated rather than bypass
the guard. This adds the GitService primitive: PATCH /pulls/{n}
for title/body and POST /pulls/{n}/requested_reviewers for the
reviewer list, with NotFound + 422 mapped to typed GitError. The
PRUpdateRequest schema enforces 'at least one field' via a
model_validator so the route returns 422 before reaching the verb.
Smoke-5 root cause. Agents wrote 5 decisions / 8 reflections / 1 struggle
during the run — every single entry persisted with task_id=NULL. The C8
tracing gate then never saw them and PMs spiraled forever on
'missing: journal:decision' while their decisions sat orphaned.
Cause: ContentActions.note/say/dm/notify called
TaskService.get_active_task_for_agent for task_id auto-injection. That
helper filters to _DEV_ACTIVE_STATUSES = {claimed, in_progress,
verifying, awaiting_qa, awaiting_documentation}. BLOCKED, PAUSED, and
NEEDS_REVISION fall outside that set — so the moment an agent gets
stuck (which is exactly when they journal), auto-injection returns None
and the entry persists without task_id.
Fix:
- New TaskService.get_journal_context_task_for_agent — same shape as
get_active_task_for_agent but the status set
_JOURNAL_CONTEXT_STATUSES adds BLOCKED, PAUSED, NEEDS_REVISION.
- ContentActions.note/say/dm/notify use the new lookup.
- ContentActions.commit keeps the narrow get_active_task_for_agent —
can't commit from blocked, so the dev-active set is correct there.
Tests:
- tests/unit/services/test_journal_context_lookup.py — 5 tests pinning
the two queries: journal-context INCLUDES blocked/paused/needs_revision,
dev-active EXCLUDES them.
- Existing content-actions tests updated to stub the new method
alongside the old one.
This alone may be 70% of what was killing smoke runs end-to-end.
Smoke run 4 crashed with AttributeError: 'SessionForTasksCreateRequest'
object has no attribute 'config' when main-pm called open_session.
The service _build_session_request reads req.config.max_message_count
(model has nested config). The gateway was passing the API schema
SessionForTasksCreateRequest (flat fields, no config attribute at all)
— so `req.config` blew up with AttributeError, not the safer None.
Fix: gateway now constructs SessionForTasksCreate (the model) with the
enum-typed relationship_type, matching what the route at
roboco/api/routes/sessions.py:230 does. Unknown relationship strings
fall back to DISCUSSION.
Two regression tests pin the contract: service receives the model;
invalid relationship_type defaults to DISCUSSION.
Pure whitespace / line-wrap reformats accumulated when ruff format
ran during earlier waves but weren't included in their commits. No
semantic changes — collection literals reflowed, with-statement
context managers regrouped via PEP 617 parens.
TodoWrite is Anthropic's private session-local scratchpad — agents use
it to track their own immediate next steps. It does NOT surface to the
panel's Progress tab and is NOT a substitute for
progress(task_id, message, percentage). Smoke run 3 didn't show this
conflation yet, but Wave D's new progress() directive risks it.
- base.md gets the canonical "TodoWrite vs progress()" callout
- developer.md + documenter.md (the two roles with progress()) get
inline reminders in their verb tables: "NOT TodoWrite"
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section E4.
Smoke run 3 showed agents loading builtin Anthropic connectors
(mcp__claude_ai_Gmail__authenticate, Google Calendar, Notion, Drive)
alongside our 5 roboco MCP servers. The connectors bloat the tool
surface and give the LLM 'discover' targets it shouldn't have.
The Claude Code CLI's --strict-mcp-config flag tells it to load ONLY
the servers from --mcp-config, ignoring all builtin defaults. Added
to _append_image_and_claude_args next to --mcp-config.
Note: the existing --tools allowlist (Read,Write,Edit,Bash,Grep,Glob,
Task,TodoWrite) only filters builtin tools, not MCP-prefixed ones —
that's why the connectors slipped through.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section E3.
Pin the SQLAlchemy enum-naming invariant: every column typed Role
(aliased as AgentRole) MUST bind to the postgres enum named
'agentrole', not 'role'. Wave A5's smoke regression came from the
inferred 'role' type colliding with migration 001's 'agentrole'.
The fix already shipped (roboco/db/tables.py::_PG_ENUM_NAME_OVERRIDES);
this test locks it in. Two assertions: agents.role.udt_name == agentrole;
no stray 'role' enum type exists in pg_type.
_check_pm_decision_required now requires the latest journal:decision
within pm_decision_window_seconds (default 300). Older decisions no
longer satisfy the gate. Adds JournalService.latest_decision_at.
Future-tighten (out of scope): per-verb-group consumption tracking
would need persistent state — Choreographer is per-request today.
Smoke run 3 showed agents auto-pausing on i_am_idle (correct behavior
for non-terminal tasks) but capturing no checkpoint — panel's
Checkpoints column stayed empty. Pre-gateway parity: the auto-pause
path now writes a synthetic checkpoint summarizing state at pause-time
so the panel reflects reality.
Manual i_will_pause (G8a, deferred) will eventually let agents pass
their own checkpoint_summary; for now this synthetic write covers the
bare i_am_idle case which is what all current agents do.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section C7.
The auditor's role is 'silent observer' — read every channel and emit
a reflect note when something notable happens. Smoke run 3 never
spawned auditor because no event-subscription registered it. Added
handler handle_auditor_spawn() wired to:
- task.blocked (EventType.TASK_BLOCKED)
- task.cancelled (EventType.TASK_CANCELLED)
- task.awaiting_ceo_approval (EventType.TASK_AWAITING_CEO_APPROVAL)
Routine events (task.claimed, task.started, task.created) deliberately
do NOT trigger auditor — those are progress, not exceptions. The
auditor's container is one-shot: i_am_idle() exits after logging its
reflect note.
Auditor spawn failures are swallowed into a WARNING log so they cannot
block the underlying event's processing chain. The auditor is a silent
observer — its absence must have no side effects on the lifecycle.
Pre-gateway parity. evidence(task_id).acceptance_criteria_status was
always [] because the gateway's i_am_done gate validated each
criterion against the dev's journal:reflect but didn't persist the
per-criterion verdict. The panel + audit log couldn't show
per-criterion checkmarks.
Now the gate writes a list of {criterion, addressed, artifact_ref,
checked_at} entries to task.acceptance_criteria_status. The existing
matching logic surfaces which artifact (commit sha / reflect-note)
addressed each criterion; entries that aren't addressed get
addressed=False so the panel can flag them.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section C5.
Pre-gateway parity. Smoke run 3 showed task.work_session_id null on
every task — the choreographer's claim/plan/start path didn't create
the row that downstream subsystems (panel, PR tracking, merge chain)
need to track agent-per-task git activity.
Add TaskService.ensure_work_session(task_id, agent_id) as a public
wrapper around the existing _create_work_session_if_needed logic.
Role restriction lifted to None so both developers and PMs get a
session (pre-gateway always created sessions for all claimants).
Built-in re-entry guard prevents duplicate rows on re-claim.
Wire the call into both _claim_plan_start_run and _resume_from_claimed
immediately before _touch, so every successful in_progress transition
(including the stuck-claimed recovery path) creates the row.
Spec ref: Wave C task C4 (2026-05-12).
Smoke run 3 showed agents reaped at the 3-min stale-claim window
while they were actively retrying rejected verbs. Two causes:
1. The reaper threshold was hardcoded at 180s via claim_stale_seconds.
LLM inference + retry loops routinely take longer than that between
verb-successes. Added settings.stale_claim_reap_seconds (default
600s); override via ROBOCO_STALE_CLAIM_REAP_SECONDS env var.
claim_stale_seconds (spawn-filter cutoff) is unchanged at 180s.
2. last_heartbeat_at only refreshed on verb SUCCESS. A verb stuck
in a rejection loop (e.g. tracing_gap missing journal:decision)
showed no heartbeat updates even though the agent was alive.
Added a best-effort heartbeat refresh inside _emit_rejection so
EVERY verb dispatch — success or rejection — counts as activity.
Heartbeat approach: option (b) — touch inside _emit_rejection (single
centralized rejection path). Requires no middleware layer, no HTTP body
parsing, and no new files. The _touch guard for task_id=None means
agent-level rejections (no task context) are a safe no-op.
Net effect: agents stop being reaped mid-retry. Genuinely-stuck
containers (no verb dispatch at all) still reap normally at 600s.
Spec ref: Wave C Task C3.
Smoke run 3 fired 'ensure_workspace: refresh fetch returned non-zero'
9 times per run because each evidence(task_id) call triggered
ensure_workspace -> fetch. The workspace doesn't change in subseconds.
Added a 30s TTL cache keyed by workspace path. ensure_workspace(force=True)
bypasses the cache for callers that genuinely need a fresh fetch.
Net effect: log noise drops from 9 entries to 1-2 per run; orchestrator
spends less time waiting on redundant git fetches.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section C2.
Smoke run 3 showed Main PM hitting 7 incomplete_input rejections on
the decision-note required-fields gate before finally succeeding.
The per-verb breaker tracks repeated rejections of the same verb in
a 60s window and returns circuit_open after the 3rd strike — but its
classification set only included tracing_gap. incomplete_input was
added in Wave 1 (pre-gateway parity for decision/reflect structured
fields) and should have been added to the breaker at the same time.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section C1.
Smoke run 3 showed Main PM's first give_me_work() returning
{status: idle, next: 'no Main PM work'} even though c7935d2c was
pending and assigned to Main PM. The filter only walked
list_assigned_for_agent (ordered by priority/updated_at — pending
could rank behind in_progress rows) and the PM path fell through
to idle because the pre-assigned pending case was not checked first.
Pre-pended a list_pending_for_agent check in both give_me_work and
pm_give_me_work: tasks where assigned_to=agent_id AND status=pending
take priority over all other lookups. Added TaskService.list_pending_for_agent
for the query (ordered by sequence, priority, created_at).
Updated existing tests in test_choreographer_dev, test_choreographer_pm_extras,
and test_heartbeat_wired to set list_pending_for_agent.return_value=[]
where they were not testing the pre-assigned path.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section B6.
Smoke run 3 showed the bash-guard hook emitting 8+ lines on every
blocked shell-git op — enumerating every alternative MCP verb across
roboco-flow / roboco-do / roboco-git-readonly. That's repeated token
spend on every refused retry; the LLM doesn't need the full alt-list
inline, it has the role prompt + the MCP tool schema for that.
Trimmed to 2 lines: denial reason + a one-line pointer to the role's
State→Verb table. Test asserts <= 3 echo lines in any denial block.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section B5.
Smoke run 3 showed Main PM taking 7 attempts to satisfy the decision-note
required-fields contract — the remediate listed which fields were
missing but didn't show what a fully-formed call looks like. The LLM
pattern-matches examples better than field-list prose; each retry it
dropped a different field.
Added a literal note(scope='decision', ...) / note(scope='reflect', ...)
call template to the rejection remediate so the agent sees the canonical
shape with named-keyword args and example values. The missing-fields
list stays — both pieces of information are useful, but the example is
what actually drives convergence.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section B4.
Three classes of failure, all surfaced once Wave A's plan-required gate
and the migration 013 went in. Per project standing rule: pre-existing
errors are not a free pass — fix them.
1. Wave A1 ordering (32 lifecycle parity failures + 1 full-pipe test).
_pm_sub_tasks_gate fired BEFORE _claim_plan_start_gate, so wrong-state
PMs got `incomplete_input` (the gate's verdict) when the spec's
lifecycle gate should have returned `invalid_state` first. Swapped:
re-entry check → spec lifecycle gate → sub_tasks gate → claim_plan_run.
Parity test now sees the spec's verdict as expected.
2. E2 enum naming (2 migration_013 failures + ripple).
_str_enum in roboco/db/tables.py didn't pass name=… to SQLAlchemy
Enum(...), so Base.metadata.create_all in test setup inferred
`role` from the Python class `Role` while the alembic migrations
declare `agentrole`. Tests saw two enums for the same class and
hit `agentrole = role` operator errors. Fixed: default name to
lower(class_name) (matches every migration), override `Role` →
`agentrole`. One dict entry; no class-by-class registration needed.
3. _MockContentActions.note() signature drift.
Wave 2 G4 added `structured` kwarg to ContentActions.note().
The integration mock at tests/integration/v2/test_full_pending_to_completed.py
didn't accept the new kwarg → 1 test failed on the very first call
from the v2 do/note route. Added `structured: object = None` and
left it unused (the test asserts lifecycle, not journal rendering).
Plus three ruff E501 line-length fixes in the test files I touched.
Quality: ruff + mypy clean. pytest 6690 passed / 0 failed / 274 skipped.
Smoke run 3 showed inconsistent return strings — main-pm got
status='sent', be-pm got status='posted' for the same verb.
Confirmed say() already returns 'posted' at its sole success exit.
Added test_say_status.py to pin the canonical past-tense pattern
(note->'noted', say->'posted', notify_ack->'acked') and prevent
regression. dm() and notify() retain 'sent' — different verbs,
different semantics.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section B3.
Smoke run analysis initially flagged three Task fields as unused
(pm_approvals, quick_context, proactive_context). A follow-up audit
found quick_context (stores original_developer marker + doc notes +
PR creator + escalation notes) and proactive_context (RAG injection)
are actively used. Only pm_approvals is truly orphaned.
Migration 014 drops pm_approvals; downgrade() recreates it if ever
needed. The two false-positive fields stay untouched.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section B2 (re-scoped 2026-05-12).
Smoke run 3 showed stop-hook.sh complaining 'Denied: you stopped
without calling a terminal tool' AFTER agents successfully called
i_am_idle() — because the hook listed 9 pre-gateway verb names
(roboco_agent_idle, roboco_task_substitute, etc.) that no longer
exist. Same staleness in bash-guard-hook.sh.
Both hooks now reference current gateway verbs only. stop-hook
branches its suggestion by ROBOCO_AGENT_ROLE so devs see
i_am_done/i_am_blocked, QAs see pass/fail, PMs see complete/escalate_up.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section B1.
Smoke run 2 (2026-05-11) produced 'UndefinedFunctionError: operator
does not exist: agentrole = role' because postgres had two enums
(role, agentrole) for the same Python class. Information_schema check
confirms no column uses role; migration drops it. Upgrade() raises if
that ever stops being true. Downgrade() recreates the enum with the
foundation's Role values.
Investigation of WHY a second enum got created is tracked in spec E2;
this migration handles the symptom.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section A5.
Smoke run 3 fired the same workspace.py warning ~9x per run:
'ensure_workspace: refresh fetch returned non-zero'
stderr: 'fatal: could not read Username for https://github.com'
This is EXPECTED behavior, not a bug. The docstring on
_fetch_origin_best_effort explains that credentials are deliberately
scrubbed from .git/config after the initial clone (part of the secret-
exfiltration mitigation) and refresh fetches are best-effort. For
private repos the auth-fail is the documented outcome.
The original A4 spec proposed re-injecting the PAT -- that would have
violated _assert_no_pat_leak and the URL-scrub mitigation. Re-scoped
to: silence the known-benign signature at DEBUG, keep WARNING for
genuine failures (network errors, broken remotes, repo-not-found).
No behavior change. No security boundary touched. Just log level.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
A4 (re-scoped 2026-05-12 after investigation showed the original spec
proposed reintroducing a documented security regression).
Fixes 2 important + 1 minor issue from the code-quality review of 5adb4ff:
1. Formula duplication: the workspace path string was inlined at two
sites in orchestrator.py (the canonical _prepare_agent_spawn and the
new _build_mount_args -w logic). Extracted to module-level helpers
_agent_workspace_path(project, team, agent_id) and
_cell_workspace_path(project, team) so both callers share the same
formula. Future path changes only land in one place.
Also extracted _resolve_project_slug_from_git_context() as the
module-level counterpart to the instance method, called by the static
_build_mount_args site that cannot access self.
2. Test consistency: test_workdir_matches_edit_allowlist_path now
extracts the Edit(<prefix>/**) value from _get_role_permissions and
asserts the spawn cmd's -w value equals that prefix. The test would
actually catch a drift where _build_mount_args and _get_role_permissions
use different formulas — previously it just compared two copies of
the same string.
3. Test coverage: added test cases for product_owner and head_marketing
spawns (both share the per-agent workspace path), so all roles that
_get_role_permissions distinguishes are covered.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
A2+A3 (re-scoped 2026-05-12).
Smoke run 3 surfaced two bugs that share a root cause:
- Edit(/app/README.md) → 'Edit exists but is not enabled in this context'
- commit(files=['/app/README.md']) → 'outside repository at <workspace>'
Both happened because the container's WORKDIR is /app (roboco package
source) while the agent's task workspace is bind-mounted at
/data/workspaces/<project>/<team>/<agent>/. The Dev role's
Edit/Write permission allowlist scopes to the workspace, so any Edit
call from /app fails the path match.
Adds '-w {workspace_path}' to the docker run command so the container
starts with cwd = task workspace. Edit(README.md) and git add README.md
now resolve inside the workspace clone.
Mirrors _get_role_permissions path selection exactly:
- developer / product_owner / head_marketing: per-agent workspace
- documenter: cell workspace (matches its Write/Edit allowlist)
- qa / cell_pm / main_pm / auditor: omit -w, fall back to /app
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
sections A2 + A3 (re-scoped per investigation 2026-05-12).
Three fixes from the code-quality review of a1009c0:
1. Critical: _pm_sub_tasks_gate ran before _handle_pm_reentry, breaking
idempotent re-entry for PMs whose containers crashed mid-run. Moved
the gate to after the re-entry short-circuit so initial-claim is the
only path that hits the gate.
2. Critical: gate had no direct unit test (the HTTP-layer test mocked
the choreographer). Added tests/unit/gateway/test_i_will_plan_sub_tasks_gate.py
with six tests: empty sub_tasks → incomplete_input, missing rich_plan
→ incomplete_input, filled sub_tasks → gate passes, developer with
empty sub_tasks → gate passes (devs don't decompose), sub_tasks filled
but approach empty → incomplete_input, in_progress re-entry short-
circuits before gate even with no sub_tasks.
3. Important: approach was only enforced at the HTTP Pydantic boundary.
Direct service-layer callers (MCP, test fixtures, orchestrator-
internal) could persist a plan with no approach. Gate now also checks
approach >= 20 chars (_PM_APPROACH_MIN_LEN constant) and includes it
in the rejection's missing list when absent.
Plus: stale docstring on i_will_plan corrected; _handle_pm_reentry
docstring rewritten to lead with the domain reason (re-entry contracts)
not the PLR0911 linter justification; unused pytest import removed from
test_i_will_plan_rich_required.py; pre-existing PLR2004/PLC0415 issues
in that file fixed.
i_will_plan now requires approach (min_length=20) at the schema and
non-empty sub_tasks at the gateway when the caller is a PM role. Restores
pre-gateway parity for _validate_claimed_start — agents could not
transition claimed -> in_progress without filling the rich plan.
Smoke run 3 (2026-05-11) showed PMs calling i_will_plan with just
plan='paragraph' and the gateway accepting it; Plan tab stayed empty
because no agent filled approach/sub_tasks/risks/open_questions.
Spec ref: docs/superpowers/specs/2026-05-12-post-smoke-3-fixes-design.md
section A1.
Wave 1 wired notify_list/get/ack into ContentActions but pointed them at
`self.notifications` (which is NotificationService — sender side, with
send_blocker_notification / send_qa_ready_notification / etc.). The
read methods (list_for_agent, get_for_recipient_and_mark_read,
acknowledge) live on `NotificationDeliveryService` instead.
Smoke run 2026-05-11 surfaced this immediately:
AttributeError: 'NotificationService' object has no attribute 'list_for_agent'
Fixes:
- roboco/api/deps.py — import NotificationDeliveryService and wire it
in as a new ContentActionsDeps field `notification_delivery`.
- roboco/services/gateway/content_actions.py — add notification_delivery
to ContentActionsDeps (Optional with `None` default for back-compat
with any tests that don't supply it). Point notify_list, notify_get,
notify_ack at self._deps.notification_delivery.
- tests/unit/gateway/test_content_actions.py — _make_deps adds a default
AsyncMock for notification_delivery so existing tests continue to pass.
Quality: ruff + mypy clean. 505 unit tests pass.
Pre-gateway parity (G8 part b of the 2026-05-11 design). The pre-gateway
TaskBlockInput at 254cc93:roboco/mcp/schemas/__init__.py required
blocker_type (external|internal|question|dependency) and what_needed
so PMs could triage their inbox by class. Current i_am_blocked dropped
both fields — every blocked task looked the same to the PM.
Now i_am_blocked accepts both as optional kwargs:
- Back-compat: callers that omit them still work (blocker_type defaults
to None → rendered as flat reason in the struggle entry).
- New: when supplied, the struggle journal entry body is structured
markdown (## Blocker Type / ## What Needed sections) so the panel's
journal view renders named blocks instead of one flat sentence.
Validator on blocker_type enforces the enum at the Pydantic boundary
with a clear "must be one of: ..." error if the agent invents a value
(same pattern as the Wave 3 G7 validators).
G8 part a — typed `pause(checkpoint_summary, remaining_work)` — defers.
That gap needs a new IntentSpec in foundation/policy/lifecycle.py
(currently pause is an ActionSpec only; agents auto-pause via i_am_idle)
plus checkpoint wiring through TaskService.add_checkpoint. Material
work, deferred until after the user has deployed and verified G7 +
G8b lands cleanly.
Wired:
- roboco/api/schemas/v2/flow.py — IAmBlockedRequest gains optional
blocker_type + what_needed; @field_validator enforces the enum
- roboco/api/routes/v2/flow_dev.py — passes the new fields through
- roboco/services/gateway/choreographer/_impl.py — i_am_blocked signature
+ structured struggle-entry rendering
- roboco/mcp/flow_server.py — typed wrapper with the kwargs
- agents/prompts/roles/developer.md — updated verb table
- tests/unit/mcp_servers/test_flow_server.py — updated to expect the
new optional kwargs as None when omitted
Quality: ruff + mypy clean. 505 tests pass.