Files
roboco/tests/unit/gateway/test_choreographer_impl_branches.py
T
110aaa7a77 Chore: v1 removal gateway canonical (#46)
* chore(agent_sdk): remove dead /traceability/remind endpoint and reminder map

The TRACEABILITY_REMINDERS dict and its /traceability/remind endpoint were
keyed entirely on pre-gateway tool names (roboco_task_*, roboco_journal_*,
roboco_message_send, roboco_session_create_for_tasks) deleted in the gateway
cutover. The endpoint had zero callers; v2 enforces traceability server-side
in the Choreographer.

* fix(bootstrap,seeds): onboarding prompts call give_me_work(), not deleted roboco_task_scan()

The startup prompt and the seeded cell/all-hands channel onboarding messages
instructed agents to call roboco_task_scan() — a tool removed in the gateway
cutover. Point them at the live give_me_work() flow verb.

* fix: replace remaining deleted v1 tool names with gateway verbs

Spawn prompts, onboarding strings, remediation messages, and comments still
referenced pre-gateway tools deleted in the cutover (roboco_task_*,
roboco_agent_idle, roboco_notify_*, roboco_message_send,
roboco_session_create_for_tasks, roboco_journal_*, roboco_escalate). Rewrote
each to the correct role-scoped gateway verb (give_me_work/i_will_work_on for
workers, triage for PMs, i_am_done vs complete, notify/notify_ack, escalate_up,
unclaim, i_documented, open_session, note). Updated one enforcement-message
test that matched the old tool name by coincidence.

* test: guard against deleted v1 tool names reappearing in roboco/

Scans roboco/ for the deleted pre-gateway tool names; excludes the orphaned
roboco/agents/ subtree (removed in a later phase).

* chore(exceptions): drop 8 unused pre-gateway exception classes + their tests

LLMError, RAGError, AlreadyExistsError, TaskBlockedError, TaskClaimError,
AgentNotAvailableError, AgentBusyError, NotificationPermissionError were never
raised in production. SessionClosedError/DatabaseError are kept (live + tested).

* chore(models): drop unused pre-gateway notification/channel/handoff factories

Removes create_task_assignment/_blocker_escalation/_review_request/
_documentation_request/_priority_change/_alert/_broadcast, create_cell_channel/
_cross_cell_channel/_announcements_channel, create_handoff (+ HandoffParams),
ProactiveContext, and A2APartType. The gateway choreographer builds these
server-side now. Drops the matching dead-code tests.

* chore(services): drop unused pre-gateway permission/messaging/audit/optimal/remediation methods

These pre-gateway helpers (channel-permission checks, channel-membership ops,
permission-denial audit hooks, doc ingestion, two remediation hints) have no
production caller — the gateway role_config + enforcement layer replaced them.
Drops the matching dead-code tests; live methods (send_message, the SESSION_*
flow, log_task_action_denial, etc.) are untouched.

* chore(orchestrator,ws,events,config): drop unused pre-gateway lifecycle/broadcast/roster symbols

orchestrator: get_running_agents, is_agent_busy, queue_priority_work,
get_all_instances (+ their OrchestratorAccessProtocol declarations in events.py).
websocket: broadcast_new_message, broadcast_session_closed (no event type emits
them). agents_config: ALL_PMS/ALL_DEVS/ALL_QA/CELL_PMS roster constants (ALL_DOCS
stays — it gates docs-write workspace perms).

* refactor(agents): delete orphaned pre-gateway agent subtree + dead organization model

The Gateway/full cutover replaced the Python agent-class implementations with
the server-side Choreographer; the classes survived only as a self-referential
island. Removes roboco/agents/{base,mixins,factory,board,developer,documenter,
pm,qa,orchestrator}.py and roboco/agents/factories/{board,cells,developers,
documenters,pms,qa}.py, plus roboco/models/organization.py (Cell/Board/
Organization — used only by those factories). Keeps factories/_base.py
(compose_prompt — the live prompt-layer composer the orchestrator calls at
spawn) behind minimal package __init__ files.

* chore(db): drop dead tasks.execution_log + outputs columns (migration 015)

Both JSON columns had zero readers/writers in code, tests, and migrations —
execution progress is tracked via progress_updates and artifacts via
commits/documents. Removes the ORM columns, the Pydantic Task.execution_log/
outputs fields, the ExecutionLog/FileRef models (+ their __init__ exports), and
the now-invalid kwargs from test fixtures. Migration 015 (down_revision
014_drop_pm_approvals) verified live: upgrade drops, downgrade re-adds.
Apply on the NAS with 'alembic upgrade head' at next deploy.

* chore(config): drop 16 unread Settings fields

Verified unused (no settings.X, no self.X property use, no getattr-by-name):
app_name, reload, workers, openai_api_key, secret_key, access_token_expire_minutes,
algorithm, log_level, log_format, the four session_* limits, message_max_length,
commit_subject_min_chars, commit_banned_words, agent_budget_sweep_interval_seconds.
Removes the empty Logging + Sessions&Messages sections and orphaned .env.example
vars. Kept: redis_db/redis_password (redis_url property), agent_sla_* (read via
getattr in task_lifecycle), encryption_key, and all live thresholds.

NOTE: commit_banned_words/commit_subject_min_chars and
agent_budget_sweep_interval_seconds were feature-config never wired to their
consumer (commit validator / budget sweep) — removed as dead, but flagged in
case the intent was to wire them.

* test(lifecycle): give i_will_work_on calls a substantive plan (#171 contract)

The real-DB lifecycle tests called i_will_work_on with a 13-char plan and no
risks/technical_considerations, so the substantive-plan gate (#171) rejected
them with incomplete_input — failing on master. Supply a >=150-char plan plus
technical_considerations and risks (mirroring tests/unit/gateway/
test_choreographer_dev.py). All 6 now pass; gate runs with no deselect.

* feat(gateway): wire commit-validator thresholds to settings

commit_subject_min_chars and commit_banned_words were config defined but never
read — the gateway commit() gate used the validator's hardcoded module defaults.
Re-add the two Settings fields and pass them through validate_commit_message in
content_actions.commit(), so config is the source of truth (validator defaults
remain the standalone/CI fallback). Adds wiring tests that monkeypatch settings
and assert the gate honors them.

* refactor(orchestrator): retire gateway_enabled flag; trigger_filter is unconditional

The gateway_enabled Settings field gated only the trigger_filter spawn-cooldown
(never the agent tool surface). Prod ran it on; the Phase-0 'legacy dispatch
path' it guarded no longer exists. Remove the field + the early-return branch in
gateway_pre_spawn_check so the cooldown runs for every spawn, drop the now-dead
ROBOCO_GATEWAY_ENABLED from docker-compose.yml, and update the stale Phase-0
comments + cooldown test. The per-container ROBOCO_GATEWAY_ENABLED env (set by
_append_manifest_args, read by agent_sdk to load the manifest) is unaffected.

* refactor(api): relabel /api/v2 -> /api/v1 as the canonical gateway surface

The gateway is the only agent API now, so the 'v2' label (with no v1) was
misleading. Renames roboco/api/routes/v2 -> routes/v1, schemas/v2 -> schemas/v1
(+ the matching test dirs and test_v2_role_dep/test_schemas_v2_flow files),
rewrites every /api/v2 path, routes.v2/schemas.v2 import, and v2-* router tag to
v1, and refreshes the stale 'v2' comments/docstrings. The panel is untouched (it
uses the unversioned /api/* REST routes). flow_server/do_server now POST to
/api/v1/*.

* docs(scripts): reset_runtime_state header matches actual SQL behavior

The header claimed it preserves groups + journals, but the .sql wipes both
(verified live: groups 6->0, journals 5->0; only agents/projects/channels
survive). Correct the wiped/preserved lists to match.

* refactor(gateway): extract _build_rich_plan to drop i_will_work_on under the complexity gate

i_will_work_on was cyclomatic rank C (11) — one over the xenon --max-absolute B
threshold — because of the five `x or default` fallbacks in the rich_plan dict.
Move that dict into a small _build_rich_plan helper (behaviour identical); both
methods are now rank B. make quality is fully green (xenon was its last failure;
bandit already passed — its 34 findings are all LOW severity, filtered by -ll).

* feat(foundation): add canonical CELL_TEAMS set; dedupe cell-subset literals

* feat(db): add ProductTable + ProductProjectTable ORM (per-cell project map)

* feat(task): add additive nullable product_id (ORM + model + DTO + create threading)

* feat(task): thread product_id through create_subtask/route/response

* feat(db): migration 016 — products, product_projects, tasks.product_id

* fix(db): document migration 016 plan deviations (revision len, FK name)

Two values in migration 016 intentionally diverge from the Task 2.4 plan
literals; this strengthens the in-file justification so the deviations are
self-documenting and verifiable.

- revision id (plan line 623): the plan's 36-char
  "016_add_products_and_task_product_id" overflows alembic_version.version_num
  (VARCHAR(32)) — alembic upgrade head raises asyncpg
  StringDataRightTruncationError. Kept at 27 chars
  ("016_add_products_product_id") so Step 4's live round-trip stays green.
- downgrade FK name (plan line 683): roboco/db/base.py sets a metadata
  naming_convention, so the FK upgrade() creates is
  "fk_tasks_product_id_products", not the Postgres default
  "tasks_product_id_fkey". The plan literal does not exist in the DB and
  would fail the downgrade with "constraint does not exist".

Both verified via the live upgrade/downgrade round-trip on a throwaway DB.

Issue 3 note: the prior commit (b896cac) also touched
tests/unit/api/test_schemas_tasks.py (added product_id=None to the
task_to_response stub). That line is load-bearing — task_to_response reads
task.product_id (added in Task 2.3, commit 67afa6b) — and belongs to Task 2.3's
scope; it is left in place because removing it breaks 4 tests and history is
not rewritten.

* refactor(db): trim migration 016 deviation notes to plan-faithful form

Reverts the out-of-scope documentation expansion (commit 1a4f296), which
was a second undocumented commit beyond Task 2.4's single plan-specified
commit and only bloated the migration docstring/comments.

The migration file now matches the plan-specified commit (b896cac) byte for
byte: the two necessary deviations from the plan literals stay (revision id
shortened to fit alembic_version.version_num VARCHAR(32); downgrade FK name
follows db/base.py's metadata naming_convention), each kept to a concise
inline note in the plan's header style.

The Task 2.3-scoped test stub line (tests/unit/api/test_schemas_tasks.py
product_id=None) is load-bearing — task_to_response reads task.product_id —
and is left in place; history is not rewritten.

Verified: live alembic upgrade head + downgrade to 015 round-trip on a
throwaway DB drops products/product_projects/tasks.product_id cleanly, and
make quality is green.

* refactor(test): annotate db_session and drop type: ignore in migration 016 test

Annotate the test_products_tables_and_task_fk_exist param as
db_session: AsyncSession (imported under TYPE_CHECKING) and remove the
# type: ignore[no-untyped-def] suppression, matching the typed db_session
pattern used across tests/integration/.

* feat(models): Product + ProductCreate/Update + ProductCellMapping (cell-validated)

* refactor(models): minimize ProductCellMapping config override to use_enum_values

The previous override re-declared validate_assignment, populate_by_name,
and extra=forbid, which RobocoBase already supplies. Pydantic merges
model_config across inheritance, so overriding only use_enum_values=False
is sufficient to keep team as a real Team enum (required so team in
CELL_TEAMS and enum identity hold for callers) while inheriting the rest
of the base config.

* fix(models): document ProductCellMapping use_enum_values override as plan-mandated

Resolves SPEC-COMPLIANCE review notes for Task 3.1 (Product domain models).

1. The ProductCellMapping use_enum_values=False override is a deviation from a
   bare project.py mirror, but it is mandated by the plan's own Task 3.1 code:
   RobocoBase sets use_enum_values=True, which coerces team to the plain string
   "backend". The plan's Step 1 test asserts m.team is Team.BACKEND (enum
   identity) and the Step 3 validator formats its error with v.value, both of
   which require team to remain a real Team enum. The override is therefore
   necessary; this commit relabels the comment to cite the specific spec lines
   that force it instead of leaving it as an unexplained departure. Downstream
   Task 3.2 (_replace_cells / project_for) already tolerates either form and the
   ORM stores the same value regardless, so the override has no behavioral reach
   beyond the in-memory enum identity the plan's test checks.

2. test_product_model.py hoists 'from uuid import uuid4' to module level rather
   than inline (as the plan's verbatim Step 1 code shows) because the global
   Pylint PLC0415 rule (import-outside-top-level) forbids inline imports and
   there is no per-file-ignore for tests/unit/models/. The hoisted form is the
   only ruff-clean rendering of the plan's test; left unchanged here.

3. Task 3.1 landed across two commits (c616d95 create, 6ebad255 refactor) rather
   than the plan's single Step 5 commit. Earlier history is intentionally not
   rewritten; this single follow-up commit brings the model to its final
   spec-faithful, fully-documented state.

* feat(service): ProductService CRUD + project_for per-cell resolver

* feat(api): Product CRUD routes + schemas, wired into the app

* fix(api): roll back and map cell-replacement IntegrityError on product update

update_product replaced cells via ProductService._replace_cells without
any try/except, so a duplicate-team cell (uq_product_projects_product_team)
or a non-existent project_id (product_projects.project_id FK) raised an
IntegrityError at flush, poisoning the AsyncSession and surfacing an
unhandled 500 with no rollback. Wrap the update + commit in a try/except
that rolls back and maps the UNIQUE violation to 409 and the FK violation
to 422, mirroring create_product's rollback discipline. Add integration
tests covering both client-error paths.

* fix(api): map create_product cell-mapping IntegrityError to 409/422

create_product only caught the slug conflict ('already exists' in str(e))
and bare-raised everything else, so a cells entry whose project_id does not
reference any project let the product_projects.project_id FK IntegrityError
propagate out of the route as an unhandled 500. The matching update_product
path was already hardened (uq_product_projects_product_team -> 409, FK
violation -> 422); apply the same mapping in create_product so a bad
project_id (or a duplicate-team cell) is a client error, not a server error.
The slug conflict is now caught as ConflictError directly instead of via a
broad except + string match.

* feat(gateway): add optional project_id to delegate inputs/request/routes

* feat(gateway): per-cell project routing (override -> product map -> parent) + product_id inheritance

* feat(task): approve_and_start — reassign board task to Main PM (CEO gate #1)

* feat(api): POST /tasks/{id}/approve-and-start (CEO gate #1, notes-required)

* test(api): cover approve-and-start 404-before-notes-gate for missing task

* feat(panel): Product types + Task.product_id

* feat(panel): productsApi + hooks + tasksApi.approveAndStart

* feat(panel): Products management screen + sidebar nav

* feat(panel): Approve & Start button (CEO gate #1)

* fix(api): narrow delete_product to IntegrityError + cover 204/409 delete paths

* test(task): assert approve_and_start persists + appends the audit note

* refactor(db): migration 016 names the tasks.product_id FK explicitly (house style)

* fix(db): make migrations authoritative + self-heal orphan product tables

init_db() no longer silently falls back to create_all when alembic upgrade
fails. That fallback masked migration failures and, since create_all cannot
ALTER an existing table, left the schema inconsistent — turning an unapplied
migration 016 into a crash loop: 016's CREATE TABLE products failed, the
upgrade rolled back, create_all re-created an empty orphan products table, and
every later boot failed again on the now-existing table while tasks.product_id
never got added. Now a migration failure is raised so the real error surfaces.

Migration 016 additionally drops EMPTY orphan products/product_projects tables
left by the old fallback before creating them, so an already-polluted DB
self-heals on the next deploy with no manual SQL. Skipped in offline (--sql)
mode; refuses to drop a table that holds rows.

* fix(db): create_all is the schema source of truth; alembic for increments

The Alembic chain is incomplete relative to the ORM — columns/tables like
notifications.delivered_at and the RAG indexed_documents table have NO migration
and have only ever been materialized by create_all. Tests don't catch this
because the test DB is also built via create_all, so migrations are never
exercised. The prior 'migrations are authoritative' init_db (and before it, the
create_all-only-on-failure fallback) therefore left a migrate-only boot with
missing columns/tables.

init_db now reflects reality:
  - Fresh DB  -> create_all builds the full current ORM schema, then stamp
                 Alembic at head so later incremental migrations apply.
  - Existing  -> run pending migrations (a real failure is raised, not masked),
                 then create_all(checkfirst) to gap-fill any missing ORM tables.
create_all cannot add a column to an existing table, so an ORM column added
without a migration needs a fresh rebuild of that table to appear.

* fix(db): migration 017 reconciles the Alembic chain with the full ORM schema

For years the live schema was built by create_all, not migrations, so the chain
drifted — tables/columns/indexes in the ORM had no migration (the
indexed_documents table, notifications.delivered_at, ~15 indexes, plus
timestamptz/server-default metadata). With init_db no longer masking that via a
create_all fallback, a migrate-only boot was missing those objects.

017 was produced by 'alembic revision --autogenerate' against Base.metadata,
reviewed, and verified: on a fresh DB, 'alembic upgrade head' (001..017) now
reproduces the create_all schema EXACTLY — a re-run of autogenerate detects zero
changes — and the 017 upgrade/downgrade round-trips cleanly. The migration chain
is now complete: migrate-only and create_all converge.

Also updates the init_db tests to assert the new behaviour (raise on an existing
DB's migration failure; create_all + stamp head on a fresh DB) instead of the
removed silent fallback.

* feat(panel): Product picker in the New Task form (drives per-cell routing)

The Products screen and Approve & Start button shipped, but the task-creation
form had no way to attach a Product — so a human couldn't set product_id from
the UI, which is exactly what drives per-cell project routing of delegated
subtasks. Adds an optional Product dropdown (Advanced -> Git config) populated
from useProducts(); 'None' falls back to the single project.

* fix(db): seed data is preserved on a fresh DB (run migrations, not bare create_all)

The previous fresh-DB path (create_all + stamp head) built the tables but never
ran the migration chain, so migration-embedded SEED DATA was skipped — most
visibly the AI providers seeded in 004. After a DB reset that left
provider_configs empty, so PUT /api/providers/ollama-key 404'd (the handler
raises NotFoundError when the Ollama provider row is missing).

Since migration 017 made the chain reproduce the full ORM schema, init_db now
runs 'alembic upgrade head' from base on a fresh DB — building every
table/column/index AND running the seeds. Verified: a fresh upgrade head seeds
both provider rows. Existing DBs still get migrations + create_all gap-fill.
Updates the init_db fresh-DB test accordingly.

* feat(task): project_id optional when a product_id is set (board fan-out tasks)

A board task that fans out across cells via a Product has no single repo of its
own — backend/frontend/ux_ui are each wrong, because the root coordinates and
delegates. Forcing one arbitrary Project was broken design (flagged at design
time). project_id is now nullable; a task must have project_id OR product_id:
  - TaskCreate model validator + a TaskService.create() invariant (covers every
    create path).
  - ORM/DTO/schema: project_id nullable; task_to_response uses to_python_uuid.
  - Gateway: a parent with only a product can delegate (guard now needs BOTH
    project and product to be None to reject); _resolve_subtask_project resolves
    each subtask from the product map and raises a clear error if a cell has no
    mapping and no parent project.
  - Migration 018 (tasks.project_id nullable), round-trip verified; fresh
    upgrade head still seeds providers.
  - Panel: Project no longer required once a Product is selected.
  - Removed the dead, never-called a2a create_task_from_message (it could only
    ever create a repo-less task) + its two coverage-only tests.

make quality green; panel tsc/lint/build green.

* Upgrade to Minimax M3

* fix(db): seed providers on existing DBs + correct enum casing

Migration 004 created the modelprovider/assignmentscope enums and seeded
provider rows in UPPERCASE, but the ORM (_str_enum) reads/writes the
lowercase StrEnum .value — so a fresh migrate-from-base DB built an enum
the ORM cannot read. Lowercase the enum labels and seed values in 004.

Add idempotent migration 019 to (re)seed the Anthropic + Ollama Cloud
providers with ON CONFLICT (name) DO NOTHING, so an existing DB whose
provider_configs table was created by create_all (and never ran 004's
seed) gets the rows on the next `alembic upgrade head` — fixing the
/api/providers/ollama-key 404 without a volume wipe.

* fix(tasks): let board/fan-out coordination tasks flow without a repo

A coordination task (project_id NULL, product_id set) targets no repo of
its own — it fans out to cell subtasks that each resolve a real project
from the product's cell->project map. Several paths still assumed every
task does git work and blocked it:

- orchestrator: add _is_coordination_task() and exempt these tasks from
  the project/branch/git-token gates in _readiness_check_task,
  _readiness_gate, _check_stuck_conditions, _validate_task_for_spawn.
- services/task.py: _ensure_branch_for_task returns "" (no branch) for a
  coordination task instead of raising; activate requires project OR
  product. This unblocks Main PM's i_will_plan claim, which otherwise
  raised before it could delegate the fan-out.
- gateway: _pending_assignment_guard exempts advisory roles
  (product_owner/head_marketing/auditor) from the "assigned but never
  claimed" idle gate — they review without claiming, so they could not
  satisfy a claim-or-unclaim remediation.

Adds focused unit tests for each.

* fix(tasks): coordination tasks reach in_progress + team reflects Main PM

The board->cells fan-out deadlocked: a coordination/fan-out task (product set,
no project of its own) could be created and claimed, but start()'s
claimed->in_progress transition hit validate_git_requirements, which still
demanded a branch_name and raised GitRequirementError. So Main PM's i_will_plan
never completed — it looped and never delegated. c961282 exempted
_ensure_branch_for_task (branch creation) but missed this parallel git gate in
the enforcement layer.

- task_lifecycle.py: add GitContext.is_coordination; skip the
  claimed->in_progress branch_name gate when it is set.
- task.py: populate is_coordination=(project_id is None and product_id is not
  None) in _validate_and_set_status; a branchless code task is still gated.
- approve_and_start: set team=Team.MAIN_PM on hand-off so the task isn't left
  labelled team=board after it leaves the board (now assigned to main-pm).

Adds a lifecycle-gate unit test and an end-to-end integration test that
claims, plans, and starts a project-less coordination task.

* fix(hooks): remove dead traceability hook + stale deleted-verb references

The v1-removal cleanup (2cfbf39) deleted the /traceability/remind SDK endpoint
but left the PostToolUse hook that curls it, so every gateway tool call 404'd
and agents silently lost their traceability reminders. Remove the dangling hook
(registration + TRACEABILITY_TRIGGER_TOOLS + Dockerfile COPY + the script); v2
carries per-verb guidance on the Envelope. Also correct two stale pre-gateway
tool names in hook text: the budget loop-detector nudged agents toward the
deleted roboco_task_escalate() (now unclaim()/i_am_idle(), which every looping
role has), and an sdk-startup comment referenced roboco_task_scan/get.

Extends the deleted-tool-name guard to scan docker/scripts/*.sh and to assert
every $SDK_URL/<path> a hook curls is a route still served by the SDK — the
check that would have caught this class (it lives in shell, invisible to mypy
and the Python import graph).

* fix(db): backfill ORM enum values the migration chain never added

Several StrEnum values were added to the ORM over time without a matching
`ALTER TYPE ... ADD VALUE` migration; 017 was autogenerate-derived and
autogenerate does not detect added enum labels, so the drift survived. On a DB
whose enum type predates the value, binding it raises at runtime — e.g.
`invalid input value for enum notificationtype: "a2a_request"` on
GET /api/notifications (list_system_notifications), and the same class for
blockerresolvertype/handoffstatus/team.

Migration 020 adds every drifted value idempotently (ADD VALUE IF NOT EXISTS —
no-op when 009 already reconciled it). Runs on the next `alembic upgrade head`.

Detected by comparing each ORM enum's values to the labels the migration chain
produces; adds tests/unit/test_enum_migration_parity.py which renders the chain
offline and fails on any future drift — the check that would have caught both
this and the provider-enum bug.

* fix(orchestrator): stop branch auto-block, board reassign, unblock livelock, agentless claims

Cluster C1 — four coupled orchestrator/task-invariant defects:

#18: a branch is created only at claim, so a pending, never-claimed code task
legitimately has no branch_name. The stuck-detection sweep (pending-only) and
readiness gate flagged that as "Task missing branch_name" and auto-blocked the
task every tick, so it never dispatched. Centralize the gate in
_branch_is_expected (status in claimed/in_progress/verifying, never a
coordination task) and apply it in both _check_stuck_conditions and
_readiness_check_task.

#14: the main_pm -> product_owner escalation rung handed an in_progress
descendant code task to the Product Owner (a board role) and marked it BLOCKED;
the board has no verb to own code work, so the dev's finished work deadlocked.
TaskService.apply_escalation (the single write primitive — covers both the
gateway escalate verb and the HTTP escalate route) now diverts a descendant code
task targeting a board/advisory role: it releases the task to PENDING for a
role-matched cell claim instead of stranding it.

#17: a blocked task reassigned to Main PM kept respawning the ex-assignee cell
PM to unblock it, but the assignee-only pre-unblock note returned not_authorized
— a livelock. _dispatch_blocker_work now dispatches the task's CURRENT PM/board
assignee (the unblock authority), falling back to the cell PM only when no
PM/board holds it. Also: a branchless coordination parent yields no valid merge
target — resolve_parent_branch now falls back to the child's own project default
branch (e.g. master) via TaskService.project_default_branch_for_task, and
_check_parent_branch_ready no longer blocks a child on a coordination parent's
non-existent branch.

#19: a task left claimed/in_progress with an assignee but no running container
was invisibly stuck (only PENDING tasks get fresh dispatch; the heartbeat reaper
can't see a freshly-seeded claim). New _dispatch_claimed_without_agent net:
after a short grace window it respawns the assignee, or releases the claim to
pending (lifecycle-safe via unclaim_for_reaper) when the assignee is unknown.
New config ROBOCO_CLAIMED_NO_AGENT_GRACE_SECONDS (default 120).

* fix(gateway): tolerant note verb + lock evidence do-tool invariant

#15: the note verb no longer hard-rejects thin decision/reflect payloads.
List-typed fields (options, consequences, next_steps) coerce a lone scalar
into a one-element list at both the NoteRequest schema (mode=before
validator) and the service layer; missing narrative fields default to a
visible placeholder instead of returning incomplete_input. The note is
always recorded, preserving audit value, and a well-intentioned note can no
longer trip the do-server 3-strikes circuit breaker. Widen the agent-facing
do_server.note hints to accept list-or-scalar and refresh the docstrings.

#8: add regression coverage locking the invariant that every role's do_tools
carries evidence (role_config + developer spawn manifest). The current source
already registers mcp__roboco-do__evidence for developers end-to-end; the
report stemmed from a stale deployed build, and the tests prevent silent
regression.

* fix(gateway): allow UX devs to receive design tasks; surface delegation rules to cell PM

The UX/UI cell's developers (ux-dev-1/ux-dev-2, Role.DEVELOPER on
Team.UX_UI) ARE its designers, but _validate_assignee_task_type rejected
task_type='design' for every DEVELOPER, blocking the UX cell's normal
design delegation. Allow 'design' for UX-team devs only; backend/frontend
devs stay rejected (design routing belongs to the UX cell). The
orchestrator already dispatches a developer for a design task
(_dev_dispatch_role_matches returns True), so this creates no orphan like
the documentation case.

Replace the static Cell-PM 'pass planning' remediate with a per-assignee
hint so a dev/design mis-type gets a developer-class next-step instead of
an off-topic planning hint.

Surface the three delegation guardrails in the cell-PM prompt so PMs stop
probing them by trial and error: valid task_type per assignee (incl.
design for UX devs), documentation auto-creation (non-delegatable), and
the sequential single-active code-spine. Fix the delegate-row task_type
list (documentation is NOT delegatable) and update the lifecycle spec
description; regenerate the lifecycle artifacts.

* fix(orchestrator): improve agent briefings for handoff consumption, product/project model, and workspace/secret hygiene

Main PM (roles/main_pm.md):
- Require reading the upstream Product Owner / Head of Marketing handoff
  (their decision/reflect journal entries + task description) BEFORE doing
  any own research or calling i_will_plan, so the Main PM builds on the
  Board's analysis instead of duplicating it. Added a dedicated section,
  hardened workflow step 1, and added an anti-pattern.
- Add a 'Products vs Projects' section: a Product fans out to one Project
  per cell; those Projects may be the SAME repo (monorepo subtrees) or
  DIFFERENT repos (multi-repo). The Main PM coordinates across them and
  must not assume one repo or call a monorepo subtree 'a separate repo'.
  Names the Prompter monorepo case (github.com/rennf93/roboco).

Developer (roles/developer.md):
- State the exact workspace path convention
  /data/workspaces/<project-slug>/<team>/<agent-slug>/, that the cwd is
  already set there, to stay inside the own cell workspace, and to not
  probe/guess the path (ls /, find /).
- Sanctioned secret handling: env/printenv is bash-guard denied and
  reveals nothing; needed secrets arrive via the task description, else
  i_am_blocked so the PM supplies them. Added matching anti-patterns.

Tests: add tests/unit/agents/test_briefing_cluster_c4.py asserting the
composed system prompt (the text mounted into agent containers) carries
each of the above.

* fix(orchestrator): board review involves PO+HoM and notifies CEO

Cluster C5 (#2, #4): a board/coordination task was reviewed by the Product
Owner alone, and the CEO got no formal signal when the review finished —
only buried channel chatter — so the Approve & Start handoff was invisible.

#4 — Board review is now a two-reviewer gate. _handle_board_assigned_task
dispatches BOTH the Product Owner and the Head of Marketing (one-shot each),
regardless of which one holds assigned_to, and the unassigned board-routing
path delegates here instead of claiming + spawning the PO alone. Board tasks
stay pending/unassigned for the CEO's Approve & Start. The board prompt now
makes the PO+HoM pair-review model explicit (HoM owns the UX/positioning
dimension).

#2 — Once BOTH reviewers have finished (dispatched and no longer active),
the orchestrator emits exactly one formal CEO notification via
NotificationService.send_board_review_complete_notification (APPROVAL type,
ack-required, carrying related_task_id) so the handoff is an actionable
signal. One-shot per task; a notification failure clears the guard so a
later tick can retry.

To let the non-assignee board member record its review note on a task held
by the other board member, content-action ownership now exempts a board role
posting to a board/coordination task (project_id is None, product_id set).
The exemption is narrow: it does not widen ownership for any other role or
any project-backed task.

Unit tests cover both reviewers dispatched, one-shot dispatch, the CEO
notification fired exactly once when both are done (and not before), the
retry-on-failure path, the notification builder, and the board co-review
ownership exemption (allowed for board+coordination, blocked otherwise).

* fix(workspace): install dev deps post-clone + raise git commit timeout for large changesets

Cluster C6 (#10, #13, #12-investigate).

#10: per-agent workspace clones never had the project's dev dependencies
installed, so the make-quality gate (ruff/mypy/pytest for Python, the TS
toolchain for the panel) was missing and devs re-downloaded tooling per
task. WorkspaceService now runs the project's install after cloning
(`uv sync` for Python, `pnpm install`/`npm ci`/`npm install` for Node/TS,
detected by manifest/lockfile). Idempotent via a lockfile-digest marker
under .git/ so a re-entry with unchanged lockfiles is a no-op; also runs on
the healthy short-circuit so pre-existing clones get backfilled. Gated by
workspace_install_dev_deps (default on) with workspace_dep_install_timeout_seconds.

#13: the gateway commit verb timed out on the large panel changeset because
every git op used the hardcoded 30s _GIT_TIMEOUT and each call also re-walks
the tree to chown. _run_git now takes a per-call timeout override sourced
from settings (git_command_timeout_seconds default); the staging + commit
ops in commit() and create_commit() use the longer git_commit_timeout_seconds
(default 180s). httpx REST timeouts unchanged in value.

#12 (investigate only — no push, no history change): the clone base ref is
NOT hardcoded; it already comes from project.default_branch threaded through
git.get_workspace -> ensure_workspace -> _clone_repo (git clone --branch).
The stale-base problem is a deploy/process issue (GitHub master is behind the
deployed migration chain), resolvable only by pushing the chain to master.
The default_branch column is the existing configurable lever.

* fix(panel): gate Approve & Start to board coordination tasks; stop 404 storm on closed sessions

CEO gate #1 button only renders for a PENDING board coordination/fan-out
task (no project_id, has product_id) — the board-reviewed handoff that
approve_and_start accepts — instead of every PENDING board-team task.
approve_and_start requires PENDING (it re-targets to Main PM without a
status change), so the gate stays on PENDING rather than the unrelated
end-of-work awaiting_ceo_approval state.

Session/message reads now treat a 404 as terminal and never retry it: a
reaped session is gone for good, and retrying every dead session-id is
what produced the growing 404 storm on GET /api/messages. The transcript
loads once (staleTime Infinity, no focus/reconnect refetch) so closed
sessions stay viewable without re-polling.

* fix(orchestrator): role-correct respawn prompt, throttle agentless dispatch, broaden #14 guard

#19 wrong-role prompt on respawn: _get_prompt_for_agent fell through to the
developer prompt for every non-dev/doc/qa role, so a respawned PM or board
agent was told to write code and call verbs it does not own. Route by the
agent's actual role through the existing per-role prompt builders
(developer/qa/documenter/cell_pm/main_pm/product_owner/head_marketing/auditor).
Both callers benefit; _spawn_pending_dev only ever passes developer/documenter/
unknown, so its behavior is unchanged.

#19 spawn-burst: _dispatch_claimed_without_agent looped over every agentless
claimed/in_progress task and could spawn many containers in one tick. Break
after the first respawn so a restart can't trigger a burst, matching every
sibling dispatcher. The release-to-pending path spawns nothing and keeps
draining stale unknown claims.

#14 guard scope: _is_descendant_code_task only matched CODE, so a descendant
DOCUMENTATION or DESIGN task escalated to a board/advisory role was still
stranded on a role with no verb to own it. Rename to
_is_descendant_executable_task and broaden to CODE/DOCUMENTATION/DESIGN — the
cell-executed types a board role cannot own. PLANNING/RESEARCH/ADMINISTRATIVE
route to a PM, not a cell agent, and are left unchanged; root tasks are still
reviewed up the chain.

* fix(docker): add node+pnpm to orchestrator so it pre-installs frontend cell deps

* Added .github workflows

* refactor(services): extract helpers to keep install_dev_deps + developer task-type check under the xenon complexity gate

* chore(github): add launch kit — CI, GHCR release, labels, templates, funding, dependabot npm, community docs

* chore(github): bump_version — drop unused noqa, fix datetime UTC import

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-06-03 06:35:03 +02:00

1368 lines
48 KiB
Python

"""Targeted coverage for branches in roboco.services.gateway.choreographer._impl.
Each test pins one rejection-envelope branch so the larger Choreographer
verb continues to surface remediation hints rather than crash on edge
states (claim failures, start failures, missing parents, etc.).
"""
from __future__ import annotations
from datetime import UTC, datetime
from typing import Any
from unittest.mock import AsyncMock, MagicMock
from uuid import uuid4
import pytest
import structlog
from roboco.services.gateway.choreographer import Choreographer, ChoreographerDeps
from roboco.services.gateway.choreographer._impl import DelegateInputs
from roboco.services.gateway.envelope import Envelope
# #172: a developer fresh claim must carry a substantive step checklist.
# Inert on re-entry/error/non-dev paths (the gate is skipped or the call
# short-circuits before it), so it is safe to pass everywhere.
_STEPS = [
{
"title": "Implement the change",
"description": (
"edit the target file, add tests, run them, and stage the "
"change for commit on the task branch"
),
}
]
# Full parity: a fresh dev claim authors the same rich plan a PM does.
# These satisfy _dev_plan_gate (plan/approach >= 150 chars,
# technical_considerations, risks).
_GOOD_PLAN = (
"Append the timestamp HTML comment to the very bottom of README.md without "
"touching any other line, then commit it on the task branch and open a PR. "
"Verify the diff is a single-line addition before submitting for QA."
)
_GOOD_TC = ["Use a trailing newline so the comment sits on its own line."]
_GOOD_RISKS = [
{
"risk": "An accidental reformat of README.md balloons the diff.",
"mitigation": "Append only; assert the diff touches one line pre-commit.",
}
]
def _wire_dev_task_svc(
task_id, *, status: str, assigned_to=None, plan=None, parent_task_id=None
):
"""Build a TaskService AsyncMock pre-wired with claim-guard side effects.
Defaults `agent_for` → developer/backend and the three list-* methods to
empty lists so claim-guard short-circuits never fire unintentionally.
Also wires ``session.begin_nested()`` so VerbRunner's savepoint context
manager works against the mock.
"""
task_svc = AsyncMock()
task_svc.get.return_value = MagicMock(
status=status,
assigned_to=assigned_to,
plan=plan,
id=task_id,
title="t",
task_type="code",
parent_task_id=parent_task_id,
team="backend",
commits=[],
pr_number=None,
branch_name="feature/backend/abc",
quick_context=None,
)
task_svc.agent_for.return_value = MagicMock(
role="developer", team="backend", slug=None
)
task_svc.list_in_progress_for_agent.return_value = []
task_svc.list_paused_for_agent.return_value = []
task_svc.get_subtasks.return_value = []
task_svc.session = MagicMock()
task_svc.session.begin_nested = MagicMock(
return_value=MagicMock(
__aenter__=AsyncMock(return_value=None),
__aexit__=AsyncMock(return_value=False),
)
)
return task_svc
def _make_deps(**overrides: Any) -> ChoreographerDeps:
base: dict[str, Any] = {
"task": AsyncMock(),
"work_session": AsyncMock(),
"git": AsyncMock(),
"a2a": AsyncMock(),
"journal": AsyncMock(),
"audit": AsyncMock(),
"evidence_repo": AsyncMock(),
}
base.update(overrides)
# VerbRunner wraps composed atomic actions in
# ``task.session.begin_nested()``. AsyncMock auto-attribute access
# would return an unawaitable coroutine, breaking the
# ``async with`` protocol. Overwrite session with a MagicMock that
# implements the async-context-manager protocol explicitly.
task_dep = base["task"]
task_dep.session = MagicMock()
task_dep.session.begin_nested = MagicMock(
return_value=MagicMock(
__aenter__=AsyncMock(return_value=None),
__aexit__=AsyncMock(return_value=False),
)
)
repo = base["evidence_repo"]
for method in (
"list_unread_a2a",
"list_unread_mentions",
"list_pending_notifications",
"task_metadata_gaps",
"recent_team_activity",
"blockers_in_lane",
"journal_highlights_for_task",
):
getattr(repo, method).return_value = []
# C8: default-fresh journal:decision so PM-decision gate passes.
# Tests that exercise the gate boundary stub their own value.
# The check matches MagicMock and AsyncMock (the two default sentinel
# types pytest's unittest.mock leaves on un-stubbed return_values).
_ldef = base["journal"].latest_decision_at.return_value
if type(_ldef).__name__ in ("MagicMock", "AsyncMock"):
base["journal"].latest_decision_at.return_value = datetime.now(UTC)
return ChoreographerDeps(**base)
# ---------------------------------------------------------------------------
# _emit_rejection: ok envelope passes through unchanged (line 158)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_emit_rejection_passes_through_ok_envelope() -> None:
deps = _make_deps()
c = Choreographer(deps)
ok = Envelope.ok(status="x", task_id=None, next="n", context_briefing={})
result = await c._emit_rejection(ok, agent_id=uuid4(), task_id=None, verb="x")
assert result is ok
deps.audit.log_event.assert_not_called()
# ---------------------------------------------------------------------------
# i_will_work_on: claim() raises Exception → invalid_state (lines 369-378)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_i_will_work_on_pending_claim_raises_returns_invalid_state() -> None:
"""When the runner re-raises a RuntimeError from claim(), the verb body
catches it and surfaces an invalid_state envelope with the runner's
message. Pre-spec the verb body produced "claim failed during
finalization"; the spec-driven body produces "verb runner failed:
<exc>" so the agent still gets a remediation hint instead of a 500.
"""
agent_id = uuid4()
task_id = uuid4()
task_svc = _wire_dev_task_svc(task_id, status="pending")
task_svc.claim.side_effect = RuntimeError("workspace down")
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_work_on(
agent_id,
task_id,
plan=_GOOD_PLAN,
steps=_STEPS,
technical_considerations=_GOOD_TC,
risks=_GOOD_RISKS,
)
body = env.as_dict()
assert body["error"] == "invalid_state"
assert "verb runner failed" in body["message"]
assert "workspace down" in body["message"]
@pytest.mark.asyncio
async def test_i_will_work_on_pending_claim_returns_none_invalid_state() -> None:
"""Lines 379-384: claim returns None → invalid_state."""
agent_id = uuid4()
task_id = uuid4()
task_svc = _wire_dev_task_svc(task_id, status="pending")
task_svc.claim.return_value = None
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_work_on(
agent_id,
task_id,
plan=_GOOD_PLAN,
steps=_STEPS,
technical_considerations=_GOOD_TC,
risks=_GOOD_RISKS,
)
body = env.as_dict()
assert body["error"] == "invalid_state"
@pytest.mark.asyncio
async def test_i_will_work_on_pending_no_plan_tracing_gap() -> None:
"""Lines 385-393: pending task, no plan → tracing_gap."""
agent_id = uuid4()
task_id = uuid4()
task_svc = _wire_dev_task_svc(task_id, status="pending", assigned_to=agent_id)
claimed_task = MagicMock(
status="pending",
assigned_to=agent_id,
plan=None,
id=task_id,
title="t",
task_type="code",
)
task_svc.claim.return_value = claimed_task
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_work_on(agent_id, task_id, plan=None, steps=_STEPS)
body = env.as_dict()
assert body["error"] == "tracing_gap"
@pytest.mark.asyncio
async def test_i_will_work_on_start_returns_none_invalid_state() -> None:
"""Lines 396-398: start returns None → start_failed_envelope."""
agent_id = uuid4()
task_id = uuid4()
task_svc = _wire_dev_task_svc(task_id, status="pending", assigned_to=agent_id)
claimed_task = MagicMock(
status="pending",
assigned_to=agent_id,
plan="some plan",
id=task_id,
title="t",
task_type="code",
)
task_svc.claim.return_value = claimed_task
task_svc.set_plan.return_value = claimed_task
task_svc.start.return_value = None # start fails
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_work_on(
agent_id,
task_id,
plan=_GOOD_PLAN,
steps=_STEPS,
technical_considerations=_GOOD_TC,
risks=_GOOD_RISKS,
)
body = env.as_dict()
assert body["error"] == "invalid_state"
assert "start failed" in body["message"]
# ---------------------------------------------------------------------------
# _i_will_work_on_needs_revision: claim returns None → invalid_state (lines 427-433)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_needs_revision_branch_claim_fails_invalid_state() -> None:
"""Lines 427-433: needs_revision, not assigned, claim fails → invalid_state."""
agent_id = uuid4()
task_id = uuid4()
other_id = uuid4()
task_svc = _wire_dev_task_svc(
task_id, status="needs_revision", assigned_to=other_id, plan="p"
)
task_svc.claim.return_value = None
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_work_on(
agent_id,
task_id,
plan=_GOOD_PLAN,
steps=_STEPS,
technical_considerations=_GOOD_TC,
risks=_GOOD_RISKS,
)
body = env.as_dict()
assert body["error"] == "invalid_state"
@pytest.mark.asyncio
async def test_needs_revision_branch_start_fails() -> None:
"""Line 436: start returns None in needs_revision branch."""
agent_id = uuid4()
task_id = uuid4()
task_svc = _wire_dev_task_svc(
task_id, status="needs_revision", assigned_to=agent_id, plan="p"
)
task_svc.start.return_value = None
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_work_on(
agent_id,
task_id,
plan=_GOOD_PLAN,
steps=_STEPS,
technical_considerations=_GOOD_TC,
risks=_GOOD_RISKS,
)
body = env.as_dict()
assert body["error"] == "invalid_state"
# ---------------------------------------------------------------------------
# _i_will_work_on_claimed: start fails → start_failed_envelope (line 454)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_claimed_branch_returns_start_failed() -> None:
"""Line 454: start fails in claimed branch."""
agent_id = uuid4()
task_id = uuid4()
task_svc = _wire_dev_task_svc(
task_id, status="claimed", assigned_to=agent_id, plan="p"
)
task_svc.start.return_value = None
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_work_on(agent_id, task_id, plan="ok", steps=_STEPS)
body = env.as_dict()
assert body["error"] == "invalid_state"
# ---------------------------------------------------------------------------
# i_will_work_on with in_progress assigned to self → idempotent (line 491)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_i_will_work_on_in_progress_assigned_to_self_idempotent() -> None:
"""Line 491: in_progress assigned_to=agent → idempotent re-entry pass through."""
agent_id = uuid4()
task_id = uuid4()
task_svc = _wire_dev_task_svc(
task_id, status="in_progress", assigned_to=agent_id, plan="p"
)
task_svc.heartbeat = AsyncMock()
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_work_on(agent_id, task_id, plan="ok", steps=_STEPS)
body = env.as_dict()
# No error — re-entry pass.
assert "error" not in body or body.get("error") is None
# ---------------------------------------------------------------------------
# i_will_plan: pending claim returns None (line 1155)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_i_will_plan_pm_with_already_active_task_rejects() -> None:
"""The already_active_guard still fires on i_will_plan even though
pm_cannot_execute_code is skipped. Covers _impl.py:1106-1108
(with-briefing wrap of the guard rejection).
"""
pm_id = uuid4()
task_id = uuid4()
other_task_id = uuid4()
target = MagicMock(
status="pending",
assigned_to=pm_id,
plan=None,
id=task_id,
title="t",
team="backend",
parent_task_id=None,
task_type="planning",
)
busy_task = MagicMock(id=other_task_id, status="in_progress")
task_svc = AsyncMock()
task_svc.get.return_value = target
task_svc.agent_for.return_value = MagicMock(role="cell_pm", team="backend")
task_svc.list_in_progress_for_agent.return_value = [busy_task]
task_svc.list_paused_for_agent.return_value = []
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_plan(
pm_id,
task_id,
plan="x" * 30,
rich_plan={
"approach": (
"Decompose the planning task into backend and frontend "
"developer-claimable subtasks. Backend lands first, QA "
"reviews each PR after it opens, documentation follows, then "
"complete and submit up. Strict sequencing with no cross-cell "
"dependencies beyond the stated ordering."
),
"sub_tasks": [
{
"title": "Slice A",
"description": (
"be-dev-1 implements the backend API change with "
"tests and opens the leaf PR for QA review."
),
}
],
},
)
body = env.as_dict()
assert body["error"] == "invalid_state"
assert "in_progress task" in body["message"]
@pytest.mark.asyncio
async def test_i_will_plan_cell_pm_on_code_typed_parent_succeeds() -> None:
"""Regression for the smoke-test deadlock (2026-05-08 trace).
When a cell PM tries to plan a code-typed parent task, the verb must
succeed — PMs PLAN code work and DELEGATE the execution; they don't
execute. The pre-fix `pm_cannot_execute_code_guard` was wrongly fired
on `i_will_plan` (the planning verb) instead of being scoped to
`i_will_work_on` (the execution verb), causing a deadlock: cell PM
couldn't plan → couldn't transition parent to in_progress → couldn't
delegate (delegate requires parent in_progress).
"""
pm_id = uuid4()
task_id = uuid4()
task = MagicMock(
status="pending",
assigned_to=pm_id,
plan=None,
id=task_id,
title="Backend slice: Git workflow smoke test",
team="backend",
parent_task_id=uuid4(), # subtask of the main_pm root
task_type="code", # ← the trigger; pre-fix this rejected with
# "Cell Pm cannot claim code tasks"
)
task_svc = AsyncMock()
task_svc.get.return_value = task
task_svc.agent_for.return_value = MagicMock(role="cell_pm", team="backend")
task_svc.list_in_progress_for_agent.return_value = []
task_svc.list_paused_for_agent.return_value = []
task_svc.get_subtasks.return_value = []
# Claim + start succeed so we can verify the verb runs end-to-end.
started_task = MagicMock(
status="in_progress",
assigned_to=pm_id,
id=task_id,
title=task.title,
team="backend",
task_type="code",
)
task_svc.claim.return_value = task
task_svc.start.return_value = started_task
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_plan(
pm_id,
task_id,
plan="Decompose into 2 dev subtasks.",
rich_plan={
"approach": (
"Split code-typed parent into developer-claimable subtasks: "
"one for API, one for test coverage validation."
),
"sub_tasks": [
{"title": "API subtask", "description": "Implement the endpoint"},
],
},
)
body = env.as_dict()
# The PM-cannot-execute-code rejection must NOT fire on i_will_plan.
assert body.get("error") != "not_authorized", (
f"i_will_plan was rejected for a code-typed parent; envelope: {body}"
)
@pytest.mark.asyncio
async def test_i_will_plan_pending_claim_fails() -> None:
pm_id = uuid4()
task_id = uuid4()
task = MagicMock(
status="pending",
assigned_to=pm_id,
plan=None,
id=task_id,
title="t",
team="backend",
parent_task_id=None,
task_type="planning",
)
task_svc = AsyncMock()
task_svc.get.return_value = task
task_svc.agent_for.return_value = MagicMock(role="cell_pm", team="backend")
task_svc.list_in_progress_for_agent.return_value = []
task_svc.list_paused_for_agent.return_value = []
task_svc.get_subtasks.return_value = []
task_svc.claim.return_value = None # claim fails
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_plan(
pm_id,
task_id,
plan="my plan that is long enough",
rich_plan={
"approach": (
"Decompose the planning task into backend and frontend "
"developer-claimable subtasks. Backend lands first, QA "
"reviews each PR after it opens, documentation follows, then "
"complete and submit up. Strict sequencing with no cross-cell "
"dependencies beyond the stated ordering."
),
"sub_tasks": [
{
"title": "Slice A",
"description": (
"be-dev-1 implements the backend API change with "
"tests and opens the leaf PR for QA review."
),
}
],
},
)
body = env.as_dict()
assert body["error"] == "invalid_state"
# ---------------------------------------------------------------------------
# delegate: parent not found (line 1208)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_delegate_parent_not_found() -> None:
pm_id = uuid4()
parent_id = uuid4()
task_svc = AsyncMock()
task_svc.get.return_value = None
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.delegate(
pm_id,
parent_id,
DelegateInputs(
title="Implement endpoint",
description="Add /v1/foo endpoint with passing tests please",
assigned_to="be-dev-1",
team="backend",
task_type="code",
nature="technical",
acceptance_criteria=["GET /v1/foo returns 200 with body"],
),
)
body = env.as_dict()
assert body["error"] == "not_found"
# ---------------------------------------------------------------------------
# _delegate_role_guards: unknown role rejection (line 1271)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_delegate_unknown_role_rejected() -> None:
pm_id = uuid4()
parent_id = uuid4()
parent = MagicMock(
status="in_progress",
assigned_to=pm_id,
project_id=uuid4(),
title="p",
)
task_svc = AsyncMock()
task_svc.get.return_value = parent
task_svc.agent_for.return_value = MagicMock(role="developer", team="backend")
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.delegate(
pm_id,
parent_id,
DelegateInputs(
title="Implement endpoint",
description="Add /v1/foo endpoint with passing tests please",
assigned_to="be-dev-1",
team="backend",
task_type="code",
nature="technical",
acceptance_criteria=["GET /v1/foo returns 200 with body"],
),
)
body = env.as_dict()
assert body["error"] == "not_authorized"
# ---------------------------------------------------------------------------
# _delegate_static_guards: unknown agent slug → invalid_state (line 1300)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_delegate_parent_no_project_rejected() -> None:
"""A parent with NEITHER a project_id NOR a product_id → invalid_state.
(A parent with a product_id but no project is allowed — subtasks resolve a
repo from the product map — so the guard now requires both to be None.)
"""
pm_id = uuid4()
parent_id = uuid4()
parent = MagicMock(
status="in_progress",
assigned_to=pm_id,
project_id=None,
product_id=None,
title="p",
)
task_svc = AsyncMock()
task_svc.get.return_value = parent
task_svc.agent_for.return_value = MagicMock(role="cell_pm", team="backend")
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.delegate(
pm_id,
parent_id,
DelegateInputs(
title="Implement endpoint",
description="Add /v1/foo endpoint with passing tests please",
assigned_to="be-dev-1",
team="backend",
task_type="code",
nature="technical",
acceptance_criteria=["GET /v1/foo returns 200 with body"],
),
)
body = env.as_dict()
assert body["error"] == "invalid_state"
assert "project_id" in body["message"] and "product_id" in body["message"]
# ---------------------------------------------------------------------------
# _validate_delegation_chain: unknown role → "role X cannot delegate" (line 1446)
# ---------------------------------------------------------------------------
def test_validate_delegation_chain_unknown_role() -> None:
deps = _make_deps()
c = Choreographer(deps)
err = c._validate_delegation_chain("auditor", "be-dev-1")
assert err is not None
assert "auditor" in err
# ---------------------------------------------------------------------------
# submit_up: task not found (line 1458)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_submit_up_task_not_found() -> None:
pm_id = uuid4()
task_id = uuid4()
task_svc = AsyncMock()
task_svc.get.return_value = None
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.submit_up(pm_id, task_id, notes="x" * 30)
body = env.as_dict()
assert body["error"] == "not_found"
# ---------------------------------------------------------------------------
# submit_up: submit_pm_review returns None (line 1477)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_submit_up_submit_pm_review_fails() -> None:
pm_id = uuid4()
task_id = uuid4()
task = MagicMock(
status="in_progress",
assigned_to=pm_id,
branch_name="feature/backend/abc",
pr_number=None,
title="t",
team="backend",
)
task_svc = AsyncMock()
task_svc.get.return_value = task
task_svc.agent_for.return_value = MagicMock(role="cell_pm", team="backend")
task_svc.all_subtasks_terminal.return_value = True
task_svc.submit_pm_review.return_value = None # service returns None
journal = AsyncMock()
journal.has_decision_for_task.return_value = True
journal.latest_decision_at.return_value = datetime.now(UTC)
git = AsyncMock()
git.create_pr = AsyncMock()
deps = _make_deps(task=task_svc, journal=journal, git=git)
c = Choreographer(deps)
env = await c.submit_up(pm_id, task_id, notes="x" * 30)
body = env.as_dict()
assert body["error"] == "invalid_state"
# ---------------------------------------------------------------------------
# _submit_up_ownership_guard wrong role (line 1520)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_submit_up_wrong_role_rejected() -> None:
pm_id = uuid4()
task_id = uuid4()
task = MagicMock(
status="in_progress",
assigned_to=pm_id,
title="t",
team="backend",
)
task_svc = AsyncMock()
task_svc.get.return_value = task
task_svc.agent_for.return_value = MagicMock(role="main_pm", team=None)
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.submit_up(pm_id, task_id, notes="x" * 30)
body = env.as_dict()
assert body["error"] == "not_authorized"
# ---------------------------------------------------------------------------
# _submit_up_state_guard: no branch_name (line 1556)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_submit_up_no_branch_rejected() -> None:
pm_id = uuid4()
task_id = uuid4()
task = MagicMock(
status="in_progress",
assigned_to=pm_id,
branch_name=None,
pr_number=None,
title="t",
team="backend",
)
task_svc = AsyncMock()
task_svc.get.return_value = task
task_svc.agent_for.return_value = MagicMock(role="cell_pm", team="backend")
task_svc.all_subtasks_terminal.return_value = True
journal = AsyncMock()
journal.has_decision_for_task.return_value = True
journal.latest_decision_at.return_value = datetime.now(UTC)
deps = _make_deps(task=task_svc, journal=journal)
c = Choreographer(deps)
env = await c.submit_up(pm_id, task_id, notes="x" * 30)
body = env.as_dict()
assert body["error"] == "invalid_state"
assert "no branch" in body["message"]
# ---------------------------------------------------------------------------
# _pm_next_hint: each branch (lines 1612-1616)
# ---------------------------------------------------------------------------
def test_pm_next_hint_pending() -> None:
deps = _make_deps()
c = Choreographer(deps)
hint = c._pm_next_hint("pending", "tid")
assert "i_will_plan" in hint
def test_pm_next_hint_paused() -> None:
deps = _make_deps()
c = Choreographer(deps)
hint = c._pm_next_hint("paused", "tid")
assert "subtasks" in hint or "complete" in hint
def test_pm_next_hint_blocked() -> None:
deps = _make_deps()
c = Choreographer(deps)
hint = c._pm_next_hint("blocked", "tid")
assert "unblock" in hint
def test_pm_next_hint_awaiting_pm_review() -> None:
deps = _make_deps()
c = Choreographer(deps)
hint = c._pm_next_hint("awaiting_pm_review", "tid")
assert "complete" in hint
def test_pm_next_hint_unknown_status() -> None:
deps = _make_deps()
c = Choreographer(deps)
hint = c._pm_next_hint("unknown_status", "tid")
assert "unknown_status" in hint
# ---------------------------------------------------------------------------
# triage_all main_pm — awaiting Main PM tasks branch (lines 1662-1663)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_triage_all_returns_awaiting_main_pm_when_no_blocked() -> None:
pm_id = uuid4()
awaiting_task = MagicMock(
id=uuid4(), status="awaiting_pm_review", title="x", team="backend"
)
task_svc = AsyncMock()
task_svc.list_blocked_all_teams.return_value = []
task_svc.list_awaiting_main_pm_all.return_value = [awaiting_task]
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.triage_all(pm_id)
body = env.as_dict()
assert body["task_id"] == str(awaiting_task.id)
assert "complete" in body["next"]
# ---------------------------------------------------------------------------
# unblock: task not found (line 1682)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_unblock_task_not_found() -> None:
pm_id = uuid4()
task_id = uuid4()
task_svc = AsyncMock()
task_svc.get.return_value = None
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.unblock(pm_id, task_id)
body = env.as_dict()
assert body["error"] == "not_found"
# ---------------------------------------------------------------------------
# _cell_pm_complete_guard: wrong status (line 1743)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_cell_pm_complete_wrong_status() -> None:
"""Line 1743: status not awaiting_pm_review."""
pm_id = uuid4()
task_id = uuid4()
task = MagicMock(
status="in_progress",
assigned_to=pm_id,
title="t",
team="backend",
)
task_svc = AsyncMock()
task_svc.get.return_value = task
task_svc.agent_for.return_value = MagicMock(role="cell_pm", team="backend")
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.cell_pm_complete(pm_id, task_id, notes="x" * 30)
body = env.as_dict()
assert body["error"] == "invalid_state"
# ---------------------------------------------------------------------------
# cell_pm_complete: task not found (line 1782)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_cell_pm_complete_not_found() -> None:
pm_id = uuid4()
task_id = uuid4()
task_svc = AsyncMock()
task_svc.get.return_value = None
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.cell_pm_complete(pm_id, task_id, notes="x" * 30)
body = env.as_dict()
assert body["error"] == "not_found"
# ---------------------------------------------------------------------------
# _maybe_advance_parent_to_pm_review: silent skips (lines 1834, 1840, 1843)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_maybe_advance_parent_skips_when_parent_missing() -> None:
parent_id = uuid4()
task_svc = AsyncMock()
task_svc.get.return_value = None
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
await c._maybe_advance_parent_to_pm_review(parent_id, "backend")
task_svc.reassign.assert_not_called()
@pytest.mark.asyncio
async def test_maybe_advance_parent_skips_when_subtasks_not_terminal() -> None:
parent_id = uuid4()
task_svc = AsyncMock()
task_svc.get.return_value = MagicMock(team="backend")
task_svc.all_subtasks_terminal.return_value = False
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
await c._maybe_advance_parent_to_pm_review(parent_id, "backend")
task_svc.reassign.assert_not_called()
@pytest.mark.asyncio
async def test_maybe_advance_parent_skips_when_no_team() -> None:
parent_id = uuid4()
task_svc = AsyncMock()
task_svc.get.return_value = MagicMock(team=None)
task_svc.all_subtasks_terminal.return_value = True
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
# leaf_team also None → triggers line 1840 short-circuit.
await c._maybe_advance_parent_to_pm_review(parent_id, None)
task_svc.reassign.assert_not_called()
@pytest.mark.asyncio
async def test_maybe_advance_parent_skips_when_no_pm_for_team() -> None:
parent_id = uuid4()
task_svc = AsyncMock()
task_svc.get.return_value = MagicMock(team="backend")
task_svc.all_subtasks_terminal.return_value = True
task_svc.cell_pm_for_team.return_value = None
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
await c._maybe_advance_parent_to_pm_review(parent_id, "backend")
task_svc.reassign.assert_not_called()
# ---------------------------------------------------------------------------
# _main_pm_complete_guard: not assigned (1851), wrong status (1859)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_main_pm_complete_not_assigned() -> None:
main_pm_id = uuid4()
other_id = uuid4()
task_id = uuid4()
task = MagicMock(
status="awaiting_pm_review",
assigned_to=other_id,
parent_task_id=None,
title="t",
)
task_svc = AsyncMock()
task_svc.get.return_value = task
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.main_pm_complete(main_pm_id, task_id, notes="x" * 30)
body = env.as_dict()
assert body["error"] == "not_authorized"
@pytest.mark.asyncio
async def test_main_pm_complete_wrong_status() -> None:
# #183: in_progress is now an accepted source (root resumed from paused);
# use paused — a genuinely non-completable status — to exercise the guard.
main_pm_id = uuid4()
task_id = uuid4()
task = MagicMock(
status="paused",
assigned_to=main_pm_id,
parent_task_id=None,
title="t",
)
task_svc = AsyncMock()
task_svc.get.return_value = task
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.main_pm_complete(main_pm_id, task_id, notes="x" * 30)
body = env.as_dict()
assert body["error"] == "invalid_state"
# ---------------------------------------------------------------------------
# _main_pm_complete_guard: missing decision journal (lines 1885-1889)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_main_pm_complete_missing_journal_decision() -> None:
main_pm_id = uuid4()
task_id = uuid4()
task = MagicMock(
status="awaiting_pm_review",
assigned_to=main_pm_id,
parent_task_id=None,
title="t",
)
task_svc = AsyncMock()
task_svc.get.return_value = task
journal = AsyncMock()
journal.has_decision_for_task.return_value = False
journal.latest_decision_at.return_value = None
deps = _make_deps(task=task_svc, journal=journal)
c = Choreographer(deps)
env = await c.main_pm_complete(main_pm_id, task_id, notes="x" * 30)
body = env.as_dict()
assert body["error"] == "tracing_gap"
# ---------------------------------------------------------------------------
# main_pm_complete: not_found (line 1908)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_main_pm_complete_not_found() -> None:
main_pm_id = uuid4()
task_id = uuid4()
task_svc = AsyncMock()
task_svc.get.return_value = None
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.main_pm_complete(main_pm_id, task_id, notes="x" * 30)
body = env.as_dict()
assert body["error"] == "not_found"
# ---------------------------------------------------------------------------
# _emit_rejection: correlation_id from contextvars (line 170)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_emit_rejection_includes_correlation_id() -> None:
"""When structlog contextvars holds a correlation_id, it gets stamped."""
deps = _make_deps()
c = Choreographer(deps)
rejection = Envelope.invalid_state(message="m", remediate="r", context_briefing={})
structlog.contextvars.bind_contextvars(correlation_id="cid-123")
try:
await c._emit_rejection(rejection, agent_id=uuid4(), task_id=None, verb="x")
finally:
structlog.contextvars.unbind_contextvars("correlation_id")
deps.audit.log_event.assert_awaited()
call_kwargs = deps.audit.log_event.await_args.kwargs
assert call_kwargs["details"]["correlation_id"] == "cid-123"
# ---------------------------------------------------------------------------
# _i_will_work_on_claimed: guard returns (line 451) — already-active blocker
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_claimed_branch_already_active_guard() -> None:
"""Line 451: in_progress task elsewhere blocks claim of another."""
agent_id = uuid4()
task_id = uuid4()
other_id = uuid4()
in_prog = MagicMock(id=other_id, status="in_progress", title="other")
task_svc = _wire_dev_task_svc(
task_id, status="claimed", assigned_to=agent_id, plan="p"
)
task_svc.list_in_progress_for_agent.return_value = [in_prog]
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_work_on(agent_id, task_id, plan="ok", steps=_STEPS)
body = env.as_dict()
assert body["error"] == "invalid_state"
# ---------------------------------------------------------------------------
# Pure helper: skill matching string entries (lines 802-803)
# ---------------------------------------------------------------------------
def test_resolve_skill_string_entries() -> None:
deps = _make_deps()
c = Choreographer(deps)
agent = MagicMock(skills=["python", "rust"], capabilities=None)
result = c._resolve_skill(agent, ["go", "python"])
assert result == "python"
# ---------------------------------------------------------------------------
# i_will_plan: claim returns None for pending task (line 1155 — emit_rejection)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_i_will_plan_pending_claim_returns_none_emit_rejection() -> None:
"""When claim() returns None inside the runner, the savepoint rolls
back and the runner-failure path surfaces as invalid_state. Pre-spec
this branched into a hand-rolled "claim failed" message; now it is
emitted via _claim_plan_start_run's exception handler.
"""
pm_id = uuid4()
task_id = uuid4()
task = MagicMock(
status="pending",
assigned_to=pm_id,
plan=None,
id=task_id,
title="t",
team="backend",
parent_task_id=None,
task_type="planning",
quick_context=None,
)
task_svc = AsyncMock()
task_svc.get.return_value = task
task_svc.agent_for.return_value = MagicMock(
id=pm_id, role="cell_pm", team="backend", slug=None
)
task_svc.list_in_progress_for_agent.return_value = []
task_svc.list_paused_for_agent.return_value = []
task_svc.get_subtasks.return_value = []
task_svc.claim.return_value = None
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_plan(
pm_id,
task_id,
plan="my plan that is long enough",
rich_plan={
"approach": (
"Decompose the planning task into backend and frontend "
"developer-claimable subtasks. Backend lands first, QA "
"reviews each PR after it opens, documentation follows, then "
"complete and submit up. Strict sequencing with no cross-cell "
"dependencies beyond the stated ordering."
),
"sub_tasks": [
{
"title": "Slice A",
"description": (
"be-dev-1 implements the backend API change with "
"tests and opens the leaf PR for QA review."
),
}
],
},
)
body = env.as_dict()
assert body["error"] == "invalid_state"
assert "verb runner failed" in body["message"]
# ---------------------------------------------------------------------------
# _submit_up_ownership_guard: not_assigned (line 1520)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_submit_up_not_assigned_rejected() -> None:
"""Line 1520: cell_pm calling submit_up but task assigned to another agent."""
pm_id = uuid4()
task_id = uuid4()
other_id = uuid4()
task = MagicMock(
status="in_progress",
assigned_to=other_id,
title="t",
team="backend",
)
task_svc = AsyncMock()
task_svc.get.return_value = task
task_svc.agent_for.return_value = MagicMock(role="cell_pm", team="backend")
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.submit_up(pm_id, task_id, notes="x" * 30)
body = env.as_dict()
assert body["error"] == "not_authorized"
assert "not assigned" in body["message"]
# ---------------------------------------------------------------------------
# Task 3: Envelope introspection — verb returns carry current_state +
# valid_next_verbs so agents stop trial-and-erroring.
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_i_will_work_on_envelope_carries_introspection_on_success() -> None:
"""Successful claim+start path stamps current_state + valid_next_verbs."""
agent_id = uuid4()
task_id = uuid4()
task_svc = _wire_dev_task_svc(task_id, status="pending", assigned_to=agent_id)
claimed_task = MagicMock(
status="in_progress",
assigned_to=agent_id,
plan="ok plan",
id=task_id,
title="t",
task_type="code",
)
task_svc.claim.return_value = claimed_task
task_svc.set_plan.return_value = claimed_task
task_svc.start.return_value = claimed_task
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_work_on(
agent_id,
task_id,
plan=_GOOD_PLAN,
steps=_STEPS,
technical_considerations=_GOOD_TC,
risks=_GOOD_RISKS,
)
body = env.as_dict()
assert body["error"] is None
assert body["current_state"] == "in_progress"
assert isinstance(body["valid_next_verbs"], list)
# `valid_next_verbs` lists lifecycle INTENT verbs; `commit` is a
# content tool (do_server), not an intent, so the canonical spec
# excludes it. `open_pr` and `i_am_done` are the in_progress intents.
assert "open_pr" in body["valid_next_verbs"]
assert "i_am_done" in body["valid_next_verbs"]
@pytest.mark.asyncio
async def test_i_will_work_on_envelope_carries_introspection_on_rejection() -> None:
"""A wrong-state rejection still stamps current_state + valid_next_verbs
so the agent learns what verbs are actually valid right now."""
agent_id = uuid4()
task_id = uuid4()
task_svc = _wire_dev_task_svc(task_id, status="completed", assigned_to=agent_id)
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_work_on(agent_id, task_id, plan="x", steps=_STEPS)
body = env.as_dict()
assert body["error"] == "invalid_state"
assert body["current_state"] == "completed"
assert isinstance(body["valid_next_verbs"], list)
# Lifecycle verbs are NOT in the list for a completed task.
assert "i_will_work_on" not in body["valid_next_verbs"]
@pytest.mark.asyncio
async def test_open_pr_does_not_create_pr_if_no_commits() -> None:
"""Atomic invariant: if commits[] is empty, open_pr must NOT call
git.create_pr. Pre-fix this was already true at the verb level, but
this test pins it as a regression: any future refactor that
re-orders precondition vs side effect breaks the test."""
dev_id = uuid4()
task_id = uuid4()
task = MagicMock(
status="in_progress",
assigned_to=dev_id,
commits=[],
pr_number=None,
branch_name="feature/backend/abc",
id=task_id,
title="t",
team="backend",
task_type="code",
)
task_svc = AsyncMock()
task_svc.get.return_value = task
task_svc.agent_for.return_value = MagicMock(role="developer", team="backend")
git_svc = AsyncMock()
git_svc.create_pr = AsyncMock()
git_svc.push_branch = AsyncMock()
deps = _make_deps(task=task_svc, git=git_svc)
c = Choreographer(deps)
env = await c.open_pr(dev_id, task_id)
body = env.as_dict()
# Spec's PRECONDITION_COMMITS now produces tracing_gap rather than the
# previous bespoke invalid_state. The atomicity invariant the test
# pins (no git side effect when commits=[]) is unchanged.
assert body["error"] == "tracing_gap"
assert body["missing"] == ["commits>=1"]
git_svc.create_pr.assert_not_called()
git_svc.push_branch.assert_not_called()
@pytest.mark.asyncio
async def test_i_will_work_on_missing_plan_does_not_claim_pending_task() -> None:
"""Atomic invariant (Task 5 pattern, Bug A from 2026-05-09 smoke):
if `plan` is missing on the FIRST i_will_work_on call against a
pending task, the task must NOT be claimed. Pre-fix the verb ran
claim() BEFORE checking plan, leaving the task in `claimed` with
no plan — and `_i_will_work_on_claimed` had no recovery path so
the agent looped forever on `start failed`.
"""
agent_id = uuid4()
task_id = uuid4()
task_svc = _wire_dev_task_svc(task_id, status="pending")
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_work_on(agent_id, task_id, plan=None, steps=_STEPS)
body = env.as_dict()
assert body["error"] == "tracing_gap"
assert "plan" in body["missing"]
(
task_svc.claim.assert_not_called(),
("claim() ran before plan precondition was satisfied — atomicity broken"),
)
@pytest.mark.asyncio
async def test_i_will_work_on_claimed_with_no_plan_accepts_recovery_plan() -> None:
"""Recovery path (Bug A from 2026-05-09 smoke): if the task is in
`claimed` state without a plan (e.g. from a prior partial-claim race
or an orchestrator restart), a fresh i_will_work_on call WITH plan
must set the plan and then start, not just call start() against a
plan-less task.
"""
agent_id = uuid4()
task_id = uuid4()
task_svc = _wire_dev_task_svc(
task_id, status="claimed", assigned_to=agent_id, plan=None
)
started = MagicMock(
status="in_progress",
assigned_to=agent_id,
plan="recovery plan",
id=task_id,
title="t",
task_type="code",
)
task_svc.set_plan.return_value = started
task_svc.start.return_value = started
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_work_on(agent_id, task_id, plan="recovery plan", steps=_STEPS)
body = env.as_dict()
assert body["error"] is None, f"expected success, got {body}"
task_svc.set_plan.assert_awaited_once()
# ---------------------------------------------------------------------------
# _pending_assignment_guard: board/advisory roles can idle without claiming
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
@pytest.mark.parametrize("role", ["product_owner", "head_marketing", "auditor"])
async def test_pending_assignment_guard_exempts_board_roles(role: str) -> None:
"""A board/advisory agent that reviewed a still-pending coordination task
has no i_will_work_on/i_will_plan verb, so the idle gate must let it pass."""
task_svc = AsyncMock()
task_svc.list_assigned_for_agent.return_value = [
MagicMock(id=uuid4(), status="pending")
]
task_svc.agent_for.return_value = MagicMock(role=role, team=None, slug=None)
c = Choreographer(_make_deps(task=task_svc))
assert await c._pending_assignment_guard(uuid4(), {}) is None
@pytest.mark.asyncio
async def test_pending_assignment_guard_still_blocks_developer() -> None:
"""A developer holding a pending unclaimed task is still told to claim it."""
task_svc = AsyncMock()
task_svc.list_assigned_for_agent.return_value = [
MagicMock(id=uuid4(), status="pending")
]
task_svc.agent_for.return_value = MagicMock(
role="developer", team="backend", slug=None
)
c = Choreographer(_make_deps(task=task_svc))
guard = await c._pending_assignment_guard(uuid4(), {})
assert guard is not None
assert guard.as_dict()["error"] == "invalid_state"