Chore: v1 removal gateway canonical (#46)

* chore(agent_sdk): remove dead /traceability/remind endpoint and reminder map

The TRACEABILITY_REMINDERS dict and its /traceability/remind endpoint were
keyed entirely on pre-gateway tool names (roboco_task_*, roboco_journal_*,
roboco_message_send, roboco_session_create_for_tasks) deleted in the gateway
cutover. The endpoint had zero callers; v2 enforces traceability server-side
in the Choreographer.

* fix(bootstrap,seeds): onboarding prompts call give_me_work(), not deleted roboco_task_scan()

The startup prompt and the seeded cell/all-hands channel onboarding messages
instructed agents to call roboco_task_scan() — a tool removed in the gateway
cutover. Point them at the live give_me_work() flow verb.

* fix: replace remaining deleted v1 tool names with gateway verbs

Spawn prompts, onboarding strings, remediation messages, and comments still
referenced pre-gateway tools deleted in the cutover (roboco_task_*,
roboco_agent_idle, roboco_notify_*, roboco_message_send,
roboco_session_create_for_tasks, roboco_journal_*, roboco_escalate). Rewrote
each to the correct role-scoped gateway verb (give_me_work/i_will_work_on for
workers, triage for PMs, i_am_done vs complete, notify/notify_ack, escalate_up,
unclaim, i_documented, open_session, note). Updated one enforcement-message
test that matched the old tool name by coincidence.

* test: guard against deleted v1 tool names reappearing in roboco/

Scans roboco/ for the deleted pre-gateway tool names; excludes the orphaned
roboco/agents/ subtree (removed in a later phase).

* chore(exceptions): drop 8 unused pre-gateway exception classes + their tests

LLMError, RAGError, AlreadyExistsError, TaskBlockedError, TaskClaimError,
AgentNotAvailableError, AgentBusyError, NotificationPermissionError were never
raised in production. SessionClosedError/DatabaseError are kept (live + tested).

* chore(models): drop unused pre-gateway notification/channel/handoff factories

Removes create_task_assignment/_blocker_escalation/_review_request/
_documentation_request/_priority_change/_alert/_broadcast, create_cell_channel/
_cross_cell_channel/_announcements_channel, create_handoff (+ HandoffParams),
ProactiveContext, and A2APartType. The gateway choreographer builds these
server-side now. Drops the matching dead-code tests.

* chore(services): drop unused pre-gateway permission/messaging/audit/optimal/remediation methods

These pre-gateway helpers (channel-permission checks, channel-membership ops,
permission-denial audit hooks, doc ingestion, two remediation hints) have no
production caller — the gateway role_config + enforcement layer replaced them.
Drops the matching dead-code tests; live methods (send_message, the SESSION_*
flow, log_task_action_denial, etc.) are untouched.

* chore(orchestrator,ws,events,config): drop unused pre-gateway lifecycle/broadcast/roster symbols

orchestrator: get_running_agents, is_agent_busy, queue_priority_work,
get_all_instances (+ their OrchestratorAccessProtocol declarations in events.py).
websocket: broadcast_new_message, broadcast_session_closed (no event type emits
them). agents_config: ALL_PMS/ALL_DEVS/ALL_QA/CELL_PMS roster constants (ALL_DOCS
stays — it gates docs-write workspace perms).

* refactor(agents): delete orphaned pre-gateway agent subtree + dead organization model

The Gateway/full cutover replaced the Python agent-class implementations with
the server-side Choreographer; the classes survived only as a self-referential
island. Removes roboco/agents/{base,mixins,factory,board,developer,documenter,
pm,qa,orchestrator}.py and roboco/agents/factories/{board,cells,developers,
documenters,pms,qa}.py, plus roboco/models/organization.py (Cell/Board/
Organization — used only by those factories). Keeps factories/_base.py
(compose_prompt — the live prompt-layer composer the orchestrator calls at
spawn) behind minimal package __init__ files.

* chore(db): drop dead tasks.execution_log + outputs columns (migration 015)

Both JSON columns had zero readers/writers in code, tests, and migrations —
execution progress is tracked via progress_updates and artifacts via
commits/documents. Removes the ORM columns, the Pydantic Task.execution_log/
outputs fields, the ExecutionLog/FileRef models (+ their __init__ exports), and
the now-invalid kwargs from test fixtures. Migration 015 (down_revision
014_drop_pm_approvals) verified live: upgrade drops, downgrade re-adds.
Apply on the NAS with 'alembic upgrade head' at next deploy.

* chore(config): drop 16 unread Settings fields

Verified unused (no settings.X, no self.X property use, no getattr-by-name):
app_name, reload, workers, openai_api_key, secret_key, access_token_expire_minutes,
algorithm, log_level, log_format, the four session_* limits, message_max_length,
commit_subject_min_chars, commit_banned_words, agent_budget_sweep_interval_seconds.
Removes the empty Logging + Sessions&Messages sections and orphaned .env.example
vars. Kept: redis_db/redis_password (redis_url property), agent_sla_* (read via
getattr in task_lifecycle), encryption_key, and all live thresholds.

NOTE: commit_banned_words/commit_subject_min_chars and
agent_budget_sweep_interval_seconds were feature-config never wired to their
consumer (commit validator / budget sweep) — removed as dead, but flagged in
case the intent was to wire them.

* test(lifecycle): give i_will_work_on calls a substantive plan (#171 contract)

The real-DB lifecycle tests called i_will_work_on with a 13-char plan and no
risks/technical_considerations, so the substantive-plan gate (#171) rejected
them with incomplete_input — failing on master. Supply a >=150-char plan plus
technical_considerations and risks (mirroring tests/unit/gateway/
test_choreographer_dev.py). All 6 now pass; gate runs with no deselect.

* feat(gateway): wire commit-validator thresholds to settings

commit_subject_min_chars and commit_banned_words were config defined but never
read — the gateway commit() gate used the validator's hardcoded module defaults.
Re-add the two Settings fields and pass them through validate_commit_message in
content_actions.commit(), so config is the source of truth (validator defaults
remain the standalone/CI fallback). Adds wiring tests that monkeypatch settings
and assert the gate honors them.

* refactor(orchestrator): retire gateway_enabled flag; trigger_filter is unconditional

The gateway_enabled Settings field gated only the trigger_filter spawn-cooldown
(never the agent tool surface). Prod ran it on; the Phase-0 'legacy dispatch
path' it guarded no longer exists. Remove the field + the early-return branch in
gateway_pre_spawn_check so the cooldown runs for every spawn, drop the now-dead
ROBOCO_GATEWAY_ENABLED from docker-compose.yml, and update the stale Phase-0
comments + cooldown test. The per-container ROBOCO_GATEWAY_ENABLED env (set by
_append_manifest_args, read by agent_sdk to load the manifest) is unaffected.

* refactor(api): relabel /api/v2 -> /api/v1 as the canonical gateway surface

The gateway is the only agent API now, so the 'v2' label (with no v1) was
misleading. Renames roboco/api/routes/v2 -> routes/v1, schemas/v2 -> schemas/v1
(+ the matching test dirs and test_v2_role_dep/test_schemas_v2_flow files),
rewrites every /api/v2 path, routes.v2/schemas.v2 import, and v2-* router tag to
v1, and refreshes the stale 'v2' comments/docstrings. The panel is untouched (it
uses the unversioned /api/* REST routes). flow_server/do_server now POST to
/api/v1/*.

* docs(scripts): reset_runtime_state header matches actual SQL behavior

The header claimed it preserves groups + journals, but the .sql wipes both
(verified live: groups 6->0, journals 5->0; only agents/projects/channels
survive). Correct the wiped/preserved lists to match.

* refactor(gateway): extract _build_rich_plan to drop i_will_work_on under the complexity gate

i_will_work_on was cyclomatic rank C (11) — one over the xenon --max-absolute B
threshold — because of the five `x or default` fallbacks in the rich_plan dict.
Move that dict into a small _build_rich_plan helper (behaviour identical); both
methods are now rank B. make quality is fully green (xenon was its last failure;
bandit already passed — its 34 findings are all LOW severity, filtered by -ll).

* feat(foundation): add canonical CELL_TEAMS set; dedupe cell-subset literals

* feat(db): add ProductTable + ProductProjectTable ORM (per-cell project map)

* feat(task): add additive nullable product_id (ORM + model + DTO + create threading)

* feat(task): thread product_id through create_subtask/route/response

* feat(db): migration 016 — products, product_projects, tasks.product_id

* fix(db): document migration 016 plan deviations (revision len, FK name)

Two values in migration 016 intentionally diverge from the Task 2.4 plan
literals; this strengthens the in-file justification so the deviations are
self-documenting and verifiable.

- revision id (plan line 623): the plan's 36-char
  "016_add_products_and_task_product_id" overflows alembic_version.version_num
  (VARCHAR(32)) — alembic upgrade head raises asyncpg
  StringDataRightTruncationError. Kept at 27 chars
  ("016_add_products_product_id") so Step 4's live round-trip stays green.
- downgrade FK name (plan line 683): roboco/db/base.py sets a metadata
  naming_convention, so the FK upgrade() creates is
  "fk_tasks_product_id_products", not the Postgres default
  "tasks_product_id_fkey". The plan literal does not exist in the DB and
  would fail the downgrade with "constraint does not exist".

Both verified via the live upgrade/downgrade round-trip on a throwaway DB.

Issue 3 note: the prior commit (b896cac) also touched
tests/unit/api/test_schemas_tasks.py (added product_id=None to the
task_to_response stub). That line is load-bearing — task_to_response reads
task.product_id (added in Task 2.3, commit 67afa6b) — and belongs to Task 2.3's
scope; it is left in place because removing it breaks 4 tests and history is
not rewritten.

* refactor(db): trim migration 016 deviation notes to plan-faithful form

Reverts the out-of-scope documentation expansion (commit 1a4f296), which
was a second undocumented commit beyond Task 2.4's single plan-specified
commit and only bloated the migration docstring/comments.

The migration file now matches the plan-specified commit (b896cac) byte for
byte: the two necessary deviations from the plan literals stay (revision id
shortened to fit alembic_version.version_num VARCHAR(32); downgrade FK name
follows db/base.py's metadata naming_convention), each kept to a concise
inline note in the plan's header style.

The Task 2.3-scoped test stub line (tests/unit/api/test_schemas_tasks.py
product_id=None) is load-bearing — task_to_response reads task.product_id —
and is left in place; history is not rewritten.

Verified: live alembic upgrade head + downgrade to 015 round-trip on a
throwaway DB drops products/product_projects/tasks.product_id cleanly, and
make quality is green.

* refactor(test): annotate db_session and drop type: ignore in migration 016 test

Annotate the test_products_tables_and_task_fk_exist param as
db_session: AsyncSession (imported under TYPE_CHECKING) and remove the
# type: ignore[no-untyped-def] suppression, matching the typed db_session
pattern used across tests/integration/.

* feat(models): Product + ProductCreate/Update + ProductCellMapping (cell-validated)

* refactor(models): minimize ProductCellMapping config override to use_enum_values

The previous override re-declared validate_assignment, populate_by_name,
and extra=forbid, which RobocoBase already supplies. Pydantic merges
model_config across inheritance, so overriding only use_enum_values=False
is sufficient to keep team as a real Team enum (required so team in
CELL_TEAMS and enum identity hold for callers) while inheriting the rest
of the base config.

* fix(models): document ProductCellMapping use_enum_values override as plan-mandated

Resolves SPEC-COMPLIANCE review notes for Task 3.1 (Product domain models).

1. The ProductCellMapping use_enum_values=False override is a deviation from a
   bare project.py mirror, but it is mandated by the plan's own Task 3.1 code:
   RobocoBase sets use_enum_values=True, which coerces team to the plain string
   "backend". The plan's Step 1 test asserts m.team is Team.BACKEND (enum
   identity) and the Step 3 validator formats its error with v.value, both of
   which require team to remain a real Team enum. The override is therefore
   necessary; this commit relabels the comment to cite the specific spec lines
   that force it instead of leaving it as an unexplained departure. Downstream
   Task 3.2 (_replace_cells / project_for) already tolerates either form and the
   ORM stores the same value regardless, so the override has no behavioral reach
   beyond the in-memory enum identity the plan's test checks.

2. test_product_model.py hoists 'from uuid import uuid4' to module level rather
   than inline (as the plan's verbatim Step 1 code shows) because the global
   Pylint PLC0415 rule (import-outside-top-level) forbids inline imports and
   there is no per-file-ignore for tests/unit/models/. The hoisted form is the
   only ruff-clean rendering of the plan's test; left unchanged here.

3. Task 3.1 landed across two commits (c616d95 create, 6ebad255 refactor) rather
   than the plan's single Step 5 commit. Earlier history is intentionally not
   rewritten; this single follow-up commit brings the model to its final
   spec-faithful, fully-documented state.

* feat(service): ProductService CRUD + project_for per-cell resolver

* feat(api): Product CRUD routes + schemas, wired into the app

* fix(api): roll back and map cell-replacement IntegrityError on product update

update_product replaced cells via ProductService._replace_cells without
any try/except, so a duplicate-team cell (uq_product_projects_product_team)
or a non-existent project_id (product_projects.project_id FK) raised an
IntegrityError at flush, poisoning the AsyncSession and surfacing an
unhandled 500 with no rollback. Wrap the update + commit in a try/except
that rolls back and maps the UNIQUE violation to 409 and the FK violation
to 422, mirroring create_product's rollback discipline. Add integration
tests covering both client-error paths.

* fix(api): map create_product cell-mapping IntegrityError to 409/422

create_product only caught the slug conflict ('already exists' in str(e))
and bare-raised everything else, so a cells entry whose project_id does not
reference any project let the product_projects.project_id FK IntegrityError
propagate out of the route as an unhandled 500. The matching update_product
path was already hardened (uq_product_projects_product_team -> 409, FK
violation -> 422); apply the same mapping in create_product so a bad
project_id (or a duplicate-team cell) is a client error, not a server error.
The slug conflict is now caught as ConflictError directly instead of via a
broad except + string match.

* feat(gateway): add optional project_id to delegate inputs/request/routes

* feat(gateway): per-cell project routing (override -> product map -> parent) + product_id inheritance

* feat(task): approve_and_start — reassign board task to Main PM (CEO gate #1)

* feat(api): POST /tasks/{id}/approve-and-start (CEO gate #1, notes-required)

* test(api): cover approve-and-start 404-before-notes-gate for missing task

* feat(panel): Product types + Task.product_id

* feat(panel): productsApi + hooks + tasksApi.approveAndStart

* feat(panel): Products management screen + sidebar nav

* feat(panel): Approve & Start button (CEO gate #1)

* fix(api): narrow delete_product to IntegrityError + cover 204/409 delete paths

* test(task): assert approve_and_start persists + appends the audit note

* refactor(db): migration 016 names the tasks.product_id FK explicitly (house style)

* fix(db): make migrations authoritative + self-heal orphan product tables

init_db() no longer silently falls back to create_all when alembic upgrade
fails. That fallback masked migration failures and, since create_all cannot
ALTER an existing table, left the schema inconsistent — turning an unapplied
migration 016 into a crash loop: 016's CREATE TABLE products failed, the
upgrade rolled back, create_all re-created an empty orphan products table, and
every later boot failed again on the now-existing table while tasks.product_id
never got added. Now a migration failure is raised so the real error surfaces.

Migration 016 additionally drops EMPTY orphan products/product_projects tables
left by the old fallback before creating them, so an already-polluted DB
self-heals on the next deploy with no manual SQL. Skipped in offline (--sql)
mode; refuses to drop a table that holds rows.

* fix(db): create_all is the schema source of truth; alembic for increments

The Alembic chain is incomplete relative to the ORM — columns/tables like
notifications.delivered_at and the RAG indexed_documents table have NO migration
and have only ever been materialized by create_all. Tests don't catch this
because the test DB is also built via create_all, so migrations are never
exercised. The prior 'migrations are authoritative' init_db (and before it, the
create_all-only-on-failure fallback) therefore left a migrate-only boot with
missing columns/tables.

init_db now reflects reality:
  - Fresh DB  -> create_all builds the full current ORM schema, then stamp
                 Alembic at head so later incremental migrations apply.
  - Existing  -> run pending migrations (a real failure is raised, not masked),
                 then create_all(checkfirst) to gap-fill any missing ORM tables.
create_all cannot add a column to an existing table, so an ORM column added
without a migration needs a fresh rebuild of that table to appear.

* fix(db): migration 017 reconciles the Alembic chain with the full ORM schema

For years the live schema was built by create_all, not migrations, so the chain
drifted — tables/columns/indexes in the ORM had no migration (the
indexed_documents table, notifications.delivered_at, ~15 indexes, plus
timestamptz/server-default metadata). With init_db no longer masking that via a
create_all fallback, a migrate-only boot was missing those objects.

017 was produced by 'alembic revision --autogenerate' against Base.metadata,
reviewed, and verified: on a fresh DB, 'alembic upgrade head' (001..017) now
reproduces the create_all schema EXACTLY — a re-run of autogenerate detects zero
changes — and the 017 upgrade/downgrade round-trips cleanly. The migration chain
is now complete: migrate-only and create_all converge.

Also updates the init_db tests to assert the new behaviour (raise on an existing
DB's migration failure; create_all + stamp head on a fresh DB) instead of the
removed silent fallback.

* feat(panel): Product picker in the New Task form (drives per-cell routing)

The Products screen and Approve & Start button shipped, but the task-creation
form had no way to attach a Product — so a human couldn't set product_id from
the UI, which is exactly what drives per-cell project routing of delegated
subtasks. Adds an optional Product dropdown (Advanced -> Git config) populated
from useProducts(); 'None' falls back to the single project.

* fix(db): seed data is preserved on a fresh DB (run migrations, not bare create_all)

The previous fresh-DB path (create_all + stamp head) built the tables but never
ran the migration chain, so migration-embedded SEED DATA was skipped — most
visibly the AI providers seeded in 004. After a DB reset that left
provider_configs empty, so PUT /api/providers/ollama-key 404'd (the handler
raises NotFoundError when the Ollama provider row is missing).

Since migration 017 made the chain reproduce the full ORM schema, init_db now
runs 'alembic upgrade head' from base on a fresh DB — building every
table/column/index AND running the seeds. Verified: a fresh upgrade head seeds
both provider rows. Existing DBs still get migrations + create_all gap-fill.
Updates the init_db fresh-DB test accordingly.

* feat(task): project_id optional when a product_id is set (board fan-out tasks)

A board task that fans out across cells via a Product has no single repo of its
own — backend/frontend/ux_ui are each wrong, because the root coordinates and
delegates. Forcing one arbitrary Project was broken design (flagged at design
time). project_id is now nullable; a task must have project_id OR product_id:
  - TaskCreate model validator + a TaskService.create() invariant (covers every
    create path).
  - ORM/DTO/schema: project_id nullable; task_to_response uses to_python_uuid.
  - Gateway: a parent with only a product can delegate (guard now needs BOTH
    project and product to be None to reject); _resolve_subtask_project resolves
    each subtask from the product map and raises a clear error if a cell has no
    mapping and no parent project.
  - Migration 018 (tasks.project_id nullable), round-trip verified; fresh
    upgrade head still seeds providers.
  - Panel: Project no longer required once a Product is selected.
  - Removed the dead, never-called a2a create_task_from_message (it could only
    ever create a repo-less task) + its two coverage-only tests.

make quality green; panel tsc/lint/build green.

* Upgrade to Minimax M3

* fix(db): seed providers on existing DBs + correct enum casing

Migration 004 created the modelprovider/assignmentscope enums and seeded
provider rows in UPPERCASE, but the ORM (_str_enum) reads/writes the
lowercase StrEnum .value — so a fresh migrate-from-base DB built an enum
the ORM cannot read. Lowercase the enum labels and seed values in 004.

Add idempotent migration 019 to (re)seed the Anthropic + Ollama Cloud
providers with ON CONFLICT (name) DO NOTHING, so an existing DB whose
provider_configs table was created by create_all (and never ran 004's
seed) gets the rows on the next `alembic upgrade head` — fixing the
/api/providers/ollama-key 404 without a volume wipe.

* fix(tasks): let board/fan-out coordination tasks flow without a repo

A coordination task (project_id NULL, product_id set) targets no repo of
its own — it fans out to cell subtasks that each resolve a real project
from the product's cell->project map. Several paths still assumed every
task does git work and blocked it:

- orchestrator: add _is_coordination_task() and exempt these tasks from
  the project/branch/git-token gates in _readiness_check_task,
  _readiness_gate, _check_stuck_conditions, _validate_task_for_spawn.
- services/task.py: _ensure_branch_for_task returns "" (no branch) for a
  coordination task instead of raising; activate requires project OR
  product. This unblocks Main PM's i_will_plan claim, which otherwise
  raised before it could delegate the fan-out.
- gateway: _pending_assignment_guard exempts advisory roles
  (product_owner/head_marketing/auditor) from the "assigned but never
  claimed" idle gate — they review without claiming, so they could not
  satisfy a claim-or-unclaim remediation.

Adds focused unit tests for each.

* fix(tasks): coordination tasks reach in_progress + team reflects Main PM

The board->cells fan-out deadlocked: a coordination/fan-out task (product set,
no project of its own) could be created and claimed, but start()'s
claimed->in_progress transition hit validate_git_requirements, which still
demanded a branch_name and raised GitRequirementError. So Main PM's i_will_plan
never completed — it looped and never delegated. c961282 exempted
_ensure_branch_for_task (branch creation) but missed this parallel git gate in
the enforcement layer.

- task_lifecycle.py: add GitContext.is_coordination; skip the
  claimed->in_progress branch_name gate when it is set.
- task.py: populate is_coordination=(project_id is None and product_id is not
  None) in _validate_and_set_status; a branchless code task is still gated.
- approve_and_start: set team=Team.MAIN_PM on hand-off so the task isn't left
  labelled team=board after it leaves the board (now assigned to main-pm).

Adds a lifecycle-gate unit test and an end-to-end integration test that
claims, plans, and starts a project-less coordination task.

* fix(hooks): remove dead traceability hook + stale deleted-verb references

The v1-removal cleanup (2cfbf39) deleted the /traceability/remind SDK endpoint
but left the PostToolUse hook that curls it, so every gateway tool call 404'd
and agents silently lost their traceability reminders. Remove the dangling hook
(registration + TRACEABILITY_TRIGGER_TOOLS + Dockerfile COPY + the script); v2
carries per-verb guidance on the Envelope. Also correct two stale pre-gateway
tool names in hook text: the budget loop-detector nudged agents toward the
deleted roboco_task_escalate() (now unclaim()/i_am_idle(), which every looping
role has), and an sdk-startup comment referenced roboco_task_scan/get.

Extends the deleted-tool-name guard to scan docker/scripts/*.sh and to assert
every $SDK_URL/<path> a hook curls is a route still served by the SDK — the
check that would have caught this class (it lives in shell, invisible to mypy
and the Python import graph).

* fix(db): backfill ORM enum values the migration chain never added

Several StrEnum values were added to the ORM over time without a matching
`ALTER TYPE ... ADD VALUE` migration; 017 was autogenerate-derived and
autogenerate does not detect added enum labels, so the drift survived. On a DB
whose enum type predates the value, binding it raises at runtime — e.g.
`invalid input value for enum notificationtype: "a2a_request"` on
GET /api/notifications (list_system_notifications), and the same class for
blockerresolvertype/handoffstatus/team.

Migration 020 adds every drifted value idempotently (ADD VALUE IF NOT EXISTS —
no-op when 009 already reconciled it). Runs on the next `alembic upgrade head`.

Detected by comparing each ORM enum's values to the labels the migration chain
produces; adds tests/unit/test_enum_migration_parity.py which renders the chain
offline and fails on any future drift — the check that would have caught both
this and the provider-enum bug.

* fix(orchestrator): stop branch auto-block, board reassign, unblock livelock, agentless claims

Cluster C1 — four coupled orchestrator/task-invariant defects:

#18: a branch is created only at claim, so a pending, never-claimed code task
legitimately has no branch_name. The stuck-detection sweep (pending-only) and
readiness gate flagged that as "Task missing branch_name" and auto-blocked the
task every tick, so it never dispatched. Centralize the gate in
_branch_is_expected (status in claimed/in_progress/verifying, never a
coordination task) and apply it in both _check_stuck_conditions and
_readiness_check_task.

#14: the main_pm -> product_owner escalation rung handed an in_progress
descendant code task to the Product Owner (a board role) and marked it BLOCKED;
the board has no verb to own code work, so the dev's finished work deadlocked.
TaskService.apply_escalation (the single write primitive — covers both the
gateway escalate verb and the HTTP escalate route) now diverts a descendant code
task targeting a board/advisory role: it releases the task to PENDING for a
role-matched cell claim instead of stranding it.

#17: a blocked task reassigned to Main PM kept respawning the ex-assignee cell
PM to unblock it, but the assignee-only pre-unblock note returned not_authorized
— a livelock. _dispatch_blocker_work now dispatches the task's CURRENT PM/board
assignee (the unblock authority), falling back to the cell PM only when no
PM/board holds it. Also: a branchless coordination parent yields no valid merge
target — resolve_parent_branch now falls back to the child's own project default
branch (e.g. master) via TaskService.project_default_branch_for_task, and
_check_parent_branch_ready no longer blocks a child on a coordination parent's
non-existent branch.

#19: a task left claimed/in_progress with an assignee but no running container
was invisibly stuck (only PENDING tasks get fresh dispatch; the heartbeat reaper
can't see a freshly-seeded claim). New _dispatch_claimed_without_agent net:
after a short grace window it respawns the assignee, or releases the claim to
pending (lifecycle-safe via unclaim_for_reaper) when the assignee is unknown.
New config ROBOCO_CLAIMED_NO_AGENT_GRACE_SECONDS (default 120).

* fix(gateway): tolerant note verb + lock evidence do-tool invariant

#15: the note verb no longer hard-rejects thin decision/reflect payloads.
List-typed fields (options, consequences, next_steps) coerce a lone scalar
into a one-element list at both the NoteRequest schema (mode=before
validator) and the service layer; missing narrative fields default to a
visible placeholder instead of returning incomplete_input. The note is
always recorded, preserving audit value, and a well-intentioned note can no
longer trip the do-server 3-strikes circuit breaker. Widen the agent-facing
do_server.note hints to accept list-or-scalar and refresh the docstrings.

#8: add regression coverage locking the invariant that every role's do_tools
carries evidence (role_config + developer spawn manifest). The current source
already registers mcp__roboco-do__evidence for developers end-to-end; the
report stemmed from a stale deployed build, and the tests prevent silent
regression.

* fix(gateway): allow UX devs to receive design tasks; surface delegation rules to cell PM

The UX/UI cell's developers (ux-dev-1/ux-dev-2, Role.DEVELOPER on
Team.UX_UI) ARE its designers, but _validate_assignee_task_type rejected
task_type='design' for every DEVELOPER, blocking the UX cell's normal
design delegation. Allow 'design' for UX-team devs only; backend/frontend
devs stay rejected (design routing belongs to the UX cell). The
orchestrator already dispatches a developer for a design task
(_dev_dispatch_role_matches returns True), so this creates no orphan like
the documentation case.

Replace the static Cell-PM 'pass planning' remediate with a per-assignee
hint so a dev/design mis-type gets a developer-class next-step instead of
an off-topic planning hint.

Surface the three delegation guardrails in the cell-PM prompt so PMs stop
probing them by trial and error: valid task_type per assignee (incl.
design for UX devs), documentation auto-creation (non-delegatable), and
the sequential single-active code-spine. Fix the delegate-row task_type
list (documentation is NOT delegatable) and update the lifecycle spec
description; regenerate the lifecycle artifacts.

* fix(orchestrator): improve agent briefings for handoff consumption, product/project model, and workspace/secret hygiene

Main PM (roles/main_pm.md):
- Require reading the upstream Product Owner / Head of Marketing handoff
  (their decision/reflect journal entries + task description) BEFORE doing
  any own research or calling i_will_plan, so the Main PM builds on the
  Board's analysis instead of duplicating it. Added a dedicated section,
  hardened workflow step 1, and added an anti-pattern.
- Add a 'Products vs Projects' section: a Product fans out to one Project
  per cell; those Projects may be the SAME repo (monorepo subtrees) or
  DIFFERENT repos (multi-repo). The Main PM coordinates across them and
  must not assume one repo or call a monorepo subtree 'a separate repo'.
  Names the Prompter monorepo case (github.com/rennf93/roboco).

Developer (roles/developer.md):
- State the exact workspace path convention
  /data/workspaces/<project-slug>/<team>/<agent-slug>/, that the cwd is
  already set there, to stay inside the own cell workspace, and to not
  probe/guess the path (ls /, find /).
- Sanctioned secret handling: env/printenv is bash-guard denied and
  reveals nothing; needed secrets arrive via the task description, else
  i_am_blocked so the PM supplies them. Added matching anti-patterns.

Tests: add tests/unit/agents/test_briefing_cluster_c4.py asserting the
composed system prompt (the text mounted into agent containers) carries
each of the above.

* fix(orchestrator): board review involves PO+HoM and notifies CEO

Cluster C5 (#2, #4): a board/coordination task was reviewed by the Product
Owner alone, and the CEO got no formal signal when the review finished —
only buried channel chatter — so the Approve & Start handoff was invisible.

#4 — Board review is now a two-reviewer gate. _handle_board_assigned_task
dispatches BOTH the Product Owner and the Head of Marketing (one-shot each),
regardless of which one holds assigned_to, and the unassigned board-routing
path delegates here instead of claiming + spawning the PO alone. Board tasks
stay pending/unassigned for the CEO's Approve & Start. The board prompt now
makes the PO+HoM pair-review model explicit (HoM owns the UX/positioning
dimension).

#2 — Once BOTH reviewers have finished (dispatched and no longer active),
the orchestrator emits exactly one formal CEO notification via
NotificationService.send_board_review_complete_notification (APPROVAL type,
ack-required, carrying related_task_id) so the handoff is an actionable
signal. One-shot per task; a notification failure clears the guard so a
later tick can retry.

To let the non-assignee board member record its review note on a task held
by the other board member, content-action ownership now exempts a board role
posting to a board/coordination task (project_id is None, product_id set).
The exemption is narrow: it does not widen ownership for any other role or
any project-backed task.

Unit tests cover both reviewers dispatched, one-shot dispatch, the CEO
notification fired exactly once when both are done (and not before), the
retry-on-failure path, the notification builder, and the board co-review
ownership exemption (allowed for board+coordination, blocked otherwise).

* fix(workspace): install dev deps post-clone + raise git commit timeout for large changesets

Cluster C6 (#10, #13, #12-investigate).

#10: per-agent workspace clones never had the project's dev dependencies
installed, so the make-quality gate (ruff/mypy/pytest for Python, the TS
toolchain for the panel) was missing and devs re-downloaded tooling per
task. WorkspaceService now runs the project's install after cloning
(`uv sync` for Python, `pnpm install`/`npm ci`/`npm install` for Node/TS,
detected by manifest/lockfile). Idempotent via a lockfile-digest marker
under .git/ so a re-entry with unchanged lockfiles is a no-op; also runs on
the healthy short-circuit so pre-existing clones get backfilled. Gated by
workspace_install_dev_deps (default on) with workspace_dep_install_timeout_seconds.

#13: the gateway commit verb timed out on the large panel changeset because
every git op used the hardcoded 30s _GIT_TIMEOUT and each call also re-walks
the tree to chown. _run_git now takes a per-call timeout override sourced
from settings (git_command_timeout_seconds default); the staging + commit
ops in commit() and create_commit() use the longer git_commit_timeout_seconds
(default 180s). httpx REST timeouts unchanged in value.

#12 (investigate only — no push, no history change): the clone base ref is
NOT hardcoded; it already comes from project.default_branch threaded through
git.get_workspace -> ensure_workspace -> _clone_repo (git clone --branch).
The stale-base problem is a deploy/process issue (GitHub master is behind the
deployed migration chain), resolvable only by pushing the chain to master.
The default_branch column is the existing configurable lever.

* fix(panel): gate Approve & Start to board coordination tasks; stop 404 storm on closed sessions

CEO gate #1 button only renders for a PENDING board coordination/fan-out
task (no project_id, has product_id) — the board-reviewed handoff that
approve_and_start accepts — instead of every PENDING board-team task.
approve_and_start requires PENDING (it re-targets to Main PM without a
status change), so the gate stays on PENDING rather than the unrelated
end-of-work awaiting_ceo_approval state.

Session/message reads now treat a 404 as terminal and never retry it: a
reaped session is gone for good, and retrying every dead session-id is
what produced the growing 404 storm on GET /api/messages. The transcript
loads once (staleTime Infinity, no focus/reconnect refetch) so closed
sessions stay viewable without re-polling.

* fix(orchestrator): role-correct respawn prompt, throttle agentless dispatch, broaden #14 guard

#19 wrong-role prompt on respawn: _get_prompt_for_agent fell through to the
developer prompt for every non-dev/doc/qa role, so a respawned PM or board
agent was told to write code and call verbs it does not own. Route by the
agent's actual role through the existing per-role prompt builders
(developer/qa/documenter/cell_pm/main_pm/product_owner/head_marketing/auditor).
Both callers benefit; _spawn_pending_dev only ever passes developer/documenter/
unknown, so its behavior is unchanged.

#19 spawn-burst: _dispatch_claimed_without_agent looped over every agentless
claimed/in_progress task and could spawn many containers in one tick. Break
after the first respawn so a restart can't trigger a burst, matching every
sibling dispatcher. The release-to-pending path spawns nothing and keeps
draining stale unknown claims.

#14 guard scope: _is_descendant_code_task only matched CODE, so a descendant
DOCUMENTATION or DESIGN task escalated to a board/advisory role was still
stranded on a role with no verb to own it. Rename to
_is_descendant_executable_task and broaden to CODE/DOCUMENTATION/DESIGN — the
cell-executed types a board role cannot own. PLANNING/RESEARCH/ADMINISTRATIVE
route to a PM, not a cell agent, and are left unchanged; root tasks are still
reviewed up the chain.

* fix(docker): add node+pnpm to orchestrator so it pre-installs frontend cell deps

* Added .github workflows

* refactor(services): extract helpers to keep install_dev_deps + developer task-type check under the xenon complexity gate

* chore(github): add launch kit — CI, GHCR release, labels, templates, funding, dependabot npm, community docs

* chore(github): bump_version — drop unused noqa, fix datetime UTC import

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
This commit is contained in:
Renzo F
2026-06-03 06:35:03 +02:00
committed by GitHub
co-authored by Renn F
parent f83f930323
commit 110aaa7a77
211 changed files with 10852 additions and 9560 deletions
@@ -600,13 +600,18 @@ async def test_delegate_unknown_role_rejected() -> None:
@pytest.mark.asyncio
async def test_delegate_parent_no_project_rejected() -> None:
"""Line 1306: parent.project_id is None → invalid_state."""
"""A parent with NEITHER a project_id NOR a product_id → invalid_state.
(A parent with a product_id but no project is allowed — subtasks resolve a
repo from the product map — so the guard now requires both to be None.)
"""
pm_id = uuid4()
parent_id = uuid4()
parent = MagicMock(
status="in_progress",
assigned_to=pm_id,
project_id=None,
product_id=None,
title="p",
)
task_svc = AsyncMock()
@@ -629,7 +634,7 @@ async def test_delegate_parent_no_project_rejected() -> None:
)
body = env.as_dict()
assert body["error"] == "invalid_state"
assert "no project_id" in body["message"]
assert "project_id" in body["message"] and "product_id" in body["message"]
# ---------------------------------------------------------------------------
@@ -1325,3 +1330,38 @@ async def test_i_will_work_on_claimed_with_no_plan_accepts_recovery_plan() -> No
body = env.as_dict()
assert body["error"] is None, f"expected success, got {body}"
task_svc.set_plan.assert_awaited_once()
# ---------------------------------------------------------------------------
# _pending_assignment_guard: board/advisory roles can idle without claiming
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
@pytest.mark.parametrize("role", ["product_owner", "head_marketing", "auditor"])
async def test_pending_assignment_guard_exempts_board_roles(role: str) -> None:
"""A board/advisory agent that reviewed a still-pending coordination task
has no i_will_work_on/i_will_plan verb, so the idle gate must let it pass."""
task_svc = AsyncMock()
task_svc.list_assigned_for_agent.return_value = [
MagicMock(id=uuid4(), status="pending")
]
task_svc.agent_for.return_value = MagicMock(role=role, team=None, slug=None)
c = Choreographer(_make_deps(task=task_svc))
assert await c._pending_assignment_guard(uuid4(), {}) is None
@pytest.mark.asyncio
async def test_pending_assignment_guard_still_blocks_developer() -> None:
"""A developer holding a pending unclaimed task is still told to claim it."""
task_svc = AsyncMock()
task_svc.list_assigned_for_agent.return_value = [
MagicMock(id=uuid4(), status="pending")
]
task_svc.agent_for.return_value = MagicMock(
role="developer", team="backend", slug=None
)
c = Choreographer(_make_deps(task=task_svc))
guard = await c._pending_assignment_guard(uuid4(), {})
assert guard is not None
assert guard.as_dict()["error"] == "invalid_state"
+105 -46
View File
@@ -6,6 +6,7 @@ from unittest.mock import AsyncMock, MagicMock
from uuid import uuid4
import pytest
from roboco.config import settings
from roboco.services.gateway.content_actions import ContentActions, ContentActionsDeps
@@ -270,8 +271,13 @@ async def test_note_reflect_scope_succeeds() -> None:
@pytest.mark.asyncio
async def test_note_reflect_missing_required_fields_returns_incomplete_input() -> None:
"""Pre-gateway parity: reflect without structured fields fails fast."""
async def test_note_reflect_missing_fields_records_with_placeholder() -> None:
"""Issue #15: a thin reflect note is recorded, not rejected.
Missing narrative fields are defaulted to a visible placeholder so the
entry still lands (audit value preserved) and the do-server circuit
breaker never fires on a well-intentioned note.
"""
agent_id = uuid4()
task_id = uuid4()
task_svc = AsyncMock()
@@ -292,15 +298,21 @@ async def test_note_reflect_missing_required_fields_returns_incomplete_input() -
)
body = env.as_dict()
assert body["error"] == "incomplete_input"
missing = set(body["missing"])
assert {"what_done", "what_learned", "what_struggled"}.issubset(missing)
journal_svc.write_entry.assert_not_awaited()
assert body["error"] is None
assert body["status"] == "noted"
journal_svc.write_entry.assert_awaited_once()
content = journal_svc.write_entry.call_args.kwargs["content"]
# Missing what_done/what_learned/what_struggled render as the placeholder.
assert "(not provided)" in content
@pytest.mark.asyncio
async def test_note_decision_requires_options_and_more() -> None:
"""Pre-gateway parity: decision requires context/options(>=2)/chosen/rationale."""
async def test_note_decision_thin_payload_records_not_rejected() -> None:
"""Issue #15: decision with missing/thin fields is recorded, not rejected.
A single option is kept as-is (the min-2 gate no longer hard-blocks);
missing context/chosen/rationale default to a placeholder.
"""
agent_id = uuid4()
task_id = uuid4()
task_svc = AsyncMock()
@@ -313,7 +325,7 @@ async def test_note_decision_requires_options_and_more() -> None:
deps = _make_deps(task=task_svc, journal=journal_svc)
ca = ContentActions(deps)
# Missing everything structured → incomplete_input listing all required.
# Bare decision: no structured fields → still recorded with placeholders.
env = await ca.note(
agent_id=agent_id,
text="bare decision",
@@ -321,10 +333,10 @@ async def test_note_decision_requires_options_and_more() -> None:
task_id=task_id,
)
body = env.as_dict()
assert body["error"] == "incomplete_input"
assert {"context", "options", "chosen", "rationale"}.issubset(set(body["missing"]))
assert body["error"] is None
assert body["status"] == "noted"
# Single option still fails (min 2).
# Single option no longer fails — recorded as-is.
env = await ca.note(
agent_id=agent_id,
text="decision with one option",
@@ -338,10 +350,10 @@ async def test_note_decision_requires_options_and_more() -> None:
},
)
body = env.as_dict()
assert body["error"] == "incomplete_input"
assert "options" in body["missing"]
assert body["error"] is None
assert body["status"] == "noted"
# Two options + all required → success.
# Fully-filled decision still succeeds (no regression).
env = await ca.note(
agent_id=agent_id,
text="real decision",
@@ -361,26 +373,48 @@ async def test_note_decision_requires_options_and_more() -> None:
assert body["error"] is None
assert body["status"] == "noted"
# Three+ options also pass — 2 is the floor, not the ceiling.
@pytest.mark.asyncio
async def test_note_decision_scalar_list_fields_are_coerced() -> None:
"""Issue #15: lone-scalar options/consequences are wrapped into lists.
An agent that passes a single option dict or a single consequences
string must not be rejected the value is wrapped into a one-element
list and rendered into the entry.
"""
agent_id = uuid4()
task_id = uuid4()
task_svc = AsyncMock()
task_svc.get_active_task_for_agent.return_value = None
task_svc.get.return_value = MagicMock(
id=task_id, assigned_to=agent_id, status="in_progress"
)
journal_svc = AsyncMock()
deps = _make_deps(task=task_svc, journal=journal_svc)
ca = ContentActions(deps)
env = await ca.note(
agent_id=agent_id,
text="three-way decision",
text="decision with scalar fields",
scope="decision",
task_id=task_id,
structured={
"context": "queue tech choice",
"options": [
{"name": "redis", "pros": "fast", "cons": "ephemeral"},
{"name": "postgres", "pros": "durable", "cons": "slower"},
{"name": "rabbitmq", "pros": "ordered", "cons": "ops overhead"},
],
"chosen": "rabbitmq",
"rationale": "ordering matters more than raw speed here",
"context": "queue tech",
# single dict instead of a list
"options": {"name": "redis", "pros": "fast", "cons": "ephemeral"},
"chosen": "redis",
"rationale": "speed",
# single string instead of a list
"consequences": "we lose durability across restarts",
},
)
body = env.as_dict()
assert body["error"] is None
assert body["status"] == "noted"
content = journal_svc.write_entry.call_args.kwargs["content"]
assert "redis" in content
assert "we lose durability across restarts" in content
@pytest.mark.asyncio
@@ -659,17 +693,19 @@ async def test_verify_explicit_task_ownership_returns_not_found() -> None:
# ---------------------------------------------------------------------------
# B4 — decision/reflect remediate includes a literal call example
# Issue #15 — thin decision/reflect notes are recorded, never rejected
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_decision_incomplete_input_includes_call_example() -> None:
"""Decision rejection includes a literal note(scope='decision', ...) template."""
async def test_decision_thin_note_records_without_rejection() -> None:
"""A bare decision (no structured fields) is recorded, not rejected."""
task = AsyncMock()
task.get_active_task_for_agent.return_value = None
task.get_journal_context_task_for_agent.return_value = None
task.agent_for.return_value = MagicMock(role="cell_pm")
deps = _make_deps(task=task)
journal_svc = AsyncMock()
deps = _make_deps(task=task, journal=journal_svc)
ca = ContentActions(deps)
env = await ca.note(
@@ -678,23 +714,20 @@ async def test_decision_incomplete_input_includes_call_example() -> None:
scope="decision",
)
body = env.as_dict()
assert body["error"] == "incomplete_input"
remediate = body.get("remediate", "")
# Must include a literal call template, not just a field list
assert "note(scope='decision'" in remediate, remediate
assert "context=" in remediate, remediate
assert "options=[" in remediate, remediate
assert "chosen=" in remediate, remediate
assert "rationale=" in remediate, remediate
assert body["error"] is None
assert body["status"] == "noted"
journal_svc.write_entry.assert_awaited_once()
@pytest.mark.asyncio
async def test_reflect_incomplete_input_includes_call_example() -> None:
"""Reflect rejection includes a literal note(scope='reflect', ...) call template."""
async def test_reflect_thin_note_records_without_rejection() -> None:
"""A bare reflect (no structured fields) is recorded, not rejected."""
task = AsyncMock()
task.get_active_task_for_agent.return_value = None
task.get_journal_context_task_for_agent.return_value = None
task.agent_for.return_value = MagicMock(role="developer")
deps = _make_deps(task=task)
journal_svc = AsyncMock()
deps = _make_deps(task=task, journal=journal_svc)
ca = ContentActions(deps)
env = await ca.note(
@@ -703,12 +736,9 @@ async def test_reflect_incomplete_input_includes_call_example() -> None:
scope="reflect",
)
body = env.as_dict()
assert body["error"] == "incomplete_input"
remediate = body.get("remediate", "")
assert "note(scope='reflect'" in remediate, remediate
assert "what_done=" in remediate, remediate
assert "what_learned=" in remediate, remediate
assert "what_struggled=" in remediate, remediate
assert body["error"] is None
assert body["status"] == "noted"
journal_svc.write_entry.assert_awaited_once()
# ---------------------------------------------------------------------------
@@ -798,3 +828,32 @@ async def test_progress_ownership_enforced() -> None:
env = await actions.progress(agent_id=agent_id, task_id=t.id, message="x")
assert env.as_dict()["error"] is not None
task.record_plan_progress.assert_not_awaited()
@pytest.mark.asyncio
async def test_commit_gate_reads_settings_min_chars(
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""The commit gate reads settings.commit_subject_min_chars."""
monkeypatch.setattr(settings, "commit_subject_min_chars", 500)
ca = ContentActions(_make_deps())
# Normally a fine descriptive subject, now shorter than the bumped minimum.
env = await ca.commit(
agent_id=uuid4(),
message="feat(api): add /healthz endpoint for liveness checks",
)
assert env.as_dict()["error"] == "invalid_state"
@pytest.mark.asyncio
async def test_commit_gate_reads_settings_banned_words(
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""The commit gate reads settings.commit_banned_words."""
monkeypatch.setattr(settings, "commit_banned_words", ("bananaword",))
monkeypatch.setattr(settings, "commit_subject_min_chars", 1)
ca = ContentActions(_make_deps())
env = await ca.commit(agent_id=uuid4(), message="bananaword")
assert env.as_dict()["error"] == "invalid_state"
@@ -336,6 +336,110 @@ async def test_evidence_blocks_when_not_assignee() -> None:
git_svc.diff.assert_not_awaited()
# ---------------------------------------------------------------------------
# Board co-review exemption (cluster C5): a board role may record its review
# note/say on a board/coordination task held by the OTHER board member.
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_board_role_may_note_coordination_task_held_by_other_board() -> None:
"""A board/coordination task (project_id=None, product_id set) is reviewed
by BOTH board members; the non-assignee board reviewer may still note it."""
agent_id = uuid4()
other_board_id = uuid4()
task_id = uuid4()
coord_task = MagicMock(
id=task_id,
status="pending",
assigned_to=other_board_id,
project_id=None,
product_id=uuid4(),
)
task_svc = AsyncMock()
task_svc.get.return_value = coord_task
task_svc.get_active_task_for_agent.return_value = None
journal_svc = AsyncMock()
deps = _make_deps(task=task_svc, journal=journal_svc)
# _make_deps defaults agent_for to a developer; override AFTER so the
# board co-review exemption sees a board role.
task_svc.agent_for.return_value = MagicMock(role="head_marketing")
ca = ContentActions(deps)
env = await ca.note(
agent_id=agent_id,
text="UX + positioning review of the board task",
scope="note",
task_id=task_id,
)
assert env.error is None
journal_svc.write_entry.assert_awaited_once()
@pytest.mark.asyncio
async def test_board_role_blocked_on_project_task_held_by_other() -> None:
"""The exemption is narrow: a board role still cannot post to a normal
project-backed task assigned to someone else."""
agent_id = uuid4()
other_id = uuid4()
task_id = uuid4()
project_task = MagicMock(
id=task_id,
status="in_progress",
assigned_to=other_id,
project_id=uuid4(),
product_id=None,
)
task_svc = AsyncMock()
task_svc.get.return_value = project_task
journal_svc = AsyncMock()
deps = _make_deps(task=task_svc, journal=journal_svc)
# Board role, but a project-backed task — exemption must NOT apply.
task_svc.agent_for.return_value = MagicMock(role="product_owner")
ca = ContentActions(deps)
env = await ca.note(
agent_id=agent_id,
text="trying to note a code task I don't own",
scope="note",
task_id=task_id,
)
body = env.as_dict()
assert body["error"] == "not_authorized"
journal_svc.write_entry.assert_not_awaited()
@pytest.mark.asyncio
async def test_non_board_role_blocked_on_coordination_task_held_by_other() -> None:
"""The exemption is board-only: a developer cannot piggyback on it."""
agent_id = uuid4()
other_id = uuid4()
task_id = uuid4()
coord_task = MagicMock(
id=task_id,
status="pending",
assigned_to=other_id,
project_id=None,
product_id=uuid4(),
)
task_svc = AsyncMock()
task_svc.get.return_value = coord_task
task_svc.agent_for.return_value = MagicMock(role="developer")
journal_svc = AsyncMock()
deps = _make_deps(task=task_svc, journal=journal_svc)
ca = ContentActions(deps)
env = await ca.note(
agent_id=agent_id,
text="dev trying to note a coordination task",
scope="note",
task_id=task_id,
)
body = env.as_dict()
assert body["error"] == "not_authorized"
journal_svc.write_entry.assert_not_awaited()
@pytest.mark.asyncio
async def test_evidence_unassigned_task_allows_inspection() -> None:
"""A task with assigned_to=None (between handoffs) may be inspected.
@@ -0,0 +1,196 @@
"""#7: assignee-vs-task_type matrix for delegate().
Cell-PM delegation friction: agents burned turns probing the
``_validate_assignee_task_type`` guard by trial and error. Two behaviors
are pinned here:
1. The UX/UI cell's developers ARE its designers — ``task_type='design'``
is legitimate work for ``ux-dev-1``/``ux-dev-2`` (Role.DEVELOPER on
Team.UX_UI), so delegate must accept it. Backend/frontend devs are NOT
design assignees and stay rejected.
2. The rejection ``remediate`` is per-assignee-class: a developer mis-type
gets a developer hint (not the generic 'pass planning to a Cell PM').
"""
from __future__ import annotations
from datetime import UTC, datetime
from typing import Any
from unittest.mock import AsyncMock, MagicMock
from uuid import uuid4
import pytest
from roboco.services.gateway.choreographer import Choreographer, ChoreographerDeps
from roboco.services.gateway.choreographer._impl import DelegateInputs
def _make_deps(**overrides: Any) -> ChoreographerDeps:
base: dict[str, Any] = {
"task": AsyncMock(),
"work_session": AsyncMock(),
"git": AsyncMock(),
"a2a": AsyncMock(),
"journal": AsyncMock(),
"audit": AsyncMock(),
"evidence_repo": AsyncMock(),
"messaging": AsyncMock(),
}
base.update(overrides)
repo = base["evidence_repo"]
for m in (
"list_unread_a2a",
"list_unread_mentions",
"list_pending_notifications",
"task_metadata_gaps",
"recent_team_activity",
"blockers_in_lane",
"journal_highlights_for_task",
):
getattr(repo, m).return_value = []
_ldef = base["journal"].latest_decision_at.return_value
if type(_ldef).__name__ in ("MagicMock", "AsyncMock"):
base["journal"].latest_decision_at.return_value = datetime.now(UTC)
return ChoreographerDeps(**base)
def _parent(parent_id: object, team: str) -> MagicMock:
return MagicMock(
id=parent_id,
project_id=uuid4(),
team=team,
status="in_progress",
task_type="planning",
sequence=0,
assigned_to=uuid4(),
)
def _inputs(assigned_to: str, team: str, task_type: str) -> DelegateInputs:
return DelegateInputs(
title="Cell work item",
description="A real unit of work for the cell to deliver end to end",
acceptance_criteria=["the deliverable exists", "it is linked to a PR"],
assigned_to=assigned_to,
team=team,
task_type=task_type,
nature="technical",
estimated_complexity="medium",
)
# --- Behavior 1: validator matrix (pure function, no DB) ----------------
# The static-guard rejection for documentation lives in
# `_delegate_static_guards` (Task #163), NOT in `_validate_assignee_task_type`
# — the validator allows `documentation` for any dev role. So
# `documentation` is exercised separately in the static-guard tests below.
_CELL_PM_PLANNING_HINT = "delegating to a Cell PM"
@pytest.mark.parametrize(
("assigned_to", "task_type", "allowed"),
[
# UX devs are designers — design is legitimate.
("ux-dev-1", "design", True),
("ux-dev-2", "design", True),
("ux-dev-1", "code", True),
("ux-dev-1", "research", True),
# Backend/frontend devs are NOT design assignees.
("be-dev-1", "design", False),
("fe-dev-2", "design", False),
# Code is fine for any dev.
("be-dev-1", "code", True),
("fe-dev-1", "research", True),
],
)
def test_validate_assignee_task_type_matrix(
assigned_to: str, task_type: str, allowed: bool
) -> None:
err = Choreographer._validate_assignee_task_type(assigned_to, task_type)
if allowed:
assert err is None, f"{assigned_to}/{task_type} should be allowed, got {err!r}"
else:
assert err is not None, f"{assigned_to}/{task_type} should be rejected"
assert assigned_to in err and task_type in err
def test_remediate_for_ux_dev_mentions_design() -> None:
hint = Choreographer._assignee_task_type_remediate("ux-dev-1")
assert "design" in hint.lower()
# Not the generic Cell-PM planning hint.
assert _CELL_PM_PLANNING_HINT not in hint
def test_remediate_for_backend_dev_routes_design_to_ux() -> None:
hint = Choreographer._assignee_task_type_remediate("be-dev-1")
assert "ux" in hint.lower()
assert "code" in hint.lower()
# Not the generic Cell-PM planning hint.
assert _CELL_PM_PLANNING_HINT not in hint
def test_remediate_for_cell_pm_keeps_planning_hint() -> None:
hint = Choreographer._assignee_task_type_remediate("be-pm")
assert _CELL_PM_PLANNING_HINT in hint
# --- Behavior 2: static-guard envelope wiring ---------------------------
@pytest.mark.asyncio
async def test_static_guards_allow_ux_design_subtask() -> None:
"""ux-pm delegating a design subtask to its designer passes the guard."""
pm_id = uuid4()
parent_id = uuid4()
c = Choreographer(_make_deps())
env = await c._delegate_static_guards(
pm_id,
parent_id,
_parent(parent_id, "ux_ui"),
_inputs("ux-dev-1", "ux_ui", "design"),
)
assert env is None, f"UX design subtask must pass static guards, got {env}"
@pytest.mark.asyncio
async def test_static_guards_reject_backend_design_with_dev_remediate() -> None:
"""be-pm handing a design subtask to a backend dev is rejected, and the
remediate is the developer-class hint (not the planning/Cell-PM hint)."""
pm_id = uuid4()
parent_id = uuid4()
c = Choreographer(_make_deps())
env = await c._delegate_static_guards(
pm_id,
parent_id,
_parent(parent_id, "backend"),
_inputs("be-dev-1", "backend", "design"),
)
assert env is not None, "design subtask for a backend dev must be rejected"
body = env.as_dict()
assert body["error"] == "invalid_state", body
remediate = body["remediate"] or ""
assert "UX" in remediate
# The developer-class hint, NOT the generic Cell-PM planning hint.
assert _CELL_PM_PLANNING_HINT not in remediate
@pytest.mark.asyncio
async def test_static_guards_reject_documentation_for_ux_dev() -> None:
"""documentation stays non-delegatable even for a UX dev (Task #163):
the lifecycle auto-creates the doc phase after the code subtask."""
pm_id = uuid4()
parent_id = uuid4()
c = Choreographer(_make_deps())
env = await c._delegate_static_guards(
pm_id,
parent_id,
_parent(parent_id, "ux_ui"),
_inputs("ux-dev-1", "ux_ui", "documentation"),
)
assert env is not None, "documentation subtask must be rejected"
body = env.as_dict()
assert body["error"] == "invalid_state", body
assert "not PM-" in (body["message"] or "") or "documenter" in (
body["remediate"] or ""
)
@@ -0,0 +1,125 @@
from __future__ import annotations
from datetime import UTC, datetime
from typing import Any
from unittest.mock import AsyncMock, MagicMock
from uuid import uuid4
import pytest
from roboco.services.gateway.choreographer import (
Choreographer,
ChoreographerDeps,
DelegateInputs,
)
def _make_deps(**overrides: Any) -> ChoreographerDeps:
base: dict[str, Any] = {
"task": AsyncMock(),
"work_session": AsyncMock(),
"git": AsyncMock(),
"a2a": AsyncMock(),
"journal": AsyncMock(),
"audit": AsyncMock(),
"evidence_repo": AsyncMock(),
}
base.update(overrides)
repo = base["evidence_repo"]
for m in (
"list_unread_a2a",
"list_unread_mentions",
"list_pending_notifications",
"task_metadata_gaps",
"recent_team_activity",
"blockers_in_lane",
"journal_highlights_for_task",
):
getattr(repo, m).return_value = []
_ldef = base["journal"].latest_decision_at.return_value
if type(_ldef).__name__ in ("MagicMock", "AsyncMock"):
base["journal"].latest_decision_at.return_value = datetime.now(UTC)
return ChoreographerDeps(**base)
def _parent(pm_id, product_id=None, project_id=None):
return MagicMock(
id=uuid4(),
project_id=project_id or uuid4(),
product_id=product_id,
status="in_progress",
assigned_to=pm_id,
)
def _inputs(**kw: Any) -> DelegateInputs:
base: dict[str, Any] = {
"title": "Implement endpoint",
"description": "Add /v1/foo endpoint with tests",
"assigned_to": "be-dev-1",
"team": "backend",
"task_type": "code",
"nature": "technical",
"acceptance_criteria": ["GET /v1/foo returns 200 with body"],
}
base.update(kw)
return DelegateInputs(**base)
async def _run(parent, inputs, product=None):
pm_id = parent.assigned_to
task_svc = AsyncMock()
task_svc.get.return_value = parent
task_svc.agent_for.return_value = MagicMock(role="cell_pm", team="backend")
task_svc.get_subtasks.return_value = []
task_svc.create_subtask.return_value = MagicMock(id=uuid4())
deps = _make_deps(task=task_svc, **({"product": product} if product else {}))
c = Choreographer(deps)
env = await c.delegate(pm_id, parent.id, inputs)
return env, task_svc
@pytest.mark.asyncio
async def test_explicit_project_id_overrides_everything() -> None:
override = uuid4()
parent = _parent(uuid4(), product_id=uuid4())
env, task_svc = await _run(parent, _inputs(project_id=override))
assert env.error is None, env.as_dict()
req = task_svc.create_subtask.call_args.args[0]
assert req.project_id == override
@pytest.mark.asyncio
async def test_product_map_resolves_project_when_no_override() -> None:
mapped = uuid4()
product_id = uuid4()
parent = _parent(uuid4(), product_id=product_id)
product = AsyncMock()
product.project_for.return_value = mapped
env, task_svc = await _run(parent, _inputs(), product=product)
assert env.error is None, env.as_dict()
req = task_svc.create_subtask.call_args.args[0]
assert req.project_id == mapped
assert req.product_id == product_id # inherited onto the subtask
product.project_for.assert_awaited_once()
@pytest.mark.asyncio
async def test_falls_back_to_parent_project_when_no_product() -> None:
parent = _parent(uuid4(), product_id=None)
env, task_svc = await _run(parent, _inputs())
assert env.error is None, env.as_dict()
req = task_svc.create_subtask.call_args.args[0]
assert req.project_id == parent.project_id
@pytest.mark.asyncio
async def test_partial_product_map_degrades_to_parent_project() -> None:
product_id = uuid4()
parent = _parent(uuid4(), product_id=product_id)
product = AsyncMock()
product.project_for.return_value = None # no mapping for this cell
env, task_svc = await _run(parent, _inputs(), product=product)
assert env.error is None, env.as_dict()
req = task_svc.create_subtask.call_args.args[0]
assert req.project_id == parent.project_id
assert req.product_id == product_id
+22 -2
View File
@@ -80,14 +80,34 @@ class TestResolveParentBranch:
task_service.get.assert_not_called()
@pytest.mark.asyncio
async def test_falls_back_when_parent_has_no_branch(self) -> None:
async def test_branchless_parent_uses_project_default_branch(self) -> None:
# #17: a branchless coordination parent never gets a branch. The child
# was cut from its own project's default branch, so that is the real
# merge target — NOT a string-derived ref the parent never created
# (which would have no valid merge target and wedge the cell↔Main-PM
# loop).
task = MagicMock(
parent_task_id=uuid4(),
branch_name="feature/backend/ROOT0001--CELL0001",
)
task_service = AsyncMock()
task_service.get = AsyncMock(return_value=MagicMock(branch_name=None))
# Parent exists but has no branch yet → string derivation.
task_service.project_default_branch_for_task = AsyncMock(return_value="master")
result = await resolve_parent_branch(task, task_service)
assert result == "master"
task_service.project_default_branch_for_task.assert_awaited_once_with(task)
@pytest.mark.asyncio
async def test_branchless_parent_falls_back_to_string_when_no_project(self) -> None:
# No project to consult (resolver returns None) → string derivation
# remains the last-resort fallback.
task = MagicMock(
parent_task_id=uuid4(),
branch_name="feature/backend/ROOT0001--CELL0001",
)
task_service = AsyncMock()
task_service.get = AsyncMock(return_value=MagicMock(branch_name=None))
task_service.project_default_branch_for_task = AsyncMock(return_value=None)
result = await resolve_parent_branch(task, task_service)
assert result == "feature/backend/ROOT0001"
-15
View File
@@ -3,21 +3,12 @@
from __future__ import annotations
from roboco.services.gateway.remediation import (
hint_for_missing_plan,
hint_for_missing_progress,
hint_for_missing_reflect,
hint_for_unaddressed_acceptance_criteria,
hint_for_unread_a2a,
)
def test_missing_plan_hint() -> None:
h = hint_for_missing_plan(task_id="abc-123")
assert "i_will_work_on" in h
assert "abc-123" in h
assert "plan=" in h
def test_missing_progress_hint() -> None:
h = hint_for_missing_progress()
assert "commit" in h.lower() or "progress" in h.lower()
@@ -36,9 +27,3 @@ def test_unaddressed_criteria_hint() -> None:
assert "criterion 1" in h
assert "criterion 3" in h
assert "t-1" in h
def test_unread_a2a_hint() -> None:
h = hint_for_unread_a2a(count=2, task_id="t-1")
assert "2" in h
assert "t-1" in h
+9
View File
@@ -67,6 +67,15 @@ class TestRoleConfigCatalog:
assert "ToolSearch" not in cfg.flow_tools
assert "ToolSearch" not in cfg.do_tools
def test_every_role_has_evidence(self) -> None:
# Issue #8: a developer container shipped without
# mcp__roboco-do__evidence. `evidence` is a read-only inspection
# tool every role needs (devs read their own PR diff, QA/PM review,
# the auditor inspects). Lock the invariant so no role's do-tool
# tuple can silently drop it again.
for role, cfg in ROLE_CONFIGS.items():
assert "evidence" in cfg.do_tools, f"{role} missing evidence do-tool"
def test_dev_flow_matches_spec_intents_for_role() -> None:
"""role_config._DEV_FLOW must equal spec.intents_for_role(Role.DEVELOPER)."""