Files
roboco/scripts/build_lifecycle_artifacts.py
207aaecd72 Feature: lifecycle canonical spec (#14)
* chore: clean make quality baseline on feature/lifecycle-canonical-spec

Three classes of pre-existing issues blocking `make quality`:

1. Alembic migrations 002/009/011 used runtime introspection
   (op.get_bind() + inspect / bind.execute) without guarding for
   offline (--sql) mode. `alembic upgrade head --sql` is part of
   `make quality`; in offline mode `op.get_bind()` returns a
   MockConnection with no inspection system, so the migrations
   crashed before emitting their SQL stubs. Each migration now
   short-circuits or simplifies in `context.is_offline_mode()` —
   live-DB behavior is unchanged.

2. ruff format drift on three files left over from prior in-flight
   edits (choreographer/_impl.py, content_actions.py, and one test
   file). `ruff format` applied.

3. vulture flagged two unused `tb` parameters in async __aexit__
   stubs in test_task_service_lifecycle_misc.py. The parameter is
   protocol-required but unused by the body — renamed to `_tb`
   (vulture treats underscore-prefixed names as intentionally unused).

`make quality` is now green from this branch's HEAD; subsequent
lifecycle-spec work can use it as the per-task gate.

* feat(lifecycle): canonical spec package + Role/Status/TaskType enums

Foundation for the canonical lifecycle/permissions module. Enums
mirror docs/internal/old/workflows/STATUS_TRANSITIONS.md +
PERMISSIONS.md. Tests pin enum membership against both the
predecessor canon and roboco.models.base.TaskType.

* feat(lifecycle): Decision dataclass with allow/reject/tracing_gap constructors

Single rejection shape every consumer maps to its native format
(Envelope, HTTP code, prompt hint). __post_init__ enforces the
allowed/rejection_kind invariants so a malformed Decision can't reach
a consumer.

* fix(lifecycle): tighten Decision invariants per Task 2 review

Two reviewer findings on the Task 2 Decision dataclass, addressed
in one commit:

1. The docstring promised `allowed=True ⇒ rejection_kind is None
   AND missing == [] AND remediate is None`, but __post_init__ only
   checked the rejection_kind half. A caller could construct an
   allow-shaped Decision with stale missing/remediate fields and
   sneak it past validation. Tighten __post_init__ to enforce the
   full invariant. Add a regression test.

2. tracing_gap defensively copies the missing list (`list(missing)`)
   to isolate the stored list from later caller-side mutation, but
   no test pinned this. Add a regression test that mutates the source
   list after construction and asserts the stored list is unchanged.

Issue 2 from the same review (mutable list vs tuple for `missing`)
is a broader design call deferred until consumers exist; the
defensive copy is sufficient until then.

* feat(lifecycle): Precondition/ActionSpec/IntentSpec/StatusTransition dataclasses

The four dataclasses that hold the canonical tables. ActionSpec and
StatusTransition are direct ports of pre-gateway PERMISSIONS.md +
STATUS_TRANSITIONS.md rows. IntentSpec is the gateway-only addition:
each gateway intent verb declares which atomic actions it composes.

* feat(lifecycle): _STATUS_TRANSITIONS table + STATUS_GRAPH view

Direct port of STATUS_TRANSITIONS.md. Every transition records its
trigger action and (optionally) a role constraint. STATUS_GRAPH is
the precomputed source→{targets} view callers use for reachability
checks.

* fix(lifecycle): pin role_constraint values + clarify Task-5 handoff

Two reviewer findings on Task 4 _STATUS_TRANSITIONS, addressed in
one commit:

1. The original Task-4 tests verified (source, target) pairs but
   not role_constraint contents. A typo in a single role name (e.g.
   forgetting MAIN_PM from escalate_to_ceo) would have slipped past
   them silently. Add test_status_transitions_role_constraints_match_canon
   pinning every non-None constraint and the cancel-block invariant.

2. role_constraint=None on the `claim` rows from PENDING and
   NEEDS_REVISION was load-bearing — it is the explicit handoff
   point between the StatusTransition table (state machine layer)
   and CLAIM_RULES (per-role claim authority, lands in Task 5).
   The original inline comment said this in passing; expand it so
   the design choice is unmissable for a stranger reading just
   spec.py.

* feat(lifecycle): _ATOMIC_ACTIONS + CLAIM_RULES + ROLE_TEAM_RULES tables

Direct port of PERMISSIONS.md. Every task management tool gets an
ActionSpec with allowed_roles, source_statuses, target_status,
self_review_block, and needs_team_match flags. CLAIM_RULES maps each
Role to the statuses they can claim from. ROLE_TEAM_RULES is the
per-slug team restriction.

* fix(lifecycle): tighten ActionSpec contracts per Task 5 review

Three reviewer findings on Task 5's _ATOMIC_ACTIONS table, addressed
in one commit:

1. set_plan.source_statuses widened to {CLAIMED, IN_PROGRESS} but
   every existing caller (i_will_work_on / i_will_plan compositions)
   runs set_plan while CLAIMED, between claim and start. Narrow to
   {CLAIMED} only. If a future "edit plan mid-flight" feature lands,
   widen explicitly with test coverage at that time.

2. needs_team_match was set True only on claim/qa_pass/qa_fail/
   docs_complete. Defense-in-depth says every role-scoped task
   action should re-assert team match (don't rely on the inheritance
   chain through assigned_to alone). Flip to True on: start,
   set_plan, block, pause, submit_verification, submit_qa,
   submit_pm_review, complete, create_subtask. Leave False on
   board/CEO actions and PM cross-cell interventions (unblock,
   resume, cancel) where the cross-cell semantics are intentional.

3. claim.source_statuses is intentionally a SUPERSET of any single
   role's CLAIM_RULES allowance (the table holds the union; CLAIM_RULES
   holds the per-role authority). Add an inline comment above the
   claim ActionSpec so a future reader doesn't conclude the two
   tables disagree — they don't, they encode overlapping facts at
   different grains.

* feat(lifecycle): _INTENT_VERBS table — every gateway verb declared

Each gateway intent verb is now a named composition of atomic actions
plus optional side effects. i_will_work_on = (claim, set_plan, start);
i_am_done = (submit_verification, submit_qa); open_pr is pure side
effects (push_branch, create_pr); etc.

* fix(lifecycle): widen block.allowed_roles to include QA + Documenter

Task 6 review caught a role-set inconsistency: i_am_blocked.allowed_roles
admits dev/QA/doc, but the underlying block.allowed_roles only allowed
dev+PM. Result: a QA or documenter calling i_am_blocked would pass the
IntentSpec gate and then be rejected by the composed ActionSpec gate
when Task 7 wires can_invoke_intent.

Widen block to include QA + Documenter. The semantic case is sound: a
QA reviewing a task can discover an external blocker; a documenter
writing docs may need PM intervention. Predecessor PERMISSIONS.md
restricted block to dev+PM, but with the gateway exposing i_am_blocked
to all worker roles, the underlying atomic must agree.

The deeper unclaim/escalate_up "imperative verb" concern from the same
review (composes=() but mutates state) is deferred to Task 8 where the
validator design lands.

* feat(lifecycle): public lookup functions + Context + preconditions

can_claim, can_invoke_action, can_invoke_intent, valid_next_verbs,
composed_actions_for, intents_for_role, status_after — the entire
public surface every consumer will use. Context carries the
caller-supplied state preconditions need (plan, journal-decision
flag, etc.). Preconditions for plan/commits/no_pr/ownership are
declared once and wired into the relevant IntentSpecs.

* fix(lifecycle): wire PRECONDITION_OWNERSHIP through Context.actor_id

Task 7 review found _p_owns_task reads agent.id but every call site
passes None for the agent arg. Result: getattr(None, "id", object())
returns a fresh sentinel, task.assigned_to == <sentinel> is always
False, and open_pr / i_am_done would reject every owner the moment
Task 9 wires consumers.

Fix: thread identity through Context.actor_id (new UUID field) and
rewrite _p_owns_task to read from the context. Both call sites already
pass the Context — no signature changes elsewhere. Add green-path
test exercising the owner-can-open-pr case the existing tests
missed (the Task 7 plan only tested precondition-failure paths,
which masked the bug).

Plus surface hygiene: STATUS_GRAPH, CLAIM_RULES, ROLE_TEAM_RULES,
and the four PRECONDITION_* constants are now in
roboco.lifecycle.__init__.__all__ so consumers in Tasks 8/9 don't
depend on the implicit `from roboco.lifecycle.spec import ...`
backdoor.

* feat(lifecycle): import-time self-consistency validators

10 validators run at module import; first failure raises
LifecycleSpecError and prevents the package from loading. Covers
status enum coverage, reachability, terminal exits, intent
compositions, status chain consistency, claim-rule role/status
coverage, self-review symmetry, team-rule slug existence, and
StatusTransition action references.

* fix(lifecycle): close validator gaps; resolve BACKLOG-claim and submit_qa IN_PROGRESS-shortcut ambiguity

Three reviewer follow-ups on Task 8's _validate.py, plus two real
data corrections the new action-target-reachability validator
surfaced.

1. Design spec §9 calls for "every ActionSpec.target_status, when
   set, is reachable from each source_status via STATUS_GRAPH" —
   missing from Task 8's 10 validators. Add
   _check_action_target_reachable_from_source.

2. _check_role_team_rules_slugs verified slug existence in
   AGENT_UUIDS but NOT that the cell team in ROLE_TEAM_RULES
   matches the seed. Add _check_role_team_rules_team_match,
   scoped to non-None entries only — None means "exempt from
   team-match enforcement" (cross-cell roles), not "no team in
   org chart".

3. test_validators_pass_on_real_spec was ceremonial. Add
   test_run_all_validators_raises_on_unknown_intent_action,
   a deliberate-break regression that monkeypatches _INTENT_VERBS
   to inject a fake action and asserts LifecycleSpecError raises.

The new action-target-reachability validator caught two real
data inconsistencies between the predecessor canon docs and the
spec tables:

A. claim.source_statuses listed BACKLOG and CLAIM_RULES[*PM]
   listed BACKLOG, but STATUS_GRAPH[BACKLOG] = {PENDING, CANCELLED}
   only. Resolution: PMs use the explicit \`activate\` action to
   move BACKLOG → PENDING, then claim from PENDING. Drop BACKLOG
   from claim.source_statuses and CLAIM_RULES.

B. submit_qa.source_statuses listed IN_PROGRESS, but
   STATUS_GRAPH[IN_PROGRESS] does NOT include AWAITING_QA. The
   intent verb i_am_done composes (submit_verification, submit_qa)
   which forces IN_PROGRESS → VERIFYING → AWAITING_QA — no
   shortcut. Drop the stale IN_PROGRESS entry from
   submit_qa.source_statuses.

Both corrections tighten the canonical state machine to a strict
no-skip transition graph. Pre-gateway PERMISSIONS.md/STATUS_TRANSITIONS.md
disagreements are resolved here; spec.py is the canon now.

* feat(gateway): Envelope.from_decision maps lifecycle Decisions to envelopes

Single shape adapter so verb bodies stop hand-composing rejection
envelopes. Each rejection_kind maps to a specific envelope flavor;
'self_review' folds into 'not_authorized' with a parenthetical hint;
constructing from an allow Decision raises (programmer error).

* feat(gateway): VerbRunner for atomic composed-action dispatch

Wraps spec.composed_actions_for(intent) in session.begin_nested()
so mid-sequence failures roll the DB back. Side effects run AFTER
the savepoint commits. Each atomic action name dispatches to a
TaskService method via a single, exhaustive _dispatch_atomic
mapping. New verbs slot in by adding an IntentSpec entry + a
_dispatch_atomic case if a new atomic is needed.

* refactor(gateway): i_will_work_on uses spec.can_invoke_intent + VerbRunner

Replace the bespoke status-branch dispatcher in i_will_work_on with the
spec-driven flow: load task -> load agent -> build spec.Context ->
spec.can_invoke_intent (and spec.can_claim for per-role status authority)
-> Envelope.from_decision on rejection -> VerbRunner.run_intent on success.

The _i_will_work_on_pending, _i_will_work_on_claimed,
_i_will_work_on_needs_revision, and _start_failed_envelope helpers are
removed; the runner replaces them. Two narrow verb-body re-entry blocks
remain for behaviors the spec does not yet model:

  1. in_progress + same agent -> idempotent heartbeat-only return
  2. claimed + same agent -> _resume_from_claimed (set_plan + start)
     to recover from a stuck mid-claim crash without re-running claim
     against a state the spec excludes.

The behavioral claim guards (already_active / paused / sibling_sequence)
also stay imperative for now -- they're not in the spec yet and migrate
into spec.extra_preconditions in a later task. Per-role claim authority
is enforced via spec.can_claim because the atomic claim action's
source_statuses are the union across roles; CLAIM_RULES narrows.

Parity test in tests/lifecycle/test_consumer_parity.py runs the verb
against every (role x status x task_type='code') combo (112 rows) and
asserts the envelope error matches the spec's Decision (or can_claim's
Decision when the intent gate passes but per-role claim authority does
not). This is the contract that makes spec/verb drift impossible.

Existing tests updated where rejection-message text changed (the spec
now produces the messages, e.g. "role 'cell_pm' may not call
'i_will_work_on'" instead of "PM cannot execute code") or where the
spec's stricter view ("invalid_state" -> "not_authorized" for a dev
trying to claim awaiting_qa) is more accurate. Test fixtures were
updated to wire task.session.begin_nested as a proper async context
manager (required by VerbRunner) and to set agent_for().id so runner-
driven calls line up with assert_awaited_with(task_id, agent_id).

* refactor(lifecycle): push CLAIM_RULES enforcement into can_invoke_action

Task 11's i_will_work_on migration had to call spec.can_claim()
separately after spec.can_invoke_intent() because the claim action's
source_statuses is the union across all claim-eligible roles —
can_invoke_intent alone would let a developer pass for claiming
awaiting_qa (a QA-only state).

The retrofit pattern would repeat in every claim-composing verb
(i_will_plan, claim_review, claim_doc_task). Push the per-role
narrowing inside can_invoke_action when the action is "claim",
using the same not_authorized vs invalid_state disambiguation
can_claim already implemented (status-reserved-for-another-role
returns not_authorized; status-no-role-can-claim returns
invalid_state). Extracted the body to _check_claim_rules_narrow
to keep can_invoke_action under xenon's complexity threshold.

Update _i_will_work_on_gate to drop the redundant spec.can_claim
call. Update test_consumer_parity.py to assert only against
can_invoke_intent's Decision.

Tasks 12-22 will inherit the cleaner pattern: spec.can_invoke_intent
is the single gate; verb bodies don't need per-action retrofits.

* refactor(gateway): i_will_plan uses spec.can_invoke_intent + VerbRunner

Migrates i_will_plan to the spec-driven pattern Task 11 set up for
i_will_work_on. The verb body now: (1) loads task + agent, (2) builds
Context, (3) checks idempotent/recovery re-entry, (4) calls
spec.can_invoke_intent, (5) returns Envelope.from_decision on
rejection, (6) delegates composition to VerbRunner. The
_i_will_plan_* helpers are removed — the runner replaces them.

Parity test in tests/lifecycle/test_consumer_parity.py runs the verb
against every (role × status × task_type) combo and asserts the
envelope matches spec.Decision.

* refactor(gateway): delegate uses spec.can_invoke_intent for role/state gate

Migrates delegate to the spec-driven role/state gate. The chain
validation (main_pm->cell_pm, cell_pm->its team's devs), the
assignee-vs-task_type rule (Cell PMs receive planning-typed only),
the enum coercion, and the parent-lifecycle/cap guards STAY in the
verb body — they encode delegate-specific semantics the spec
doesn't model.

Parity test in tests/lifecycle/test_consumer_parity.py asserts the
spec's role+state rejection is correctly surfaced. Chain/assignee
rejections continue to be tested in test_choreographer_pm_extras.

* refactor(gateway): open_pr uses spec.can_invoke_intent + VerbRunner

Migrates open_pr to spec-driven gating. The spec's
extra_preconditions (PRECONDITION_OWNERSHIP, PRECONDITION_COMMITS,
PRECONDITION_NO_PR) handle all three precondition checks; the verb
body delegates side-effect dispatch (push_branch, create_pr) to
VerbRunner.

Idempotent re-entry retained: an open_pr call against a task that
already has a PR (and the caller owns it) returns OK without
re-opening, rather than the tracing_gap the spec would otherwise
produce. This preserves agent ergonomics — two calls in a row
shouldn't surface a misleading "no_prior_pr" hint.

Parity test in tests/lifecycle/test_consumer_parity.py runs the verb
against representative (status x commits x pr_number) combos and
asserts the envelope matches spec.Decision.

* refactor(gateway): i_am_done uses spec.can_invoke_intent + VerbRunner

Migrates i_am_done to spec-driven gating. The spec's
extra_preconditions (PRECONDITION_OWNERSHIP, PRECONDITION_COMMITS)
handle ownership and commit-count checks; VerbRunner dispatches
the (submit_verification, submit_qa) atomic chain.

The tracing-gate preconditions (progress entry, journal:reflect,
acceptance criteria) and the field-level submit-qa gates stay in
the verb body — they model gates the spec doesn't yet cover.
Defense-in-depth: those gates run after the spec accepts the
ownership/commits checks.

Parity test in tests/lifecycle/test_consumer_parity.py runs the
verb against (role × status × ownership × commits) and asserts
the envelope matches spec.Decision.

* refactor(gateway): i_am_blocked uses spec.can_invoke_intent + VerbRunner

Migrates i_am_blocked to spec-driven gating. The journal:struggle
write stays in the verb body (it's a side effect outside the
lifecycle action). VerbRunner dispatches the `block` atomic action
via task_service.escalate.

Parity test in tests/lifecycle/test_consumer_parity.py.

* refactor(gateway): unclaim and resume use spec.can_invoke_intent

Migrates both verbs to the spec-driven gate. unclaim's verb body
keeps its dispatch (task.unclaim_for_agent) because composes=();
resume goes through VerbRunner with composes=("resume",).

The reassignment-rejection branch (introduced in 19f27b4 for the
2026-05-08 trace's "not your claim" case) stays - the spec doesn't
model "task got reassigned out from under you by an upstream verb,"
and the existing envelope text ("current owner: X - call
give_me_work() to find your current work") is the load-bearing
hint that fixed the original bug. Extracted the shared branch into
_reassigned_rejection / _ReassignedCtx so both verbs reuse it
without duplicating the envelope construction.

Parity tests in tests/lifecycle/test_consumer_parity.py.

* refactor(gateway): complete uses spec.can_invoke_intent at the dispatcher

Migrates the top-level `complete` dispatcher to gate role/state via
spec.can_invoke_intent before routing to cell_pm_complete or
main_pm_complete. The two lower-level methods keep their existing
PR-merge / CEO-escalation logic and pre-flight guards (those model
journal:decision preconditions and PR-mergeability checks the spec
doesn't model yet).

The runner pattern is NOT applied here — `complete` has two divergent
runtime paths (Cell PM merges leaf into parent branch; Main PM opens
master PR + escalates to CEO) that don't fit the runner's
single-composition model. Verb-body-owns-dispatch is the right
pattern.

Parity test in tests/lifecycle/test_consumer_parity.py runs the
verb against (role × status) combos and asserts the dispatcher's
spec rejection is correctly surfaced.

* refactor(gateway): escalate_up, escalate_to_ceo, submit_up use spec.can_invoke_intent

Migrates the three PM-side escalation/submission verbs to
spec-driven role/state gating. The verb-specific guards
(journal:decision, escalation_target configured, _submit_up_guard's
ownership + notes-length + subtasks-terminal) STAY in the verb body
- the spec doesn't model these.

escalate_up has composes=() so the verb body owns dispatch via
task.escalate. escalate_to_ceo and submit_up route their
compositions through VerbRunner.

Parity tests in tests/lifecycle/test_consumer_parity.py.

* refactor(gateway): qa.py + doc.py role mixins use spec.can_invoke_intent

Migrates the five QA + Documenter verbs (claim_review, pass_review,
fail_review, claim_doc_task, i_documented) to spec-driven gating.

The self-review block lives at the atomic-action layer
(_ATOMIC_ACTIONS["qa_pass"|"qa_fail"|"docs_complete"].self_review_block=True)
and naturally fires when the verb body builds a Context with
actor_slug==original_developer_slug. No verb-body retrofits needed.

The verb-specific helpers (_verify_qa_owner, _qa_pass_gate_check,
_check_i_documented_inputs) STAY — they encode notes-length /
journal:learning / files-list / qa_evidence_inspected gates the
spec doesn't model.

claim_review and claim_doc_task own dispatch via task.qa_claim /
task.doc_claim respectively (not the runner) because those
specialized claim methods keep status at AWAITING_QA /
AWAITING_DOCUMENTATION, which is what the downstream qa_pass /
qa_fail / docs_complete source-status requirement expects.
The spec gate still validates role + claim source-status + task_type
before dispatch.

pass_review / fail_review / i_documented route their compositions
(qa_pass / qa_fail / docs_complete) through VerbRunner.run_intent
inside a savepoint.

Parity tests in tests/lifecycle/test_consumer_parity.py for all
five verbs.

* fix(lifecycle): claim_review and claim_doc_task have empty composes

Tasks 21-22 surfaced a real spec/runtime mismatch: both verbs were
declared composes=("claim", "start"), but the actual implementation
uses task.qa_claim / task.doc_claim which intentionally keep status
at AWAITING_QA / AWAITING_DOCUMENTATION. If the runner ever ran the
declared composition, it would transition the task to CLAIMED then
IN_PROGRESS, breaking the source-status invariants of qa_pass,
qa_fail, and docs_complete.

The spec is the canon — align it to the runtime. composes=() means
"verb body owns dispatch" (same pattern as escalate_up and unclaim).
The spec gate still validates role + AWAITING_QA / AWAITING_
DOCUMENTATION source-status via the role's CLAIM_RULES narrowing,
enforced through special handling in can_invoke_intent, so role/state
safety is preserved.

* refactor(gateway): role_config flow lists derived from spec.intents_for_role

Hand-maintained _DEV_FLOW etc. tuples replaced with calls into the
spec. Adding/removing a role from an IntentSpec.allowed_roles now
automatically updates the MCP manifest. The spec is the canon;
role_config becomes a thin shim that adds the do-tool / write /
subagent / description metadata the spec doesn't carry.

* feat(lifecycle): generators + make lifecycle for deterministic artifact regen

Renders intent-verbs.md, status-transitions.md, panel/lib/lifecycle.json,
and per-role agents/prompts/_generated/lifecycle-{role}.md fragments
from the canonical spec. `make lifecycle` runs the regenerator;
deterministic output enables CI to gate on `git diff --exit-code` after
running it. The agent prompt fragments will be injected at the top of
each role's system prompt (Task 25) so agents see the same verbs the
gateway accepts.

* feat(lifecycle): inject generated prompt fragments + CI drift gate

Each agent's system prompt now starts with the spec-generated
'verbs available to your role' fragment. CI runs make lifecycle
and fails if regeneration produces a diff — drift between spec
and artifacts cannot land on master.

* refactor(gateway): delete verb_gates.py — superseded by lifecycle.spec

verb_gates.is_verb_allowed and verb_gates.valid_next_verbs are now
spec.can_invoke_intent(...).allowed and spec.valid_next_verbs.
Importers updated to consume the canonical spec module directly.
tests/unit/gateway/test_verb_gates.py removed — coverage lives in
tests/lifecycle/test_spec.py.

envelope.with_introspection wraps spec.valid_next_verbs with role-string
coercion + best-effort try/except so malformed task fixtures (AsyncMock
status) and unknown role strings still yield [] instead of raising —
preserves the legacy verb_gates contract.

content_actions content-tool RBAC (commit/notify) is now a pair of
explicit role frozensets in this file. These are content tools, not
lifecycle intents, so they intentionally do NOT live in spec._INTENT_VERBS.

Two existing introspection tests asserted "commit" in valid_next_verbs;
fixed to assert open_pr/i_am_done — commit is correctly absent under
the canonical spec because it is a do-server content tool, not a flow
intent verb.

* refactor(gateway): collapse scattered role constants into spec

The pm_cannot_execute_code_guard and role_typed_claim_guard guards
both modeled rules the spec now handles via can_invoke_action's
CLAIM_RULES narrowing and ActionSpec.allowed_task_types. Drop them
from claim_guards.py — the choreographer's existing skip-flags on
_run_claim_guards are now permanent: those guards no longer fire.
Simplify _run_claim_guards's signature accordingly.

The concurrency-invariant guards (already_active_guard,
paused_tasks_guard, sibling_sequence_guard) STAY — the spec doesn't
model these system-level invariants. sibling_sequence_guard's loop
body extracted into _earlier_blocking_sibling helper to keep the
slimmed module under xenon's --max-modules A average.

* refactor(enforcement): task_lifecycle becomes a thin view of lifecycle.spec

VALID_TRANSITIONS and ROLE_RESTRICTED_TRANSITIONS are now derived
from roboco.lifecycle.spec — no independent tables. The 433-line
file collapses to ~30 lines of view definitions; future changes
go in spec.py. Helper functions exported by the legacy module are
preserved as thin wrappers so existing consumers don't need to
change their imports today.

A small _LEGACY_OPERATIONAL_EDGES table sits alongside the
spec-derived view to cover transitions the runtime exercises but
the spec has not yet absorbed (voluntary unclaim, reaper sweep,
PM-direct completes from in_progress, parallel-doc-PR developer
trigger). It is fenced and clearly documented; once those callers
are migrated to spec-driven dispatch the constant goes empty and
the file collapses to a pure view.

A test in test_task_service_lifecycle_misc.py was rewritten: the
predecessor asserted CEO-only authority over awaiting_ceo_approval
cancels (legacy table behavior), but the canonical spec authorizes
{CELL_PM, MAIN_PM, CEO} uniformly across all non-terminal cancel
sources. The test now exercises the broader spec-defined cascade.

* feat(lifecycle): UNMIGRATED guard pins known-debt consumers

Two pieces of debt surfaced during Task 28's collapse of
enforcement/task_lifecycle.py: (1) ~11 operational edges still in
the shim's _LEGACY_OPERATIONAL_EDGES because the spec's
_STATUS_TRANSITIONS doesn't yet model them; (2) role-gate
disagreements in _LEGACY_ROLE_GATES that the spec disagrees with.

UNMIGRATED is the named-debt set; KNOWN_UNMIGRATED_CONSUMERS pins
the catalog so a contributor adding a new entry must update both
sides. Validator (_check_unmigrated_is_subset) fires at import if
they drift. Test pins the current entries.

Phase 3's terminal invariant is `UNMIGRATED == frozenset()` —
expected when both legacy data carriers fold into spec, at which
point the assertion becomes a permanent regression guard.

* test(lifecycle): tier 3 end-to-end real-DB happy paths

Eight integration tests covering every major lifecycle path:
dev (pending → awaiting_qa), QA pass, QA fail, doc handoff,
Cell PM complete, Main PM escalate-to-CEO, block+unblock,
pause+resume. Each test drives the spec → choreographer →
TaskService → DB stack with only the git layer mocked. Catches
"spec says X, DB constraint says Y" mismatches the unit-tier
parametrized parity suite cannot detect.

* test(lifecycle): tier 4 smoke replay — pin known-bug shapes after spec migration

Synthesized fixture covering the 9 bugs from the 2026-05-08
audit-log trace + the 2 from the 2026-05-09 follow-up trace. Each
record documents (verb, role, task setup, expected post-fix
envelope shape, fix commit, spec invariant). The replay test
parametrizes over the records and asserts the spec / choreographer
behavior now matches the post-fix expectation — locks in the
fixes as permanent regressions.

The original audit log was wiped during cleanup; the fixture is
a documented synthesis, not a verbatim capture. The bug list is
faithful to the prior session's analysis of the trace.

* fix(orchestrator): silence dev-dispatcher noise for non-dev-lane tasks

Dev dispatcher fetched all pending/claimed/in_progress tasks regardless
of assignee role and warned 'role/task_type mismatch' on each pass when
it found cell_pm/main_pm/product_owner/etc. tasks — those belong to
_dispatch_pm_work, not this lane. The 30s warning loop showed up
prominently in the 2026-05-10 smoke run.

Filter at the lane boundary: silently skip when assignee role is not
developer/documenter/unknown. The D-49 misassignment warning still
fires for the legitimate cases (developer assigned a documentation
task, etc.).

* fix(gateway,prompts): unblock the three smoke-run dead-ends

Three issues surfaced by the 2026-05-10 smoke run, fixed together
because they're all blockers for end-to-end task completion:

1. Acceptance-criteria tracing gate was unsatisfiable. Nothing in the
   codebase writes to task.acceptance_criteria_status, so
   _check_acceptance_criteria always returned every criterion as
   missing. Treat a reflect note as the addressing artifact: when the
   agent has written one, the gate clears. Per-criterion citation via
   acceptance_criteria_status is still honored when populated, so the
   schema stays available for future per-criterion tracking.

2. Cell PM runaway re-decomposition. On every wake-up be-pm
   re-decomposed its parent task without checking for existing
   children, producing duplicate dev subtasks. cell_pm.md now teaches
   'list children before delegating' and 'one dev subtask is usually
   enough — QA/Documenter/PM-merge engage automatically'. Added
   anti-pattern entries for re-decomposition and over-decomposition.

3. Main PM exit/respawn loop on claimed-state tasks. The model
   cycled through delegate/resume/escalate/unblock looking for a verb
   that worked on 'claimed', and got cleanly rejected by every one.
   The right verb is i_will_plan (it composes claim+set_plan+start
   and resumes from claimed). main_pm.md now spells this out
   explicitly with a worked example of which verbs reject and why.

* feat(prompts): restore pre-gateway lifecycle scaffolding across all 6 roles

The gateway migration shrank role prompts from ~50 lines to ~15
(commit 534152c for dev; analogous shrinks for qa/doc/cell_pm/main_pm/board
in e12a596, 05ac832, 8dc381b). The verb surface got cleaner but the
prescription for using verbs through the lifecycle disappeared. The
2026-05-10 smoke run surfaced the regression: agents thrash through
verbs hoping one fits, journal sparsely, skip the dev reflect note,
and (for cell PMs) re-decompose on every wake-up.

Each role prompt now restores three sections that the pre-gateway
versions had:

1. State -> Verb table — what to call when respawned in each
   lifecycle status. Eliminates the verb-cycling antipattern: the
   agent looks up its current status and calls the one verb that
   transitions out of it.

2. Mandatory pre-handoff checklist — explicit walk-through of the
   gates the next verb will check, ordered so the agent fixes the
   missing piece before retrying:
   - developer: 7 items before i_am_done
   - qa: 8 items before pass/fail (incl. self-review forbidden,
     read dev journal not just diff, name artifact per criterion)
   - doc: 7 items before i_documented
   - cell_pm: 7 items before submit_up (incl. integration green)
   - main_pm: 7 items before complete(root)
   - board: separate checklists for escalate_to_ceo (PO/HoM) and
     reflect-note quality (Auditor — its only output)

3. Journaling cadence — when to use each of the five scopes
   (note/decision/struggle/learning/reflect). The pre-gateway
   prompts named all five scopes with role-specific examples;
   the post-gateway prompts mention 'reflect' once and skip the
   rest. Restored across every role.

Plus restored the load-bearing rules that got dropped:

- Cell PM: 'A SINGLE subtask flows through dev -> QA -> doc ->
  PM-merge. DON'T split into per-role subtasks.' This is exactly
  what be-pm violated in the smoke run, creating duplicate
  'branch naming subtask' / 'PR workflow subtask' / etc.
- QA + Doc: 'read the dev's journal, not just the diff' — pre-
  gateway forced this via roboco_journal_read_team; post-gateway
  the inline data exists but the agent isn't told to use it.
- Developer: 'every acceptance criterion gets a citation in the
  reflect note' — pairs with the tracing-gate change in 75b667d
  where the reflect note is treated as the addressing artifact.

* feat(foundation): bootstrap foundation/identity.py with Role/Team/RoleLevel

Phase 1 task 1 of the foundation canonicalization plan
(docs/superpowers/specs/2026-05-10-foundation-canonicalization-design.md).

Three enums, no consumers yet — separate tasks migrate the existing
forks (models.base.AgentRole, lifecycle.spec.Role, agents_config role
sets, services/permissions.PM_ROLES) onto this canonical surface.

* feat(foundation/identity): add AGENTS catalog (single source for slug->role+team+UUID)

Resolves head-marketing.team drift (spec §5.1) by setting Team.BOARD
authoritatively. Team.MARKETING remains in the enum for legacy seed
data but no agent claims it; flagged for removal in cleanup.

* feat(foundation/identity): add role-sets + ROLE_LEVEL hierarchy

* feat(foundation/identity): add lookups + public API re-exports

* feat(foundation): import-time validators (uniqueness, role coverage, role-level)

* chore(foundation): verify+align postgres agentrole/team enums with foundation/identity

scripts/verify_postgres_enums.py reads the live agentrole+team enums
from postgres (via asyncpg using roboco.config.settings.database_*)
and compares them against the foundation Role+Team enums. Exits 0 on
match, 1 on drift (with a per-side diff), and 1 with a clear message
if postgres is unreachable so callers like make foundation-check can
treat that as a skip.

alembic/versions/012_align_agentrole_team_with_foundation.py is the
forward-only safety-net migration. It runs ALTER TYPE agentrole ADD
VALUE IF NOT EXISTS 'system' (idempotent on postgres >= 9.6) so any
DB without the recently-added Role.SYSTEM sentinel gets it on next
upgrade. Postgres has no DROP VALUE primitive without a destructive
type recreation, so foundation keeps legacy values (e.g. Team.MARKETING)
to absorb the inverse direction; the migration's downgrade is
intentionally a no-op.

Local verification deferred: postgres is not reachable from this
workstation (role 'roboco' does not exist), so the script could not
confirm the live enum shape. The migration is idempotent and runs
unconditionally on the next alembic upgrade head, and whoever next
runs make foundation-check against a live DB will get the post-migration
proof of alignment.

* refactor(lifecycle): re-export Role from foundation.identity (single source)

* refactor(models): re-export AgentRole and Team from foundation.identity

Removes the parallel Team and AgentRole StrEnum definitions in
models/base.py. They are now bound to roboco.foundation.identity.Role
and roboco.foundation.identity.Team respectively, so AgentRole IS
identity.Role (same Python class object). SQLAlchemy column types
bound as sa.Enum(AgentRole, name='agentrole') continue to work because
identity is preserved across import paths.

Note: foundation.Team drops the legacy 'fullstack' member that lived
on models.base.Team. The two _resolve_team_dir tests that used
Team.FULLSTACK to exercise the 'fullstack' branch now pass the literal
string 'fullstack' instead — same code path, no enum-membership coupling.

Adds two identity assertions to tests/foundation/test_role_reexport.py
verifying AgentRole is identity.Role and Team is identity.Team.

* fix(foundation): correct Team enum — add FULLSTACK, remove QA

The original plan's audit incorrectly identified the models/base.Team
membership. Actual original was 7 values: backend, frontend, ux_ui,
fullstack, main_pm, board, marketing. My plan replaced fullstack with
qa and added system — but qa was never a team (only a role).

Postgres team enum has fullstack (alembic 009), and services/task.py:675
+ services/git.py:779 branch on the literal "fullstack". Without
foundation.Team.FULLSTACK, any Project row with assigned_cell="fullstack"
would fail to round-trip through the SQLAlchemy ORM.

This correction:
- Adds FULLSTACK; removes QA from foundation.Team
- Updates the 8-value test expected set
- Restores tests/integration/test_task_service_misc.py to use Team.FULLSTACK

* refactor(agents_config): derive AGENT_ROLE_MAP/AGENT_TEAM_MAP/CELL_MEMBERS from foundation

* refactor(roles): canonicalize role-sets via foundation.identity

- agents_config.PM_ROLES (5-role: PMs + board + CEO) renamed to
  TASK_CREATOR_ROLES; the name PM_ROLES is reserved for the canonical
  2-role set (CELL_PM + MAIN_PM) defined in foundation.identity.
- agents_config._BOARD_ROLES aliased to foundation.BOARD_ROLES (drops
  main_pm from the set; board A2A handler updated to keep allowing
  board -> main_pm direct messaging via explicit branch).
- services/permissions.PM_ROLES (2-role) re-exported from foundation.

Closes the silent semantic divergence flagged in spec section 3 (HIGH severity).

* refactor(seeds,orchestrator): derive agent catalogs from foundation

- seeds/initial_data.AGENT_UUIDS derived from foundation.AGENTS.
- DEFAULT_AGENTS row generation pulls slug+role+team+id from foundation;
  per-agent presentation strings (display name) stay in this file in
  _AGENT_PRESENTATION dict. The system sentinel remains a literal with
  team=None because the postgres `team` enum has no 'system' value.
- runtime/orchestrator._AGENT_TEAM_MAP and the cell-prefix table replaced
  with foundation.team_for_slug. _AGENT_TEAM_MAP is now a derived ClassVar
  covering every slug (not just management).
- head-marketing.team resolved to "board" (was "marketing" in seed +
  orchestrator, "board" in agents_config — three-way drift, now unified).
- ceo.team resolved to "board" (was None in seed; foundation declares
  board membership so the seed-bootstrapped DB row now reflects that).
- Adds tests/foundation/test_seed_orchestrator_parity.py — gate against
  future drift between seed/orchestrator and foundation.

Closes the identity sub-phase. Adding an agent edits exactly one file:
foundation/identity.py:AGENTS.

* feat(foundation/policy): task_completeness rules + denylist

Implements spec §5.2: field-level completeness rules at create/delegate
time, plus the denylist that catches the literal placeholder string from
the deleted services/task.py:5061-5062 silent fallback ("completed and
reviewed by assignee" — agents copy-paste this from old logs).

CompletenessSpec is data; check() is a pure function; field_hints map
gives the agent the literal answer key for each missing field.

* feat(envelope): add incomplete_input envelope kind for interrogation pattern

Sister to tracing_gap; distinct error code lets agent prompts teach
incomplete_input handling separately from tracing-gap recovery. Carries
missing + field_hints + remediate for the spec §5.2.1 interrogation
pattern; Task 19 will wire the gateway delegate verb to use it.

* feat(foundation/task_completeness): auto-fill helpers (team, priority, parent)

* feat(api/schemas): DelegateRequest enforces TASK_AT_CREATE constraints

Removes silent defaults for nature/task_type/estimated_complexity;
adds min_length=20 to description; requires non-empty acceptance_criteria.
Mirrors foundation.policy.task_completeness.TASK_AT_CREATE so under-filled
delegate calls fail at the request boundary (422) instead of being silently
papered over downstream.

Tests touching delegate calls updated to pass the now-required fields.

* feat(models/task): TaskCreate + TaskCreateRequest enforce TASK_AT_CREATE

* feat(api/schemas): TaskUpdate rejects blanking acceptance_criteria

Golden Rule preservation — acceptance_criteria cannot be set to []/None
via PATCH. Pydantic field min_length doesn't catch explicit None, so a
model_validator(mode='before') guards the patch payload.

* fix(services/task): delete silent acceptance_criteria fallback (skeleton-task root cause)

The fallback at services/task.py:5061-5062 silently replaced empty
acceptance_criteria with ['completed and reviewed by assignee'] -
the proximate cause of every skeleton task in the 2026-05-10 smoke
run. Removed; create_subtask now invokes foundation.policy.task_completeness
and raises TaskCompletenessError on missing fields (spec section 5.2).

Two existing transition tests relied on a 1-char description default
that the new completeness check rejects (description min_length=20);
both updated to pass an explicit valid description. New integration
test pins the rejection contract (empty list + legacy phrase both
raise).

Companion code path at services/gateway/choreographer/_impl.py:1852
(the upstream `or []` collapse) is fixed in the next task.

* fix(gateway/delegate): use task_completeness + Envelope.incomplete_input

Replaces the `acceptance_criteria=inputs.acceptance_criteria or []`
collapse at _impl.py:1852 with a foundation.policy.task_completeness
check that rejects empty / placeholder input via
Envelope.incomplete_input — the spec section 5.2.1 interrogation pattern.

Auto-fill helpers (fill_team_from_assignee + fill_priority_from_parent)
fill the unambiguous fields before the check, then anything still
missing surfaces as a structured rejection with field_hints; the agent
gets a literal answer key for what to provide on retry.

DelegateInputs gains an explicit `nature` field (no default) so the
HTTP boundary can thread DelegateRequest.nature through to the
choreographer. Route handlers (flow_cell_pm, flow_main_pm) forward it.
The hardcoded TaskNature.TECHNICAL fallback in _create_subtask_from_inputs
is removed; the helper now coerces inputs.nature to the enum or raises
TaskCompletenessError if a non-gateway caller bypassed the check.

Closes the gateway-side path to skeleton tasks. Service-layer raise
(Task 18) remains as defense-in-depth for non-gateway callers.

Existing delegate-guard tests updated to pass full payloads — the
prior `title='x', description='y'` minimal stubs now hit the
completeness gate first; the full payloads still exercise the
auth/chain/cap guards downstream.

* feat(api/routes/tasks): POST /tasks uses foundation.task_completeness check

Replace the hand-rolled acceptance_criteria non-empty check in the
POST /tasks handler with a call to task_completeness.check(TASK_AT_CREATE,
data). Route, schema (TaskCreate), and service (TaskCreateRequest) now
all share one canonical notion of 'complete' — the fourth and final
create path is now strict.

Pydantic still rejects structurally invalid payloads (empty AC list,
short title/description, missing enums) with 422. The TC check at the
route boundary additionally rejects denylisted placeholder phrases
('completed and reviewed by assignee', etc.) that pass schema validation
but signal a stub task.

Add tests/integration/test_post_tasks_completeness.py:
  - empty acceptance_criteria  -> 422 (Pydantic)
  - placeholder phrase         -> 400/422 with 'acceptance_criteria' in body

* chore(make): add foundation-check drift gate (mirrors lifecycle-check)

* test(foundation): Phase 1 smoke gate — skeleton-task path returns incomplete_input

Phase 1 closes here: identity catalogs are single-sourced; the silent
acceptance_criteria fallback is gone; gateway delegate returns
incomplete_input with populated field_hints when criteria are missing.

The 2026-05-10 smoke run that produced skeleton tasks no longer can.
Phases 2-4 (tracing, journaling, communications, agent_loop, housekeeping)
get their own plans.

* feat(foundation/policy): journaling scope catalog (5 panel-UI scopes)

* feat(foundation/policy/journaling): role read tiers + protected journals

* refactor(content_actions): derive _VALID_NOTE_SCOPES from foundation.journaling

* refactor(services/journal): derive _SCOPE_TO_TYPE from foundation.journaling

* refactor(enforcement/journal_perms): import read-tier rules from foundation

PROTECTED_JOURNALS + ROLE_READ_TIERS now sourced from foundation.policy.journaling.
The local helpers (_check_protected_access, _check_cell_pm_access,
_check_cell_member_access) are collapsed into a single tier-driven check via
_decide_protected / _decide_by_tier. Pre-Phase-2 GLOBAL_READERS that lumped
CEO/auditor/PO/HoM/main_pm together is split into ReadTier.ALL (ceo+auditor —
includes protected) vs ReadTier.ALL_CELLS (others — excludes protected).
Observable behavior preserved.

* feat(foundation/policy): tracing Requirement enum + check_requirements

19 requirements (16 from pre-Phase-2 tracing_gate + 3 pre-gateway parity:
JOURNAL_NOTE_AT_CLAIM, JOURNAL_DECISION_AT_CLAIM, JOURNAL_DURING_WORK).
GateContext expanded with the new presence flags and journal_during_work_count.
Acceptance-criteria checker keeps the spec §9 item 1 reflect-note shortcut.

* feat(foundation/policy/tracing): VERB_REQUIREMENTS table + verb parity validator

Maps every gateway intent verb to its required-set. Includes the 6 inline
journal:decision callsites (submit_up, complete, unblock, escalate_up,
escalate_to_ceo, delegate) plus the 4 pre-gateway parity additions
(NOTE_AT_CLAIM, DECISION_AT_CLAIM, REFLECT on complete, DURING_WORK).
Validator asserts every spec verb is covered or explicitly waived, and
every Requirement enum value is used by at least one verb.

PLAN added to i_will_work_on / i_will_plan (mirrors spec.PRECONDITION_PLAN
in the tracing layer). SELF_VERIFIED added to i_am_done as a defense-in-depth
backstop (auto-set by the in_progress→verifying transition).

* refactor(gateway/i_am_done): tracing gates via foundation.policy.tracing

Adds JOURNAL_DURING_WORK_AT_LEAST_ONE check (pre-gateway parity P2 —
agents must write at least one decision/learning/struggle entry between
claim and submit). Adds journal.has_struggle_for_task helper.
Replaces the pre-Phase-2 tracing_gate.check_requirements call.

SELF_VERIFIED is filtered from the pre-flight required-set: the spec
composes (submit_verification, submit_qa) for i_am_done and the
auto-run submit_verification flips self_verified=True before submit_qa
runs. The flag therefore acts as a defense-in-depth backstop AFTER the
spec, not before — checking it pre-flight would block the auto-verify
path. SELF_VERIFIED stays in the foundation required-set and is
re-asserted by the spec action's own preconditions.

Test fixtures updated: 9 i_am_done success-path tests now mock
has_decision_for_task=True (or equivalent) so the new during-work
cadence gate is satisfied. NO_PR-token assertion broadened to also
accept the foundation token "pr_open".

* refactor(gateway/qa): pass/fail gates via foundation.policy.tracing

* refactor(gateway/doc): i_documented gates via foundation.policy.tracing

Doc-specific missing-key translations (docs_notes>=min, docs_files_non_empty)
added to the central _build_tracing_gap translator established in Task 9.

* refactor(gateway): unify 6 inline journal:decision checks via tracing.check_requirements

Pre-Phase-2 inline blocks at _impl.py lines ~2230/2394/2442/2574/2814/2895
each ran the same has_decision_for_task + Envelope.tracing_gap pattern. They
now call:

- _check_pm_decision_required(verb, ...) — for unblock, escalate_up,
  escalate_to_ceo, delegate. Each declares only JOURNAL_DECISION in
  VERB_REQUIREMENTS, so a single helper consuming
  tracing.requirements_for(verb) suffices.
- _check_complete_gates — for cell_pm_complete and main_pm_complete.
  Consumes VERB_REQUIREMENTS["complete"] = JOURNAL_DECISION + JOURNAL_REFLECT
  + NOTES_MIN_CHARS. The inline _subtasks_not_terminal_envelope is kept
  because its remediation enumerates the non-terminal subtask ids — strictly
  richer than the foundation hint.
- _check_submit_up_gates — for submit_up. Consumes
  VERB_REQUIREMENTS["submit_up"] minus SUBTASKS_TERMINAL (deferred to the
  inline envelope for the same reason as complete).

Also adds the journal:decision tracing gate to the delegate verb
(VERB_REQUIREMENTS["delegate"] = {JOURNAL_DECISION}) — pre-gateway PM.md
required journal:decision before each delegate, but the gateway path had
not yet enforced it. Threaded into _delegate_extra_guards so the verb
body's return count stays under the lint cap.

_build_tracing_gap gains hint translations for journal:decision, notes>=min,
and subtasks_terminal. The body is refactored to a static dispatch table
+ acceptance-criteria batch handler so the branch count stays under the
lint cap.

PM-verb success-path tests updated to provide notes >= 20 chars (the new
NOTES_MIN_CHARS gate); has_reflect_for_task mocks added to a few tests
where they're now load-bearing (AsyncMock truthiness covers most).
6575 tests passing, mypy + ruff clean.

* feat(gateway/claim): require journal:note_at_claim and journal:decision_at_claim

Pre-gateway parity P1, P3: developers wrote a note (scope='note') on
every claim; PMs wrote a decision (scope='decision') on plan. Restored
via foundation.policy.tracing requirements wired through a new
_post_claim_journal_gate helper that runs AFTER the composed
(claim, set_plan, start) sequence completes.

Failed checks return tracing_gap with a remediate hint that tells the
agent to journal then retry. The claim itself stays — the agent
journals and re-issues the verb (idempotent re-entry shortcuts back
to OK once the entry is present).

Adds journal.has_note_for_task helper paralleling
has_decision/reflect/learning/struggle. The PLAN requirement is
filtered out of the post-claim check because spec.PRECONDITION_PLAN
already enforced it before the runner ran — re-asserting at the
tracing layer would emit a misleading hint.

Two new tests verify the gate fires for missing note/decision; existing
success-path tests already mock the journal service via AsyncMock
(returning truthy) so no regressions.

* test(foundation): Phase 2 smoke gate + tracing_gate.py deleted

Phase 2 closes here:
- foundation/policy/journaling.py owns the 5-scope catalog + read tiers
- foundation/policy/tracing.py owns Requirement enum + VERB_REQUIREMENTS
- 6 inline journal:decision checks replaced with unified helpers
- pre-gateway parity restored: NOTE_AT_CLAIM, DECISION_AT_CLAIM,
  DURING_WORK, REFLECT-on-complete
- services/gateway/tracing_gate.py deleted
- enforcement/journal_perms.py read-tier rules canonicalized

Smoke gate 2 enforces: no inline has_decision_for_task remains; every
intent verb has a tracing decision; tracing_gate module is gone.

* fix(foundation/task_completeness): align hint strings with actual enum values

_HINT_NATURE listed 5 values (technical | bugfix | feature | refactor | docs)
but TaskNature only has 2 (TECHNICAL / NON_TECHNICAL). _HINT_ESTIMATED_COMPLEXITY
listed "critical" which Complexity doesn't have. _HINT_TEAM omitted FULLSTACK
(real, used) and didn't note that MARKETING is legacy seed-data. _HINT_TASK_TYPE
was already correct.

Hints now reflect the actual enums in roboco/models/base.py and
roboco/foundation/identity.py — agents reading the gateway's incomplete_input
remediate envelopes will no longer be told to send values the enums reject.

Tests using nature="feature" (DelegateRequest's nature is `str`, not the
enum, so it accepted the fake value silently) updated to nature="technical"
so they exercise a real enum value end-to-end.

* fix(orchestrator): remove dead "critical" complexity branches

Complexity enum has only LOW / MEDIUM / HIGH — no CRITICAL value.
The three "critical" branches in dispatch logic at lines ~3032 / 3331 /
5049 were dead code (the comparison can never be true). Removed.

Surfaced during Phase 2 closeout when the foundation hint string was
audited against the actual enum.

* feat(foundation/policy/communications): Priority + NOTIFY_SENDER_ROLES + ACK_REQUIRED_BY_TYPE

* feat(foundation/policy/communications): CHANNELS catalog (channel topology)

* refactor(agents_config): derive CHANNEL_ACCESS from foundation.communications

* refactor(seeds): derive DEFAULT_CHANNELS / CHANNEL_MEMBERSHIPS from foundation

* refactor(content_actions): derive notify allowlist + priorities from foundation

Replaces _NOTIFY_ALLOWED_ROLES + _VALID_NOTIFY_PRIORITIES literals with
derivations from foundation.communications.NOTIFY_SENDER_ROLES + Priority.

Behavior change: pre-Phase-3 the literal frozenset {cell_pm, main_pm,
product_owner, head_marketing} excluded CEO. Foundation includes CEO
(per spec 5.5). The contradiction with agents_config.NOTIFICATION_PERMISSIONS
(which already granted CEO can_send=True) is now resolved.

* refactor(notification_delivery): requires_ack from foundation.ACK_REQUIRED_BY_TYPE

* refactor(enforcement,agents_config): delete dead notification policy

- enforcement/notification_perms.py deleted (dead at call-graph; only
  the enforcement/__init__.py re-export kept it reachable, and that
  re-export is gone too).
- agents_config.NOTIFICATION_PERMISSIONS dict deleted; agents_config
  .can_send_notifications now derives from
  foundation.policy.communications.NOTIFY_SENDER_ROLES (auditor
  correctly excluded — silent observer per spec §5.5).
- services/permissions.py: _can_role_send_notifications and
  can_agent_send_notifications now derive from NOTIFY_SENDER_ROLES;
  _get_notification_scope encodes the scope rule (cell/all/list)
  locally as a function-of-role and returns list[AgentRole] instead
  of list[slug]; can_notify list-scope branch updated to match.
- enforcement/__init__.py: removed the notification_perms re-export
  and the NotificationPermissionError, get_notification_scope,
  validate_notification_permission names from __all__.

Closes the spec §3 contradiction: gateway content_actions
._NOTIFY_ALLOWED_ROLES (Task 5) and the legacy
agents_config.NOTIFICATION_PERMISSIONS no longer disagree about
whether auditor may call notify(). Both now derive from
foundation.NOTIFY_SENDER_ROLES.

* fix(content_actions): runtime auditor guard in say/dm (defense in depth)

Closes the spec §5.5 gap where the auditor's silent role was enforced
ONLY by manifest exclusion. The manifest pre-filters the tool surface
exposed to the auditor agent, but if anything bypassed it, the auditor
could speak. The new runtime guard in ContentActions.say/dm refuses
with Envelope.not_authorized when the caller's role is "auditor",
regardless of how the call arrived.

* fix(a2a): pass Priority tristate end-to-end (was reduced to boolean)

Pre-Phase-3 path:
  request priority: str -> services/a2a.py reduces to urgent: bool
  -> services/notification.py maps bool back to NotificationPriority
This made Priority.HIGH unreachable through the A2A path.

After this fix the full tristate (NORMAL/HIGH/URGENT) survives end-to-end:

  * services/a2a.py:create_a2a_notification parses metadata["priority"]
    (preferred) or falls back to legacy metadata["urgent"] / config.urgent
    (URGENT-only). Unknown values fall back to NORMAL.
  * services/notification.py:send_a2a_notification now takes
    a2a_context["priority"] (NotificationPriority); a defensive bool/str
    coerce keeps legacy callers from crashing.
  * runtime/orchestrator.py:_build_a2a_prompt reads priority off the
    notification row (the source of truth) instead of a non-existent
    metadata.urgent and renders three tiers: URGENT bold, HIGH softer,
    NORMAL no prefix.

Cosmetic [URGENT] body/subject prefix stays urgent-only; HIGH gets no
prefix but is recorded as HIGH at the NotificationTable.priority column.

Tests:
  * 9 new tests in tests/integration/test_a2a_priority_tristate.py
    pinning the round-trip for HIGH/NORMAL/URGENT through both layers
    plus legacy-bool backcompat.
  * Updated tests/unit/services/test_notification.py::test_send_a2a_notification
    to the new priority= contract.

Closes the spec section 3 contradiction flagged in the audit.

* feat(foundation/policy): agent_loop BudgetPolicy + VERB_RETRY_LIMITS

* refactor(agent_sdk): import budget thresholds from foundation

* refactor(orchestrator): import _PM_RESPAWN_MAX_UNPRODUCTIVE from foundation

* fix(post-tool-budget-hook): exit 1 on loop-halt (was exit 0 / non-blocking)

Pre-Phase-3 the hook printed [Loop] and exit 0'd — agents could ignore it
and keep retrying. The 2026-05-10 smoke run showed i_am_done retried 5+
times within the global 150-tool budget, never hitting a real wall.

Now the hook reads the SDK response's loop_action field (sourced from
foundation.BudgetPolicy.loop_action; default "halt") and exits 1 to
deny the wrapping tool call when the rolling-window loop detector fires
AND loop_action is "halt". Operators can soften via env
ROBOCO_AGENT_LOOP_ACTION=warn for debugging.

Changes:
- BudgetStatus pydantic model: add loop_action: Literal["warn", "halt"]
  (default "halt") so the SDK response carries the policy.
- agent_sdk/server.py: read ROBOCO_AGENT_LOOP_ACTION env override on top
  of foundation default and surface it in _budget_snapshot().
- post-tool-budget-hook.sh: parse .loop_action, exit 1 to stderr when
  loop+halt; falls back to legacy warn-only print if the field is
  missing (older SDK / partial deploy).

* feat(agent_sdk): per-verb retry circuit breaker via foundation.VERB_RETRY_LIMITS

Pre-Phase-3 the gateway had no per-verb retry cap. The 2026-05-10 smoke
showed i_am_done retried 5+ times in 2 minutes within the global 150-tool
budget — the agent never hit a real wall.

Now the SDK tracks (verb, task_id) -> deque[timestamp] over a 60s sliding
window. When the count for a verb exceeds foundation.retry_limit_for(verb),
the next attempt receives Envelope.circuit_open with a remediate hint
pointing to i_am_blocked / i_am_idle as graceful exits.

Verbs in foundation.UNLIMITED_RETRY_VERBS (give_me_work, triage,
evidence, etc.) bypass the breaker. Only rejection envelopes
(tracing_gap, invalid_state, not_authorized, incomplete_input) feed
the counter — successful calls do not count.

Wire-up:
- Envelope.circuit_open classmethod + as_dict pass-through
- _SessionState.verb_attempts: defaultdict[(verb, task_id), deque[float]]
- Helpers _record_verb_attempt / _verb_attempt_count / _check_verb_circuit
- POST /verb/attempted: hook posts after a rejected gateway call;
  response carries breaker state + (when open) the wire-format
  Envelope.circuit_open dict the agent should surface to itself
- GET /verb/circuit_status: read-only state probe
- _state.reset() (also POST /budget/reset) wipes the tracker on spawn

* test(foundation): Phase 3 smoke gate + foundation-check extended

Phase 3 closes here:
- foundation/policy/communications.py owns Priority, NOTIFY_SENDER_ROLES,
  ACK_REQUIRED_BY_TYPE, ChannelSpec, CHANNELS, parse_priority
- foundation/policy/agent_loop.py owns BudgetPolicy, VERB_RETRY_LIMITS,
  UNLIMITED_RETRY_VERBS, retry_limit_for
- 6 channel topology fork sites collapsed to one source (CHANNELS)
- Notification sender contradiction closed (CEO included; auditor excluded)
- A2A urgency tristate restored (HIGH reachable end-to-end); A2A
  service now consumes parse_priority instead of inlining branches
  (also drops create_a2a_notification CC from C/13 to A/<10)
- Auditor silent role enforced at runtime in say/dm
- enforcement/notification_perms.py deleted (was dead code)
- 7 hand-set requires_ack callsites consolidated to ACK_REQUIRED_BY_TYPE
- post-tool-budget-hook.sh exits 1 on loop-halt
- Per-verb retry circuit breaker live in agent_sdk (60s sliding window)

make foundation-check now validates communications + tracing + journaling
+ identity drift in one command. make quality green.

* refactor(lifecycle): copy spec.py to foundation/policy/lifecycle.py + shim

Phase 4 Task 1 — relocates the canonical lifecycle spec next to its policy
siblings (task_completeness, tracing, journaling, communications, agent_loop).

The original roboco/lifecycle/spec.py is now an explicit re-export shim;
consumers continue to work unchanged. Subsequent Phase 4 tasks (2-7) migrate
the imports in batches, then Task 8 deletes the shim.

No behavior change — pure code move.

* refactor(services): import lifecycle from foundation (Phase 4 batch)

* refactor(agents,enforcement): import lifecycle from foundation (Phase 4 batch)

* refactor(tests): import lifecycle from foundation (Phase 4 batch)

* refactor(foundation): absorb lifecycle _validate + _generators

Phase 4 Tasks 9 + 10. Moves the lifecycle spec's internal validators
to roboco/foundation/_validate_lifecycle.py and its RAG/prompt artifact
emitter to roboco/foundation/_generators.py.

The lifecycle validators live in a sibling module (not merged with
foundation/_validate.py) because the lifecycle spec imports from
foundation at module load — placing the lifecycle checks alongside the
identity checks would create an import cycle between
roboco.foundation and roboco.foundation.policy.lifecycle (the latter
calls the validators at the bottom of its own definition). The
_validate_lifecycle module defers its policy.lifecycle imports to
function bodies so it loads cleanly when the spec hasn't finished
initialising yet; the per-file PLC0415 exemption in pyproject.toml
documents the reason.

Test files relocated:
- tests/lifecycle/test_spec.py        -> tests/foundation/test_lifecycle_spec.py
- tests/lifecycle/test_generators.py  -> tests/foundation/test_lifecycle_generators.py

scripts/build_lifecycle_artifacts.py now imports the generators from
roboco.foundation; the on-disk artifacts (docs/rag/lifecycle,
panel/lib/lifecycle.json, agents/prompts/_generated/lifecycle-*.md)
regenerate byte-identically.

After this commit, roboco/lifecycle/ contains only the spec.py and
__init__.py re-export shims — Task 8 deletes those.

No behavior change. 6638 tests pass; make quality green.

* refactor(lifecycle): delete legacy roboco/lifecycle/ package

Phase 4 Task 8. All consumers migrated to roboco.foundation.policy.lifecycle
in Tasks 2-7; the internal validators + generators moved to foundation in
Tasks 9-10. The legacy package contained only re-export shims.

Also trims tests/foundation/test_role_reexport.py — the two assertions that
checked the lifecycle.spec shim's object-identity are gone with the shim.
The two models.base shim assertions (AgentRole / Team) are still
meaningful and stay.

Inline docstrings / comments in enforcement/task_lifecycle.py,
services/gateway/role_config.py, services/gateway/content_actions.py,
tests/integration/test_task_service_lifecycle_misc.py and
foundation/policy/lifecycle.py that referenced the now-deleted
roboco.lifecycle.spec module are updated to point at
roboco.foundation.policy.lifecycle.

After this commit, roboco.lifecycle is gone. Lifecycle policy lives only
at roboco.foundation.policy.lifecycle. Adding new lifecycle rules edits
exactly that one file.

* refactor(api): consolidate route-guard role-sets via foundation

Replace hand-written role-name string frozensets in roboco/api/deps.py
(_PM_OR_ABOVE_ROLES, _DEVELOPER_OR_ABOVE_ROLES, _GLOBAL_CELL_ACCESS_ROLES)
and roboco/api/routes/v2/_role_dep.py (require_dev/qa/doc/cell_pm/main_pm/
board/auditor) with foundation-derived expressions over PM_ROLES,
BOARD_ROLES, DEV_ROLES, and Role enum members.

Behavior is preserved: Role is a StrEnum, so the lowercase X-Agent-Role
header still compares equal to its matching member. HEAD_MARKETING stays
excluded from every -or-above set (marketing spokesperson, not approver);
the carve-out is now expressed as (BOARD_ROLES - {Role.HEAD_MARKETING})
instead of an opaque literal.

Adds tests/foundation/test_route_guard_consolidation.py (6 tests) pinning
both the foundation-derived membership and the import contract.

* test(foundation): Phase 4 smoke gate + housekeeping closeout

Phase 4 closes the foundation canonicalization effort (Phases 1-4 spanning
2026-05-10 -> 2026-05-11):

Phase 1 - identity + task_completeness (skeleton-task bug killed)
Phase 2 - tracing + journaling (pre-gateway cadence restored)
Phase 3 - communications + agent_loop (channel/notification/A2A/circuit-breaker)
Phase 4 - housekeeping (lifecycle moved to foundation; consumers migrated)

All cross-cutting policy now lives in roboco/foundation/. Adding a policy
edits exactly one file. The legacy roboco.lifecycle package is gone.
Smoke gates 1-4 enforce: no skeleton tasks, no inline journal:decision
checks, channel topology canonical, A2A tristate preserved, auditor silent
at runtime, lifecycle module path canonical.

make quality + make foundation-check both green.

* fix(mcp/agent_sdk): wire per-verb circuit breaker into response handler

Phase 3 Task 14 added the SDK infrastructure (tracker, endpoints,
Envelope.circuit_open, retry_limit_for) but nothing was actually
recording rejections — the breaker never tripped. This commit wires
the gateway-response path so every rejection envelope (tracing_gap /
invalid_state / not_authorized / incomplete_input) hits
POST /verb/attempted, and if the breaker is open, the envelope is
substituted with the circuit_open response before the agent sees it.

Best-effort: SDK-unreachable / malformed-response failures fall open
(agent sees the original rejection), so the breaker never breaks the
gateway path.

* fix(notification_delivery): retype CEO approval-flow notifications APPROVAL

notify_assignee_of_ceo_rejection and notify_ceo_of_escalation were both
typed NotificationType.TASK_ASSIGNMENT, which the Phase 3 foundation
table (ACK_REQUIRED_BY_TYPE in roboco/foundation/policy/communications.py)
maps to requires_ack=False. Both are approval-flow notifications and
should mandate acknowledgment.

Retyped both to NotificationType.APPROVAL so the table lookup yields
requires_ack=True via ACK_REQUIRED_BY_TYPE[NotificationType.APPROVAL].

* test(foundation): move lifecycle parity + smoke-replay tests under tests/foundation/

Phase 4 Task 8 deleted roboco/lifecycle/ but tests/lifecycle/ still held
two files importing roboco.foundation.policy.lifecycle. Mirror the layout
of test_lifecycle_spec.py and test_lifecycle_generators.py (moved in
Phase 4 Tasks 9+10) by relocating them under tests/foundation/ with the
test_lifecycle_* prefix, then delete the now-empty tests/lifecycle/
package.

  tests/lifecycle/test_consumer_parity.py
    -> tests/foundation/test_lifecycle_consumer_parity.py
  tests/lifecycle/test_smoke_replay.py
    -> tests/foundation/test_lifecycle_smoke_replay.py

* build(make): consolidate ci-lifecycle-check into foundation-check

ci-lifecycle-check was a thin wrapper that regenerated lifecycle artifacts
via scripts/build_lifecycle_artifacts.py and gated on git diff. After
Phase 4 it sat alongside foundation-check covering the same drift-gate
intent. Merge the lifecycle-artifact regen + git-diff step into
foundation-check so a single 'make foundation-check' is the canonical
drift gate.

Keep ci-lifecycle-check as a phony alias forwarding to foundation-check
for any external script or CI lane still using the old target name.
Drop the redundant ci-lifecycle-check call from 'make quality'.

* ++

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-05-11 02:15:47 +02:00

57 lines
1.5 KiB
Python
Executable File

#!/usr/bin/env python3
"""Regenerate lifecycle artifacts from roboco/foundation/policy/lifecycle.py.
Outputs (deterministic):
- docs/rag/lifecycle/intent-verbs.md
- docs/rag/lifecycle/status-transitions.md
- panel/lib/lifecycle.json
- agents/prompts/_generated/lifecycle-{role}.md (one per role)
Run as part of `make lifecycle`. CI gate: `make lifecycle && git diff
--exit-code` fails if regeneration produces a diff.
"""
from __future__ import annotations
from pathlib import Path
from roboco.foundation import _generators
from roboco.foundation.policy.lifecycle import Role
REPO_ROOT = Path(__file__).resolve().parent.parent
def write(path: Path, content: str) -> None:
path.parent.mkdir(parents=True, exist_ok=True)
path.write_text(content)
print(f"wrote {path.relative_to(REPO_ROOT)}")
def main() -> int:
write(
REPO_ROOT / "docs" / "rag" / "lifecycle" / "intent-verbs.md",
_generators.render_intent_verbs_md(),
)
write(
REPO_ROOT / "docs" / "rag" / "lifecycle" / "status-transitions.md",
_generators.render_status_transitions_md(),
)
write(
REPO_ROOT / "panel" / "lib" / "lifecycle.json",
_generators.render_panel_json(),
)
for role in Role:
write(
REPO_ROOT
/ "agents"
/ "prompts"
/ "_generated"
/ f"lifecycle-{role.value}.md",
_generators.render_agent_prompt_fragment(role.value),
)
return 0
if __name__ == "__main__":
raise SystemExit(main())