Files
roboco/tests/integration/test_lifecycle_real_db.py
T
06682f33c6 Fix: agent workflow hardening (#70)
* fix(gateway): push the branch before QA handoff so reviewers see the latest commits

The commit content tool commits locally without pushing; only open_pr pushed
the branch. On the first submission that was fine, but a fix committed while
addressing needs_revision never reached origin (open_pr is skipped once the PR
exists), so QA — which reviews the remote PR branch — re-reviewed the stale
remote and re-failed the task on every cycle, a loop that never converged.

i_am_done now pushes the task branch (idempotent; a no-op when nothing is
unpushed) as part of the shared submit gate, covering both the normal and
resume-from-verifying paths. A push failure blocks the handoff with a clear
remediation rather than parking the task in awaiting_qa with commits that exist
only in the developer's local workspace.

* fix(orchestrator): don't reap a stale claim while the agent's container is alive

The stale-claim reaper released any claimed/in_progress task whose
last_heartbeat_at exceeded the TTL. The heartbeat only updates on certain
gateway calls, so a developer deep in a long edit/test cycle outran the TTL and
had its claim reaped mid-work — churning the task and risking a double spawn
against the still-running container.

The reaper now skips a task whose assignee still holds a live (ACTIVE) agent
instance, trusting container liveness — the ground truth — over the heartbeat
proxy. The check is defensive on missing fields so a heartbeat-only caller (and
the reaper's existing unit tests) behave exactly as before.

* fix(gateway): refuse to unblock a task while a dependency is unfinished

A PM unblock on a dependency-gated task moved it straight to in_progress,
overriding the dependency — letting a dependent proceed without its upstream's
work (e.g. a frontend task built before its UX design lands). A dependency
block is meant to clear on its own via _unblock_dependents the moment the
upstream reaches a terminal state.

unblock now refuses while any dependency is still non-terminal, returning a
clear remediation that the block resolves automatically. Manual unblock remains
available for genuine, non-dependency blockers.

* fix(gateway): release a dependency-blocked claim to pending instead of looping

A task that reached claimed/in_progress with an unfinished dependency was left
in that state when the claim guard rejected, so the orchestrator's respawn loop
kept reviving its assignee — which could make no progress — burning work for
nothing.

The claim guard now releases such a task back to pending. claimed -> blocked is
not a legal transition, so pending — held by the dispatch dependency filter — is
the lifecycle-correct resting state: the respawn loop ignores pending tasks, and
_unblock_dependents re-dispatches it once the upstream reaches a terminal state.
release_dependency_blocked_claim shares a _force_unclaim_to_pending core with
unclaim_for_reaper so both record a truthful work-session abandon reason.

* feat(security): warn at startup in header-trust mode + document the auth posture

When ROBOCO_AGENT_AUTH_REQUIRED is not enabled the API accepts the X-Agent-Id /
X-Agent-Role headers without a signed token, so any client that can reach it may
act as any role (including 'ceo'). The API now logs a clear warning at startup
in this mode, and the README gains a Security section documenting the auth
posture and how to harden it. Acceptable only on a trusted private network — do
not expose the API to untrusted networks.

* fix(workspace): scope the refresh fetch to current + default branch

ensure_workspace's healthy short-circuit ran an all-refs 'git fetch origin' to
keep every origin/<branch> ref current. On a monorepo with many accumulated
feature/* branches that exceeds the refresh timeout, the fetch silently fails,
and the workspace keeps a stale base — so an agent builds on an out-of-date
branch.

The refresh now fetches only the workspace's current branch and the repo's
default branch (resolved via origin/HEAD), with --no-tags --prune: it transfers
near-nothing and can't time out. Readers need their own branch and the default;
the integration branch is refreshed at branch-creation time.

* fix(git): refresh a dependency-blocked task's branch off the current integration tip

A cross-cell dependent (e.g. a frontend task waiting on the UX design) was
branched off a base captured before its upstream merged into the integration
branch, and the branch was never re-synced — so the agent built on a stale
snapshot with none of the upstream's work.

Two changes close the gap:
- release_dependency_blocked_claim now clears branch_name, so the re-claim
  (after the dependency clears) re-runs branch creation.
- create_branch, when the branch is already on disk with no commits of its own,
  resets it onto the freshly-pulled base — the dependent now builds on the
  current integration tip. A branch carrying real commits is left untouched, so
  no work is discarded; the cell->leaf cascade carries the upstream down to the
  dev branch automatically.

* refactor(gateway): drop the sibling-sequence claim guard

Sibling sequence no longer gates a claim. Cross-cell ordering is
enforced by task dependencies — a cell task that depends on another is
held until its upstream reaches a terminal state, a stronger,
status-aware gate than the sequence-number check. That check was
dormant in practice anyway: every fan-out child carries sequence 0, on
which the guard short-circuited. `sequence` stays a sibling-ordering /
dispatch-priority field (list_pending ordering and the panel).

Removes sibling_sequence_guard and its _earlier_blocking_sibling
helper, the now-unused skip_sequence parameter threaded through the
claim verbs, and the sibling fetch that fed it.

* feat(gateway): sort a cross-cell dependent after its upstream

When the frontend cell task is wired to depend on its UX/UI sibling, set
its sequence to the upstream's sequence + 1 so it sorts after the design
it waits on — list_pending ordering and the panel now show UX ahead of
the implementation it gates, in either delegation order.

Adds TaskService.set_sequence (the sibling-ordering field is a service
write; it carries no claim-gating semantics — dependencies gate claims).

* feat(gateway): make the backend cell depend on UX too

UX/UI design defines the screens and API contracts both implementation
cells build against, so the backend cell — not just the frontend — waits
on the UX/UI cell task in a product fan-out and sorts after it. Wires in
either delegation order: a backend task delegated after UX gets the
dependency directly; a UX task delegated after a still-pending backend
sibling retro-wires it.

Mirrors the existing frontend wiring (_depend_backend_on_ux and
_depend_pending_backends_on_ux). Backend is held by the same dependency
gate, so it costs no extra dispatch churn.

* fix(websocket): forward notification acks instead of logging them incomplete

The bridge handler serves both notification.sent and notification.acked,
but acked events carry `agent_id` (the acking agent) rather than
`recipient_id`, so every acknowledgement tripped the missing-field guard
and logged "Incomplete notification event" instead of reaching the panel.
Accept either field as the recipient.

* feat(api): hint the full UUID when a truncated task id fails validation

Agents copy the 8-character task prefix the system shows them (the commit
prefix, task summaries) and send it as task_id, which fails UUID
validation with an opaque "invalid length" 422 and wastes a call. The
request-validation handler now detects a task_id UUID error and attaches
a `remediate` hint telling the agent to retry with the full 36-character
UUID from its task envelope.

* fix(audit): record the blocked transition when a task is escalated

Escalation sets a task to blocked by writing task.status directly, which
bypassed the validated transition helper and so never emitted a
task.blocked audit row — the lifecycle moved but the Auditor saw nothing.
Extract the audit emit from the central transition helper into
_emit_status_transition_audit and call it from the escalate path,
capturing the prior status and outgoing owner before reassignment so the
row is attributed correctly.

* fix(docs): stop doubling the docs path so design specs index into RAG

The documenter sometimes hands a doc path already rooted at docs/, and
joining it onto DOCS_BASE_PATH (/app/docs) produced /app/docs/docs/...,
so the file was never found and the spec never indexed — the frontend
cell could not retrieve the UX design over RAG. Normalize the path
before joining: trust an absolute path, otherwise strip a single
redundant leading docs/ segment.

* feat(security): let the control panel authenticate in secure mode

With ROBOCO_AGENT_AUTH_REQUIRED=true every request must carry a valid
HMAC token, which locked the human control panel out — it sends role
headers but no token. nginx, the only trusted hop between the browser
and the API, now injects the CEO token on /api and /ws, so the browser
never holds the signing secret. The injected value is just the existing
per-agent token issued for the CEO identity (issue_panel_token), so the
token-verification path is unchanged. An empty value (dev/header-trust
mode) renders to no header.

`make panel-token` prints the value; set it as ROBOCO_PANEL_AGENT_TOKEN
in .env before enabling secure mode. .env.example and the README
Security section document the flow.

* chore(compose): consolidate the two compose files into one

docker-compose.yml and docker-compose.yaml had diverged: .yml — the file
Docker actually uses — carried ROBOCO_PUBLIC_BASE_URL but was missing the
/app/manifests bind-mount, while .yaml had the manifests mount but not
the base URL. Merge the union into docker-compose.yml and delete the
duplicate so there is one source of truth and no "multiple config files"
warning.

This activates the manifests mount in the deployed file: without it the
orchestrator writes per-agent tool manifests to its ephemeral container
fs, they never reach the host for the daemon to bind-mount, and agents
fall back to all-verbs registration. Drop the stale .yaml reference from
the config.py docstring, the labeler, and the CI path filters.

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-06-05 16:35:22 +02:00

868 lines
30 KiB
Python

"""Tier 3 — end-to-end happy paths against the real test DB.
Each test exercises the spec → choreographer → TaskService → DB stack
with Alembic migrations applied. Catches "spec says X, DB constraint
says Y" mismatches the unit-tier parametrized parity suite cannot
detect.
Companion to ``tests/integration/test_full_lifecycle_real_db.py``
(audit P2-1 deliverable). That file walks one task through the dev
chain end to end; this file isolates each major lifecycle path into
its own test so a regression on, say, QA-fail does not also blow up
the doc-handoff test.
Mocks: only the git layer (workspace + PR ops) is stubbed because the
test DB has no checkout. The spec, choreographer, VerbRunner, and
TaskService are real — those are the layers Task 30 verifies.
"""
from __future__ import annotations
from datetime import UTC, datetime
from typing import TYPE_CHECKING, Any
from unittest.mock import AsyncMock
from uuid import UUID, uuid4
import pytest
import pytest_asyncio
from roboco.db.tables import AgentTable, ProjectTable, TaskTable
from roboco.foundation.policy.lifecycle import Status
from roboco.models.base import (
AgentRole,
AgentStatus,
TaskNature,
TaskStatus,
TaskType,
Team,
)
from roboco.services.gateway.choreographer import Choreographer, ChoreographerDeps
from roboco.services.task import TaskService
# A developer fresh claim must carry a substantive step checklist.
_STEPS = [
{
"title": "Implement the change",
"description": (
"edit the target file, add tests, run them, and stage the "
"change for commit on the task branch"
),
}
]
_GOOD_PLAN = (
"Implement the task on its feature branch: edit the target module, add or "
"update unit tests covering the change, run the suite locally, then commit "
"on the branch and open a PR. Keep the diff focused on the acceptance "
"criteria and verify it before submitting for QA."
)
_GOOD_TC = ["Follow the existing module's patterns; keep the change minimal."]
_GOOD_RISKS = [
{
"risk": "Scope creep balloons the diff and slows review.",
"mitigation": "Touch only the files the acceptance criteria require.",
}
]
if TYPE_CHECKING:
from collections.abc import AsyncIterator
from sqlalchemy.ext.asyncio import AsyncSession
_BRANCH = "feature/backend/healthz"
_PR_NUMBER = 8
_PR_URL = "https://github.com/example/life/pull/8"
# QA notes must clear settings.qa_notes_min_chars (default 80).
_QA_PASS_NOTES = (
"Reviewed the diff; route returns 200 OK with timestamp. Tests cover "
"both acceptance criteria. Approving."
)
_QA_FAIL_NOTES_PREFIX = (
"Reviewed the diff; route returns 500 on the timestamp branch and the "
"second acceptance criterion is not exercised by the new tests. "
)
class _StubGit:
"""Deterministic GitService stub.
Mirrors ``test_full_lifecycle_real_db.py``'s _StubGit. Mutates the
test's TaskTable row directly so the choreographer reads consistent
pr_number / commits state without disk or network I/O.
"""
def __init__(self, session: Any, task: TaskTable) -> None:
self._session = session
self._task = task
async def commit(
self,
*,
branch_name: str,
message: str,
task_id: UUID,
files: list[str] | None = None,
actor_agent_id: Any = None,
) -> dict[str, Any]:
del branch_name, files, actor_agent_id
sha = uuid4().hex[:40]
commits = list(self._task.commits or [])
commits.append({"sha": sha, "message": message, "task_id": str(task_id)})
self._task.commits = commits
await self._session.flush()
return {
"sha": sha,
"message": message,
"files_changed": 1,
"insertions": 1,
"deletions": 0,
}
async def push_branch(
self, branch_name: str, *, actor_agent_id: Any = None
) -> tuple[str, int]:
del branch_name, actor_agent_id
return ("ok", 0)
async def push_task_branch(self, agent_id: UUID, task_id: UUID) -> int:
del agent_id, task_id
return 0
async def create_pr(
self,
branch_name: str,
*,
parent: str,
is_root_pr: bool,
actor_agent_id: Any = None,
) -> dict[str, Any]:
del branch_name, parent, actor_agent_id
self._task.pr_number = _PR_NUMBER
self._task.pr_url = _PR_URL
self._task.pr_created = True
await self._session.flush()
return {"pr_number": _PR_NUMBER, "pr_url": _PR_URL, "is_root_pr": is_root_pr}
async def diff(
self, *, branch_name: str, base: Any = None, actor_agent_id: Any = None
) -> str:
del branch_name, base, actor_agent_id
return "stub diff"
async def list_changed_files(
self, *, branch_name: str, base: Any = None, actor_agent_id: Any = None
) -> list[str]:
del branch_name, base, actor_agent_id
return []
async def pr_target(self, pr_number: int, *, actor_agent_id: Any = None) -> str:
del pr_number, actor_agent_id
return "main"
async def pr_merge(self, *args: Any, **kwargs: Any) -> dict[str, Any]:
del args, kwargs
return {
"merged": True,
"sha": uuid4().hex[:40],
"merge_commit_sha": uuid4().hex[:40],
}
def _mock_evidence_repo() -> Any:
repo = AsyncMock()
for method in (
"list_unread_a2a",
"list_unread_mentions",
"list_pending_notifications",
"task_metadata_gaps",
"recent_team_activity",
"blockers_in_lane",
"journal_highlights_for_task",
):
getattr(repo, method).return_value = []
return repo
def _mock_journal_with_reflect() -> Any:
"""Journal stub that reports reflect/learning/decision entries present.
``latest_decision_at`` is anchored to ``datetime.now(UTC)`` so the C8
recency window on the PM-decision gate accepts it.
"""
journal = AsyncMock()
journal.has_reflect_for_task.return_value = True
journal.has_learning_for_task.return_value = True
journal.has_decision_for_task.return_value = True
journal.has_struggle_for_task.return_value = False
journal.latest_decision_at.return_value = datetime.now(UTC)
return journal
def _mock_work_session() -> Any:
"""WorkSession stub: stable file list, no unpushed commits."""
ws = AsyncMock()
ws.files_changed.return_value = ["roboco/api/routes/health.py"]
ws.has_unpushed_commits.return_value = False
return ws
def _build_choreographer(
db_session: Any, task: TaskTable, task_service: TaskService
) -> Choreographer:
"""Wire a real Choreographer with the supplied TaskService + stubbed git.
Caller owns the TaskService so it can use it for direct DB reads
(``task_service.get(task_id)``) — sharing one instance keeps the
session contract clean and avoids "two TaskServices, two views"
surprises.
"""
deps = ChoreographerDeps(
task=task_service,
work_session=_mock_work_session(),
git=_StubGit(db_session, task),
a2a=AsyncMock(),
journal=_mock_journal_with_reflect(),
audit=AsyncMock(),
evidence_repo=_mock_evidence_repo(),
)
return Choreographer(deps)
async def _seed_agents_and_project(
db_session: AsyncSession,
) -> dict[str, Any]:
"""Seed system + project + dev/qa/doc/cell_pm agents.
Slugs match ``agents_config.ESCALATION_CHAIN`` so ``i_am_blocked``
and ``escalate_up`` find a real escalation target. The agents-config
chain is the source of truth at runtime; matching it here exercises
the same lookup the gateway uses in production.
"""
system_agent = AgentTable(
id=uuid4(),
name="System",
slug=f"system-{uuid4().hex[:8]}",
role=AgentRole.SYSTEM,
team=None,
status=AgentStatus.ACTIVE,
model_config={},
system_prompt="system",
capabilities=[],
permissions={},
metrics={},
)
db_session.add(system_agent)
await db_session.flush()
project = ProjectTable(
id=uuid4(),
name="Lifecycle Test Project",
slug=f"life-{uuid4().hex[:8]}",
git_url="https://github.com/example/life.git",
default_branch="main",
protected_branches=["main"],
assigned_cell=Team.BACKEND,
created_by=system_agent.id,
is_active=True,
)
db_session.add(project)
await db_session.flush()
dev_agent = AgentTable(
id=uuid4(),
name="BE Dev 1",
slug="be-dev-1",
role=AgentRole.DEVELOPER,
team=Team.BACKEND,
status=AgentStatus.ACTIVE,
model_config={},
system_prompt="dev",
capabilities=["python"],
permissions={},
metrics={},
)
qa_agent = AgentTable(
id=uuid4(),
name="BE QA",
slug="be-qa",
role=AgentRole.QA,
team=Team.BACKEND,
status=AgentStatus.ACTIVE,
model_config={},
system_prompt="qa",
capabilities=["review"],
permissions={},
metrics={},
)
doc_agent = AgentTable(
id=uuid4(),
name="BE Doc",
slug="be-doc",
role=AgentRole.DOCUMENTER,
team=Team.BACKEND,
status=AgentStatus.ACTIVE,
model_config={},
system_prompt="doc",
capabilities=["docs"],
permissions={},
metrics={},
)
cell_pm_agent = AgentTable(
id=uuid4(),
name="BE Cell PM",
slug="be-pm",
role=AgentRole.CELL_PM,
team=Team.BACKEND,
status=AgentStatus.ACTIVE,
model_config={},
system_prompt="cell_pm",
capabilities=["coord"],
permissions={},
metrics={},
)
db_session.add_all([dev_agent, qa_agent, doc_agent, cell_pm_agent])
await db_session.flush()
return {
"system_agent": system_agent,
"project": project,
"dev_agent": dev_agent,
"qa_agent": qa_agent,
"doc_agent": doc_agent,
"cell_pm_agent": cell_pm_agent,
}
def _build_task(
*,
project_id: UUID,
creator_id: UUID,
assignee_id: UUID | None,
status: TaskStatus,
) -> TaskTable:
"""Construct a backend code-typed task pinned to ``status``.
``acceptance_criteria_status`` carries the stub artefact rows the
pre-merge gate inspects; supplying them here keeps the per-test
setup readable.
"""
return TaskTable(
id=uuid4(),
title="Add /healthz endpoint",
description="Return 200 OK from /healthz",
status=status,
priority=2,
task_type=TaskType.CODE,
nature=TaskNature.TECHNICAL,
team=Team.BACKEND,
project_id=project_id,
created_by=creator_id,
assigned_to=assignee_id,
branch_name=_BRANCH,
acceptance_criteria=["Returns 200", "Includes timestamp"],
acceptance_criteria_status=[
{"criterion": "Returns 200", "referencing_artifact_id": "stub"},
{"criterion": "Includes timestamp", "referencing_artifact_id": "stub"},
],
)
@pytest_asyncio.fixture
async def lifecycle_setup(
db_session: AsyncSession,
) -> AsyncIterator[dict[str, Any]]:
"""Seed agents + project + a single PENDING task assigned to the dev."""
seeded = await _seed_agents_and_project(db_session)
task = _build_task(
project_id=seeded["project"].id,
creator_id=seeded["system_agent"].id,
assignee_id=seeded["dev_agent"].id,
status=TaskStatus.PENDING,
)
db_session.add(task)
await db_session.flush()
seeded["task"] = task
yield seeded
# ---------------------------------------------------------------------------
# 1. Dev path: pending → claimed → in_progress → verifying → awaiting_qa
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_dev_full_chain_through_awaiting_qa(
db_session: AsyncSession, lifecycle_setup: dict[str, Any]
) -> None:
"""pending → claimed → in_progress → verifying → awaiting_qa.
Drives ``i_will_work_on`` (claim+plan+start), a stubbed commit,
``open_pr``, then ``i_am_done`` which auto-runs submit_verification
+ submit_qa. Asserts the final DB row sits at ``awaiting_qa``.
"""
task = lifecycle_setup["task"]
dev_agent = lifecycle_setup["dev_agent"]
task_service = TaskService(db_session)
stub_git = _StubGit(db_session, task)
deps = ChoreographerDeps(
task=task_service,
work_session=_mock_work_session(),
git=stub_git,
a2a=AsyncMock(),
journal=_mock_journal_with_reflect(),
audit=AsyncMock(),
evidence_repo=_mock_evidence_repo(),
)
c = Choreographer(deps)
env = await c.i_will_work_on(
dev_agent.id,
task.id,
plan=_GOOD_PLAN,
steps=_STEPS,
technical_considerations=_GOOD_TC,
risks=_GOOD_RISKS,
)
assert env.error is None, f"i_will_work_on failed: {env.message}"
assert env.status == Status.IN_PROGRESS.value
# Commit + record progress so open_pr's commits-precondition holds.
await stub_git.commit(
branch_name=_BRANCH,
message=f"[{str(task.id)[:8]}] feat(api): add /healthz",
task_id=task.id,
)
await task_service.add_progress(task.id, dev_agent.id, "implemented /healthz")
env = await c.open_pr(dev_agent.id, task.id)
assert env.error is None, f"open_pr failed: {env.message}"
env = await c.i_am_done(dev_agent.id, task.id, "tests pass; route works")
assert env.error is None, f"i_am_done failed: {env.message}"
assert env.status == Status.AWAITING_QA.value
final = await task_service.get(task.id)
assert final is not None
assert str(final.status) == Status.AWAITING_QA.value
# i_am_done auto-runs submit_qa, which hands the task off to the
# backend QA agent (production behaviour — see ``_notify_qa``).
# Resolve via the same lookup the choreographer uses so the
# assertion is robust to other test fixtures that may have seeded
# additional QA agents on the BACKEND team (e.g. smoke_test_batch
# commits its agents, so they outlive their session).
resolved_qa = await task_service.qa_agent_for_team(Team.BACKEND)
assert resolved_qa is not None
assert final.assigned_to == resolved_qa.id
# ---------------------------------------------------------------------------
# 2. QA pass path: awaiting_qa → claimed → awaiting_documentation
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_qa_pass_path(
db_session: AsyncSession, lifecycle_setup: dict[str, Any]
) -> None:
"""awaiting_qa → claim_review → pass_review → awaiting_documentation.
``claim_review`` keeps status at AWAITING_QA (specialised qa_claim
sets assignment without transitioning) so the spec's source-status
requirement on ``qa_pass`` still matches downstream.
"""
task = lifecycle_setup["task"]
qa_agent = lifecycle_setup["qa_agent"]
doc_agent = lifecycle_setup["doc_agent"]
# Pin the task at awaiting_qa with a PR + commit so the QA gates
# (pr exists, commits non-empty) all pass.
task.status = TaskStatus.AWAITING_QA
task.pr_number = _PR_NUMBER
task.pr_url = _PR_URL
task.commits = [
{"sha": uuid4().hex[:40], "message": "feat: /healthz", "task_id": str(task.id)}
]
task.self_verified = True
await db_session.flush()
task_service = TaskService(db_session)
c = _build_choreographer(db_session, task, task_service)
env = await c.claim_review(qa_agent.id, task.id)
assert env.error is None, f"claim_review failed: {env.message}"
after_claim = await task_service.get(task.id)
assert after_claim is not None
assert str(after_claim.status) == Status.AWAITING_QA.value
assert after_claim.assigned_to == qa_agent.id
env = await c.pass_review(qa_agent.id, task.id, notes=_QA_PASS_NOTES)
assert env.error is None, f"pass_review failed: {env.message}"
assert env.status == Status.AWAITING_DOCUMENTATION.value
final = await task_service.get(task.id)
assert final is not None
assert str(final.status) == Status.AWAITING_DOCUMENTATION.value
# pass_review reassigns to the team's documenter for handoff. Look
# up via the same path the choreographer uses (robust to other
# tests' committed BACKEND documenters).
resolved_doc = await task_service.documenter_for_team(Team.BACKEND)
assert resolved_doc is not None
assert final.assigned_to == resolved_doc.id
del doc_agent # asserted indirectly via documenter_for_team.
# ---------------------------------------------------------------------------
# 3. QA fail path: awaiting_qa → claimed → needs_revision
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_qa_fail_path(
db_session: AsyncSession, lifecycle_setup: dict[str, Any]
) -> None:
"""awaiting_qa → claim_review → fail_review(issues) → needs_revision.
``fail_review`` reassigns to the original developer so they can
revise; that lookup walks ``quick_context``'s
``original_developer:<slug>`` marker, which the dev path stamps
on i_will_work_on. Here we set it directly so the assertion is
deterministic without re-running the dev chain.
"""
task = lifecycle_setup["task"]
qa_agent = lifecycle_setup["qa_agent"]
dev_agent = lifecycle_setup["dev_agent"]
task.status = TaskStatus.AWAITING_QA
task.pr_number = _PR_NUMBER
task.pr_url = _PR_URL
task.commits = [
{"sha": uuid4().hex[:40], "message": "feat: /healthz", "task_id": str(task.id)}
]
task.self_verified = True
# ``extract_original_developer`` parses a UUID off this line; the
# spec layer's slug-based self-review check is a separate code path
# (``_extract_original_developer`` in qa.py) which only fires when
# an actor's slug equals this value, so a UUID here doesn't trip it.
task.quick_context = f"original_developer:{dev_agent.id}"
await db_session.flush()
task_service = TaskService(db_session)
c = _build_choreographer(db_session, task, task_service)
env = await c.claim_review(qa_agent.id, task.id)
assert env.error is None, f"claim_review failed: {env.message}"
issues = ["Returns 500 on the timestamp branch", "Missing test for the second AC"]
# fail_review's notes are derived from issues; QA pass-gate also
# requires notes >= 80 chars, so we send a leading explanation as
# the issues list — the verb concatenates them and easily clears
# the threshold.
long_issues = [_QA_FAIL_NOTES_PREFIX + issues[0], issues[1]]
env = await c.fail_review(qa_agent.id, task.id, issues=long_issues)
assert env.error is None, f"fail_review failed: {env.message}"
assert env.status == Status.NEEDS_REVISION.value
final = await task_service.get(task.id)
assert final is not None
assert str(final.status) == Status.NEEDS_REVISION.value
# fail_qa reassigns to the original developer so they can revise.
assert final.assigned_to == dev_agent.id
# ---------------------------------------------------------------------------
# 4. Doc path: awaiting_documentation → claimed → awaiting_pm_review
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_doc_path(
db_session: AsyncSession, lifecycle_setup: dict[str, Any]
) -> None:
"""awaiting_documentation → claim_doc_task → i_documented → awaiting_pm_review.
``claim_doc_task`` keeps status at AWAITING_DOCUMENTATION (doc_claim
is assignment-only, mirroring qa_claim). ``i_documented`` flips to
AWAITING_PM_REVIEW and reassigns to the cell PM for that team.
"""
task = lifecycle_setup["task"]
doc_agent = lifecycle_setup["doc_agent"]
cell_pm_agent = lifecycle_setup["cell_pm_agent"]
task.status = TaskStatus.AWAITING_DOCUMENTATION
task.pr_number = _PR_NUMBER
task.pr_url = _PR_URL
task.pr_created = True
task.qa_verified = True
task.assigned_to = None # documenter must claim from unassigned.
task.commits = [
{"sha": uuid4().hex[:40], "message": "feat: /healthz", "task_id": str(task.id)}
]
await db_session.flush()
task_service = TaskService(db_session)
c = _build_choreographer(db_session, task, task_service)
env = await c.claim_doc_task(doc_agent.id, task.id)
assert env.error is None, f"claim_doc_task failed: {env.message}"
after_claim = await task_service.get(task.id)
assert after_claim is not None
assert str(after_claim.status) == Status.AWAITING_DOCUMENTATION.value
assert after_claim.assigned_to == doc_agent.id
env = await c.i_documented(
doc_agent.id,
task.id,
notes="Documented /healthz behaviour in docs/api/health.md",
files=["docs/api/health.md"],
)
assert env.error is None, f"i_documented failed: {env.message}"
assert env.status == Status.AWAITING_PM_REVIEW.value
final = await task_service.get(task.id)
assert final is not None
assert str(final.status) == Status.AWAITING_PM_REVIEW.value
# i_documented hands off to the cell PM for the team. Resolve via
# the same lookup the choreographer uses (robust to other tests'
# committed BACKEND cell PMs).
resolved_pm = await task_service.cell_pm_for_team(Team.BACKEND)
assert resolved_pm is not None
assert final.assigned_to == resolved_pm.id
del cell_pm_agent # asserted indirectly via cell_pm_for_team.
# ---------------------------------------------------------------------------
# 5. PM complete (Cell PM, simple task): awaiting_pm_review → completed
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_pm_complete_simple_task(
db_session: AsyncSession, lifecycle_setup: dict[str, Any]
) -> None:
"""awaiting_pm_review → cell_pm complete → completed.
A cell PM completing a non-root awaiting_pm_review task transitions
straight to COMPLETED — there is no cell→main escalation in
``complete`` (the cell→main hand-off, when intended, uses
``submit_up``).
"""
task = lifecycle_setup["task"]
cell_pm_agent = lifecycle_setup["cell_pm_agent"]
task.status = TaskStatus.AWAITING_PM_REVIEW
task.pr_number = _PR_NUMBER
task.pr_url = _PR_URL
task.pr_created = True
task.qa_verified = True
task.docs_complete = True
task.assigned_to = cell_pm_agent.id
task.commits = [
{"sha": uuid4().hex[:40], "message": "feat: /healthz", "task_id": str(task.id)}
]
await db_session.flush()
task_service = TaskService(db_session)
c = _build_choreographer(db_session, task, task_service)
env = await c.complete(
cell_pm_agent.id, task.id, notes="LGTM — merging the leaf PR."
)
assert env.error is None, f"complete failed: {env.message}"
assert env.status == Status.COMPLETED.value
final = await task_service.get(task.id)
assert final is not None
assert str(final.status) == Status.COMPLETED.value
# ---------------------------------------------------------------------------
# 6. PM escalate: awaiting_pm_review → awaiting_ceo_approval
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_pm_escalate_to_ceo_path(
db_session: AsyncSession, lifecycle_setup: dict[str, Any]
) -> None:
"""awaiting_pm_review → main_pm complete on a root task → awaiting_ceo_approval.
Main PM completing a root task (no parent) opens the master PR if
needed and calls ``task.escalate_to_ceo``, leaving the task at
AWAITING_CEO_APPROVAL with ``assigned_to=None``. CEO approval is
human-in-the-loop (UI-driven), so the test stops there.
"""
task = lifecycle_setup["task"]
project = lifecycle_setup["project"]
system_agent = lifecycle_setup["system_agent"]
main_pm_agent = AgentTable(
id=uuid4(),
name="Main PM",
slug="main-pm",
role=AgentRole.MAIN_PM,
team=None,
status=AgentStatus.ACTIVE,
model_config={},
system_prompt="main_pm",
capabilities=["coord"],
permissions={},
metrics={},
)
db_session.add(main_pm_agent)
await db_session.flush()
del project, system_agent # only needed for fixture wiring above.
task.status = TaskStatus.AWAITING_PM_REVIEW
task.pr_number = _PR_NUMBER
task.pr_url = _PR_URL
task.pr_created = True
task.qa_verified = True
task.docs_complete = True
task.parent_task_id = None # explicit — escalate_to_ceo refuses subtasks.
task.assigned_to = main_pm_agent.id
task.commits = [
{"sha": uuid4().hex[:40], "message": "feat: /healthz", "task_id": str(task.id)}
]
await db_session.flush()
task_service = TaskService(db_session)
c = _build_choreographer(db_session, task, task_service)
# SQLAlchemy column-typed `id` attributes need an explicit UUID
# cast for mypy under the project's strict config — the values are
# already real ``uuid.UUID`` at runtime.
env = await c.complete(
UUID(str(main_pm_agent.id)),
UUID(str(task.id)),
notes="Root task ready for CEO approval — escalating.",
)
assert env.error is None, f"complete failed: {env.message}"
assert env.status == Status.AWAITING_CEO_APPROVAL.value
final = await task_service.get(task.id)
assert final is not None
assert str(final.status) == Status.AWAITING_CEO_APPROVAL.value
# main_pm_complete clears assigned_to so the orchestrator does not
# respawn an agent while the task waits on the human CEO.
assert final.assigned_to is None
# ---------------------------------------------------------------------------
# 7. Block + unblock + restore: in_progress → blocked → in_progress
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_block_then_unblock_restore(
db_session: AsyncSession, lifecycle_setup: dict[str, Any]
) -> None:
"""in_progress → i_am_blocked → unblock(restore=True) → in_progress.
``i_am_blocked`` runs the spec's ``block`` action which delegates
to ``task_service.escalate``: the task is reassigned to the dev's
escalation target (``be-pm`` per ``ESCALATION_CHAIN``) and marked
BLOCKED with the original dev stashed in ``blocker_raised_by``.
PM ``unblock(restore=True)`` then falls through to the legacy
``unblock`` path (no ``pre_block_state`` snapshot exists for chain
escalations), restoring assignment to the dev and flipping back
to IN_PROGRESS.
"""
task = lifecycle_setup["task"]
dev_agent = lifecycle_setup["dev_agent"]
cell_pm_agent = lifecycle_setup["cell_pm_agent"]
task_service = TaskService(db_session)
c = _build_choreographer(db_session, task, task_service)
# Drive into in_progress via the real claim+start sequence.
env = await c.i_will_work_on(
dev_agent.id,
task.id,
plan=_GOOD_PLAN,
steps=_STEPS,
technical_considerations=_GOOD_TC,
risks=_GOOD_RISKS,
)
assert env.error is None, f"i_will_work_on failed: {env.message}"
assert env.status == Status.IN_PROGRESS.value
env = await c.i_am_blocked(
dev_agent.id,
task.id,
reason="external dependency on auth library upgrade",
)
assert env.error is None, f"i_am_blocked failed: {env.message}"
assert env.status == Status.BLOCKED.value
blocked = await task_service.get(task.id)
assert blocked is not None
assert str(blocked.status) == Status.BLOCKED.value
# Escalation reassigns to the cell PM (be-pm) and stashes the dev
# as blocker_raised_by so unblock can hand the task back.
assert blocked.assigned_to == cell_pm_agent.id
assert blocked.blocker_raised_by == dev_agent.id
env = await c.unblock(cell_pm_agent.id, task.id, restore=True)
assert env.error is None, f"unblock failed: {env.message}"
assert env.status == Status.IN_PROGRESS.value
restored = await task_service.get(task.id)
assert restored is not None
assert str(restored.status) == Status.IN_PROGRESS.value
# legacy unblock restores assigned_to from blocker_raised_by — the
# original dev gets the task back so the orchestrator respawns them.
assert restored.assigned_to == dev_agent.id
# ---------------------------------------------------------------------------
# 8. Pause + resume: in_progress → paused → in_progress
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_pause_then_resume(
db_session: AsyncSession, lifecycle_setup: dict[str, Any]
) -> None:
"""in_progress → i_am_idle (auto-pause) → resume → in_progress.
There is no agent-driven ``pause`` verb. ``i_am_idle`` auto-pauses
every in_progress task the agent owns so the closure dispatcher can
wake them on respawn. ``resume`` (composes=("resume",)) then flips
the same task back to IN_PROGRESS for the same assignee.
"""
task = lifecycle_setup["task"]
dev_agent = lifecycle_setup["dev_agent"]
task_service = TaskService(db_session)
c = _build_choreographer(db_session, task, task_service)
env = await c.i_will_work_on(
dev_agent.id,
task.id,
plan=_GOOD_PLAN,
steps=_STEPS,
technical_considerations=_GOOD_TC,
risks=_GOOD_RISKS,
)
assert env.error is None, f"i_will_work_on failed: {env.message}"
assert env.status == Status.IN_PROGRESS.value
env = await c.i_am_idle(dev_agent.id)
assert env.error is None, f"i_am_idle failed: {env.message}"
assert env.status == "idle"
paused = await task_service.get(task.id)
assert paused is not None
assert str(paused.status) == Status.PAUSED.value
# Auto-pause keeps assigned_to so resume can find the same claimant.
assert paused.assigned_to == dev_agent.id
env = await c.resume(dev_agent.id, task.id)
assert env.error is None, f"resume failed: {env.message}"
assert env.status == Status.IN_PROGRESS.value
resumed = await task_service.get(task.id)
assert resumed is not None
assert str(resumed.status) == Status.IN_PROGRESS.value
assert resumed.assigned_to == dev_agent.id