mirror of
https://github.com/rennf93/roboco.git
synced 2026-08-03 07:23:24 +02:00
* fix(gateway): push the branch before QA handoff so reviewers see the latest commits The commit content tool commits locally without pushing; only open_pr pushed the branch. On the first submission that was fine, but a fix committed while addressing needs_revision never reached origin (open_pr is skipped once the PR exists), so QA — which reviews the remote PR branch — re-reviewed the stale remote and re-failed the task on every cycle, a loop that never converged. i_am_done now pushes the task branch (idempotent; a no-op when nothing is unpushed) as part of the shared submit gate, covering both the normal and resume-from-verifying paths. A push failure blocks the handoff with a clear remediation rather than parking the task in awaiting_qa with commits that exist only in the developer's local workspace. * fix(orchestrator): don't reap a stale claim while the agent's container is alive The stale-claim reaper released any claimed/in_progress task whose last_heartbeat_at exceeded the TTL. The heartbeat only updates on certain gateway calls, so a developer deep in a long edit/test cycle outran the TTL and had its claim reaped mid-work — churning the task and risking a double spawn against the still-running container. The reaper now skips a task whose assignee still holds a live (ACTIVE) agent instance, trusting container liveness — the ground truth — over the heartbeat proxy. The check is defensive on missing fields so a heartbeat-only caller (and the reaper's existing unit tests) behave exactly as before. * fix(gateway): refuse to unblock a task while a dependency is unfinished A PM unblock on a dependency-gated task moved it straight to in_progress, overriding the dependency — letting a dependent proceed without its upstream's work (e.g. a frontend task built before its UX design lands). A dependency block is meant to clear on its own via _unblock_dependents the moment the upstream reaches a terminal state. unblock now refuses while any dependency is still non-terminal, returning a clear remediation that the block resolves automatically. Manual unblock remains available for genuine, non-dependency blockers. * fix(gateway): release a dependency-blocked claim to pending instead of looping A task that reached claimed/in_progress with an unfinished dependency was left in that state when the claim guard rejected, so the orchestrator's respawn loop kept reviving its assignee — which could make no progress — burning work for nothing. The claim guard now releases such a task back to pending. claimed -> blocked is not a legal transition, so pending — held by the dispatch dependency filter — is the lifecycle-correct resting state: the respawn loop ignores pending tasks, and _unblock_dependents re-dispatches it once the upstream reaches a terminal state. release_dependency_blocked_claim shares a _force_unclaim_to_pending core with unclaim_for_reaper so both record a truthful work-session abandon reason. * feat(security): warn at startup in header-trust mode + document the auth posture When ROBOCO_AGENT_AUTH_REQUIRED is not enabled the API accepts the X-Agent-Id / X-Agent-Role headers without a signed token, so any client that can reach it may act as any role (including 'ceo'). The API now logs a clear warning at startup in this mode, and the README gains a Security section documenting the auth posture and how to harden it. Acceptable only on a trusted private network — do not expose the API to untrusted networks. * fix(workspace): scope the refresh fetch to current + default branch ensure_workspace's healthy short-circuit ran an all-refs 'git fetch origin' to keep every origin/<branch> ref current. On a monorepo with many accumulated feature/* branches that exceeds the refresh timeout, the fetch silently fails, and the workspace keeps a stale base — so an agent builds on an out-of-date branch. The refresh now fetches only the workspace's current branch and the repo's default branch (resolved via origin/HEAD), with --no-tags --prune: it transfers near-nothing and can't time out. Readers need their own branch and the default; the integration branch is refreshed at branch-creation time. * fix(git): refresh a dependency-blocked task's branch off the current integration tip A cross-cell dependent (e.g. a frontend task waiting on the UX design) was branched off a base captured before its upstream merged into the integration branch, and the branch was never re-synced — so the agent built on a stale snapshot with none of the upstream's work. Two changes close the gap: - release_dependency_blocked_claim now clears branch_name, so the re-claim (after the dependency clears) re-runs branch creation. - create_branch, when the branch is already on disk with no commits of its own, resets it onto the freshly-pulled base — the dependent now builds on the current integration tip. A branch carrying real commits is left untouched, so no work is discarded; the cell->leaf cascade carries the upstream down to the dev branch automatically. * refactor(gateway): drop the sibling-sequence claim guard Sibling sequence no longer gates a claim. Cross-cell ordering is enforced by task dependencies — a cell task that depends on another is held until its upstream reaches a terminal state, a stronger, status-aware gate than the sequence-number check. That check was dormant in practice anyway: every fan-out child carries sequence 0, on which the guard short-circuited. `sequence` stays a sibling-ordering / dispatch-priority field (list_pending ordering and the panel). Removes sibling_sequence_guard and its _earlier_blocking_sibling helper, the now-unused skip_sequence parameter threaded through the claim verbs, and the sibling fetch that fed it. * feat(gateway): sort a cross-cell dependent after its upstream When the frontend cell task is wired to depend on its UX/UI sibling, set its sequence to the upstream's sequence + 1 so it sorts after the design it waits on — list_pending ordering and the panel now show UX ahead of the implementation it gates, in either delegation order. Adds TaskService.set_sequence (the sibling-ordering field is a service write; it carries no claim-gating semantics — dependencies gate claims). * feat(gateway): make the backend cell depend on UX too UX/UI design defines the screens and API contracts both implementation cells build against, so the backend cell — not just the frontend — waits on the UX/UI cell task in a product fan-out and sorts after it. Wires in either delegation order: a backend task delegated after UX gets the dependency directly; a UX task delegated after a still-pending backend sibling retro-wires it. Mirrors the existing frontend wiring (_depend_backend_on_ux and _depend_pending_backends_on_ux). Backend is held by the same dependency gate, so it costs no extra dispatch churn. * fix(websocket): forward notification acks instead of logging them incomplete The bridge handler serves both notification.sent and notification.acked, but acked events carry `agent_id` (the acking agent) rather than `recipient_id`, so every acknowledgement tripped the missing-field guard and logged "Incomplete notification event" instead of reaching the panel. Accept either field as the recipient. * feat(api): hint the full UUID when a truncated task id fails validation Agents copy the 8-character task prefix the system shows them (the commit prefix, task summaries) and send it as task_id, which fails UUID validation with an opaque "invalid length" 422 and wastes a call. The request-validation handler now detects a task_id UUID error and attaches a `remediate` hint telling the agent to retry with the full 36-character UUID from its task envelope. * fix(audit): record the blocked transition when a task is escalated Escalation sets a task to blocked by writing task.status directly, which bypassed the validated transition helper and so never emitted a task.blocked audit row — the lifecycle moved but the Auditor saw nothing. Extract the audit emit from the central transition helper into _emit_status_transition_audit and call it from the escalate path, capturing the prior status and outgoing owner before reassignment so the row is attributed correctly. * fix(docs): stop doubling the docs path so design specs index into RAG The documenter sometimes hands a doc path already rooted at docs/, and joining it onto DOCS_BASE_PATH (/app/docs) produced /app/docs/docs/..., so the file was never found and the spec never indexed — the frontend cell could not retrieve the UX design over RAG. Normalize the path before joining: trust an absolute path, otherwise strip a single redundant leading docs/ segment. * feat(security): let the control panel authenticate in secure mode With ROBOCO_AGENT_AUTH_REQUIRED=true every request must carry a valid HMAC token, which locked the human control panel out — it sends role headers but no token. nginx, the only trusted hop between the browser and the API, now injects the CEO token on /api and /ws, so the browser never holds the signing secret. The injected value is just the existing per-agent token issued for the CEO identity (issue_panel_token), so the token-verification path is unchanged. An empty value (dev/header-trust mode) renders to no header. `make panel-token` prints the value; set it as ROBOCO_PANEL_AGENT_TOKEN in .env before enabling secure mode. .env.example and the README Security section document the flow. * chore(compose): consolidate the two compose files into one docker-compose.yml and docker-compose.yaml had diverged: .yml — the file Docker actually uses — carried ROBOCO_PUBLIC_BASE_URL but was missing the /app/manifests bind-mount, while .yaml had the manifests mount but not the base URL. Merge the union into docker-compose.yml and delete the duplicate so there is one source of truth and no "multiple config files" warning. This activates the manifests mount in the deployed file: without it the orchestrator writes per-agent tool manifests to its ephemeral container fs, they never reach the host for the daemon to bind-mount, and agents fall back to all-verbs registration. Drop the stale .yaml reference from the config.py docstring, the labeler, and the CI path filters. --------- Co-authored-by: Renn F <rennf93@users.noreply.github.com>
868 lines
30 KiB
Python
868 lines
30 KiB
Python
"""Tier 3 — end-to-end happy paths against the real test DB.
|
|
|
|
Each test exercises the spec → choreographer → TaskService → DB stack
|
|
with Alembic migrations applied. Catches "spec says X, DB constraint
|
|
says Y" mismatches the unit-tier parametrized parity suite cannot
|
|
detect.
|
|
|
|
Companion to ``tests/integration/test_full_lifecycle_real_db.py``
|
|
(audit P2-1 deliverable). That file walks one task through the dev
|
|
chain end to end; this file isolates each major lifecycle path into
|
|
its own test so a regression on, say, QA-fail does not also blow up
|
|
the doc-handoff test.
|
|
|
|
Mocks: only the git layer (workspace + PR ops) is stubbed because the
|
|
test DB has no checkout. The spec, choreographer, VerbRunner, and
|
|
TaskService are real — those are the layers Task 30 verifies.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
from datetime import UTC, datetime
|
|
from typing import TYPE_CHECKING, Any
|
|
from unittest.mock import AsyncMock
|
|
from uuid import UUID, uuid4
|
|
|
|
import pytest
|
|
import pytest_asyncio
|
|
from roboco.db.tables import AgentTable, ProjectTable, TaskTable
|
|
from roboco.foundation.policy.lifecycle import Status
|
|
from roboco.models.base import (
|
|
AgentRole,
|
|
AgentStatus,
|
|
TaskNature,
|
|
TaskStatus,
|
|
TaskType,
|
|
Team,
|
|
)
|
|
from roboco.services.gateway.choreographer import Choreographer, ChoreographerDeps
|
|
from roboco.services.task import TaskService
|
|
|
|
# A developer fresh claim must carry a substantive step checklist.
|
|
_STEPS = [
|
|
{
|
|
"title": "Implement the change",
|
|
"description": (
|
|
"edit the target file, add tests, run them, and stage the "
|
|
"change for commit on the task branch"
|
|
),
|
|
}
|
|
]
|
|
|
|
_GOOD_PLAN = (
|
|
"Implement the task on its feature branch: edit the target module, add or "
|
|
"update unit tests covering the change, run the suite locally, then commit "
|
|
"on the branch and open a PR. Keep the diff focused on the acceptance "
|
|
"criteria and verify it before submitting for QA."
|
|
)
|
|
_GOOD_TC = ["Follow the existing module's patterns; keep the change minimal."]
|
|
_GOOD_RISKS = [
|
|
{
|
|
"risk": "Scope creep balloons the diff and slows review.",
|
|
"mitigation": "Touch only the files the acceptance criteria require.",
|
|
}
|
|
]
|
|
|
|
if TYPE_CHECKING:
|
|
from collections.abc import AsyncIterator
|
|
|
|
from sqlalchemy.ext.asyncio import AsyncSession
|
|
|
|
|
|
_BRANCH = "feature/backend/healthz"
|
|
_PR_NUMBER = 8
|
|
_PR_URL = "https://github.com/example/life/pull/8"
|
|
# QA notes must clear settings.qa_notes_min_chars (default 80).
|
|
_QA_PASS_NOTES = (
|
|
"Reviewed the diff; route returns 200 OK with timestamp. Tests cover "
|
|
"both acceptance criteria. Approving."
|
|
)
|
|
_QA_FAIL_NOTES_PREFIX = (
|
|
"Reviewed the diff; route returns 500 on the timestamp branch and the "
|
|
"second acceptance criterion is not exercised by the new tests. "
|
|
)
|
|
|
|
|
|
class _StubGit:
|
|
"""Deterministic GitService stub.
|
|
|
|
Mirrors ``test_full_lifecycle_real_db.py``'s _StubGit. Mutates the
|
|
test's TaskTable row directly so the choreographer reads consistent
|
|
pr_number / commits state without disk or network I/O.
|
|
"""
|
|
|
|
def __init__(self, session: Any, task: TaskTable) -> None:
|
|
self._session = session
|
|
self._task = task
|
|
|
|
async def commit(
|
|
self,
|
|
*,
|
|
branch_name: str,
|
|
message: str,
|
|
task_id: UUID,
|
|
files: list[str] | None = None,
|
|
actor_agent_id: Any = None,
|
|
) -> dict[str, Any]:
|
|
del branch_name, files, actor_agent_id
|
|
sha = uuid4().hex[:40]
|
|
commits = list(self._task.commits or [])
|
|
commits.append({"sha": sha, "message": message, "task_id": str(task_id)})
|
|
self._task.commits = commits
|
|
await self._session.flush()
|
|
return {
|
|
"sha": sha,
|
|
"message": message,
|
|
"files_changed": 1,
|
|
"insertions": 1,
|
|
"deletions": 0,
|
|
}
|
|
|
|
async def push_branch(
|
|
self, branch_name: str, *, actor_agent_id: Any = None
|
|
) -> tuple[str, int]:
|
|
del branch_name, actor_agent_id
|
|
return ("ok", 0)
|
|
|
|
async def push_task_branch(self, agent_id: UUID, task_id: UUID) -> int:
|
|
del agent_id, task_id
|
|
return 0
|
|
|
|
async def create_pr(
|
|
self,
|
|
branch_name: str,
|
|
*,
|
|
parent: str,
|
|
is_root_pr: bool,
|
|
actor_agent_id: Any = None,
|
|
) -> dict[str, Any]:
|
|
del branch_name, parent, actor_agent_id
|
|
self._task.pr_number = _PR_NUMBER
|
|
self._task.pr_url = _PR_URL
|
|
self._task.pr_created = True
|
|
await self._session.flush()
|
|
return {"pr_number": _PR_NUMBER, "pr_url": _PR_URL, "is_root_pr": is_root_pr}
|
|
|
|
async def diff(
|
|
self, *, branch_name: str, base: Any = None, actor_agent_id: Any = None
|
|
) -> str:
|
|
del branch_name, base, actor_agent_id
|
|
return "stub diff"
|
|
|
|
async def list_changed_files(
|
|
self, *, branch_name: str, base: Any = None, actor_agent_id: Any = None
|
|
) -> list[str]:
|
|
del branch_name, base, actor_agent_id
|
|
return []
|
|
|
|
async def pr_target(self, pr_number: int, *, actor_agent_id: Any = None) -> str:
|
|
del pr_number, actor_agent_id
|
|
return "main"
|
|
|
|
async def pr_merge(self, *args: Any, **kwargs: Any) -> dict[str, Any]:
|
|
del args, kwargs
|
|
return {
|
|
"merged": True,
|
|
"sha": uuid4().hex[:40],
|
|
"merge_commit_sha": uuid4().hex[:40],
|
|
}
|
|
|
|
|
|
def _mock_evidence_repo() -> Any:
|
|
repo = AsyncMock()
|
|
for method in (
|
|
"list_unread_a2a",
|
|
"list_unread_mentions",
|
|
"list_pending_notifications",
|
|
"task_metadata_gaps",
|
|
"recent_team_activity",
|
|
"blockers_in_lane",
|
|
"journal_highlights_for_task",
|
|
):
|
|
getattr(repo, method).return_value = []
|
|
return repo
|
|
|
|
|
|
def _mock_journal_with_reflect() -> Any:
|
|
"""Journal stub that reports reflect/learning/decision entries present.
|
|
|
|
``latest_decision_at`` is anchored to ``datetime.now(UTC)`` so the C8
|
|
recency window on the PM-decision gate accepts it.
|
|
"""
|
|
journal = AsyncMock()
|
|
journal.has_reflect_for_task.return_value = True
|
|
journal.has_learning_for_task.return_value = True
|
|
journal.has_decision_for_task.return_value = True
|
|
journal.has_struggle_for_task.return_value = False
|
|
journal.latest_decision_at.return_value = datetime.now(UTC)
|
|
return journal
|
|
|
|
|
|
def _mock_work_session() -> Any:
|
|
"""WorkSession stub: stable file list, no unpushed commits."""
|
|
ws = AsyncMock()
|
|
ws.files_changed.return_value = ["roboco/api/routes/health.py"]
|
|
ws.has_unpushed_commits.return_value = False
|
|
return ws
|
|
|
|
|
|
def _build_choreographer(
|
|
db_session: Any, task: TaskTable, task_service: TaskService
|
|
) -> Choreographer:
|
|
"""Wire a real Choreographer with the supplied TaskService + stubbed git.
|
|
|
|
Caller owns the TaskService so it can use it for direct DB reads
|
|
(``task_service.get(task_id)``) — sharing one instance keeps the
|
|
session contract clean and avoids "two TaskServices, two views"
|
|
surprises.
|
|
"""
|
|
deps = ChoreographerDeps(
|
|
task=task_service,
|
|
work_session=_mock_work_session(),
|
|
git=_StubGit(db_session, task),
|
|
a2a=AsyncMock(),
|
|
journal=_mock_journal_with_reflect(),
|
|
audit=AsyncMock(),
|
|
evidence_repo=_mock_evidence_repo(),
|
|
)
|
|
return Choreographer(deps)
|
|
|
|
|
|
async def _seed_agents_and_project(
|
|
db_session: AsyncSession,
|
|
) -> dict[str, Any]:
|
|
"""Seed system + project + dev/qa/doc/cell_pm agents.
|
|
|
|
Slugs match ``agents_config.ESCALATION_CHAIN`` so ``i_am_blocked``
|
|
and ``escalate_up`` find a real escalation target. The agents-config
|
|
chain is the source of truth at runtime; matching it here exercises
|
|
the same lookup the gateway uses in production.
|
|
"""
|
|
system_agent = AgentTable(
|
|
id=uuid4(),
|
|
name="System",
|
|
slug=f"system-{uuid4().hex[:8]}",
|
|
role=AgentRole.SYSTEM,
|
|
team=None,
|
|
status=AgentStatus.ACTIVE,
|
|
model_config={},
|
|
system_prompt="system",
|
|
capabilities=[],
|
|
permissions={},
|
|
metrics={},
|
|
)
|
|
db_session.add(system_agent)
|
|
await db_session.flush()
|
|
|
|
project = ProjectTable(
|
|
id=uuid4(),
|
|
name="Lifecycle Test Project",
|
|
slug=f"life-{uuid4().hex[:8]}",
|
|
git_url="https://github.com/example/life.git",
|
|
default_branch="main",
|
|
protected_branches=["main"],
|
|
assigned_cell=Team.BACKEND,
|
|
created_by=system_agent.id,
|
|
is_active=True,
|
|
)
|
|
db_session.add(project)
|
|
await db_session.flush()
|
|
|
|
dev_agent = AgentTable(
|
|
id=uuid4(),
|
|
name="BE Dev 1",
|
|
slug="be-dev-1",
|
|
role=AgentRole.DEVELOPER,
|
|
team=Team.BACKEND,
|
|
status=AgentStatus.ACTIVE,
|
|
model_config={},
|
|
system_prompt="dev",
|
|
capabilities=["python"],
|
|
permissions={},
|
|
metrics={},
|
|
)
|
|
qa_agent = AgentTable(
|
|
id=uuid4(),
|
|
name="BE QA",
|
|
slug="be-qa",
|
|
role=AgentRole.QA,
|
|
team=Team.BACKEND,
|
|
status=AgentStatus.ACTIVE,
|
|
model_config={},
|
|
system_prompt="qa",
|
|
capabilities=["review"],
|
|
permissions={},
|
|
metrics={},
|
|
)
|
|
doc_agent = AgentTable(
|
|
id=uuid4(),
|
|
name="BE Doc",
|
|
slug="be-doc",
|
|
role=AgentRole.DOCUMENTER,
|
|
team=Team.BACKEND,
|
|
status=AgentStatus.ACTIVE,
|
|
model_config={},
|
|
system_prompt="doc",
|
|
capabilities=["docs"],
|
|
permissions={},
|
|
metrics={},
|
|
)
|
|
cell_pm_agent = AgentTable(
|
|
id=uuid4(),
|
|
name="BE Cell PM",
|
|
slug="be-pm",
|
|
role=AgentRole.CELL_PM,
|
|
team=Team.BACKEND,
|
|
status=AgentStatus.ACTIVE,
|
|
model_config={},
|
|
system_prompt="cell_pm",
|
|
capabilities=["coord"],
|
|
permissions={},
|
|
metrics={},
|
|
)
|
|
db_session.add_all([dev_agent, qa_agent, doc_agent, cell_pm_agent])
|
|
await db_session.flush()
|
|
|
|
return {
|
|
"system_agent": system_agent,
|
|
"project": project,
|
|
"dev_agent": dev_agent,
|
|
"qa_agent": qa_agent,
|
|
"doc_agent": doc_agent,
|
|
"cell_pm_agent": cell_pm_agent,
|
|
}
|
|
|
|
|
|
def _build_task(
|
|
*,
|
|
project_id: UUID,
|
|
creator_id: UUID,
|
|
assignee_id: UUID | None,
|
|
status: TaskStatus,
|
|
) -> TaskTable:
|
|
"""Construct a backend code-typed task pinned to ``status``.
|
|
|
|
``acceptance_criteria_status`` carries the stub artefact rows the
|
|
pre-merge gate inspects; supplying them here keeps the per-test
|
|
setup readable.
|
|
"""
|
|
return TaskTable(
|
|
id=uuid4(),
|
|
title="Add /healthz endpoint",
|
|
description="Return 200 OK from /healthz",
|
|
status=status,
|
|
priority=2,
|
|
task_type=TaskType.CODE,
|
|
nature=TaskNature.TECHNICAL,
|
|
team=Team.BACKEND,
|
|
project_id=project_id,
|
|
created_by=creator_id,
|
|
assigned_to=assignee_id,
|
|
branch_name=_BRANCH,
|
|
acceptance_criteria=["Returns 200", "Includes timestamp"],
|
|
acceptance_criteria_status=[
|
|
{"criterion": "Returns 200", "referencing_artifact_id": "stub"},
|
|
{"criterion": "Includes timestamp", "referencing_artifact_id": "stub"},
|
|
],
|
|
)
|
|
|
|
|
|
@pytest_asyncio.fixture
|
|
async def lifecycle_setup(
|
|
db_session: AsyncSession,
|
|
) -> AsyncIterator[dict[str, Any]]:
|
|
"""Seed agents + project + a single PENDING task assigned to the dev."""
|
|
seeded = await _seed_agents_and_project(db_session)
|
|
task = _build_task(
|
|
project_id=seeded["project"].id,
|
|
creator_id=seeded["system_agent"].id,
|
|
assignee_id=seeded["dev_agent"].id,
|
|
status=TaskStatus.PENDING,
|
|
)
|
|
db_session.add(task)
|
|
await db_session.flush()
|
|
seeded["task"] = task
|
|
yield seeded
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# 1. Dev path: pending → claimed → in_progress → verifying → awaiting_qa
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_dev_full_chain_through_awaiting_qa(
|
|
db_session: AsyncSession, lifecycle_setup: dict[str, Any]
|
|
) -> None:
|
|
"""pending → claimed → in_progress → verifying → awaiting_qa.
|
|
|
|
Drives ``i_will_work_on`` (claim+plan+start), a stubbed commit,
|
|
``open_pr``, then ``i_am_done`` which auto-runs submit_verification
|
|
+ submit_qa. Asserts the final DB row sits at ``awaiting_qa``.
|
|
"""
|
|
task = lifecycle_setup["task"]
|
|
dev_agent = lifecycle_setup["dev_agent"]
|
|
task_service = TaskService(db_session)
|
|
stub_git = _StubGit(db_session, task)
|
|
deps = ChoreographerDeps(
|
|
task=task_service,
|
|
work_session=_mock_work_session(),
|
|
git=stub_git,
|
|
a2a=AsyncMock(),
|
|
journal=_mock_journal_with_reflect(),
|
|
audit=AsyncMock(),
|
|
evidence_repo=_mock_evidence_repo(),
|
|
)
|
|
c = Choreographer(deps)
|
|
|
|
env = await c.i_will_work_on(
|
|
dev_agent.id,
|
|
task.id,
|
|
plan=_GOOD_PLAN,
|
|
steps=_STEPS,
|
|
technical_considerations=_GOOD_TC,
|
|
risks=_GOOD_RISKS,
|
|
)
|
|
assert env.error is None, f"i_will_work_on failed: {env.message}"
|
|
assert env.status == Status.IN_PROGRESS.value
|
|
|
|
# Commit + record progress so open_pr's commits-precondition holds.
|
|
await stub_git.commit(
|
|
branch_name=_BRANCH,
|
|
message=f"[{str(task.id)[:8]}] feat(api): add /healthz",
|
|
task_id=task.id,
|
|
)
|
|
await task_service.add_progress(task.id, dev_agent.id, "implemented /healthz")
|
|
|
|
env = await c.open_pr(dev_agent.id, task.id)
|
|
assert env.error is None, f"open_pr failed: {env.message}"
|
|
|
|
env = await c.i_am_done(dev_agent.id, task.id, "tests pass; route works")
|
|
assert env.error is None, f"i_am_done failed: {env.message}"
|
|
assert env.status == Status.AWAITING_QA.value
|
|
|
|
final = await task_service.get(task.id)
|
|
assert final is not None
|
|
assert str(final.status) == Status.AWAITING_QA.value
|
|
# i_am_done auto-runs submit_qa, which hands the task off to the
|
|
# backend QA agent (production behaviour — see ``_notify_qa``).
|
|
# Resolve via the same lookup the choreographer uses so the
|
|
# assertion is robust to other test fixtures that may have seeded
|
|
# additional QA agents on the BACKEND team (e.g. smoke_test_batch
|
|
# commits its agents, so they outlive their session).
|
|
resolved_qa = await task_service.qa_agent_for_team(Team.BACKEND)
|
|
assert resolved_qa is not None
|
|
assert final.assigned_to == resolved_qa.id
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# 2. QA pass path: awaiting_qa → claimed → awaiting_documentation
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_qa_pass_path(
|
|
db_session: AsyncSession, lifecycle_setup: dict[str, Any]
|
|
) -> None:
|
|
"""awaiting_qa → claim_review → pass_review → awaiting_documentation.
|
|
|
|
``claim_review`` keeps status at AWAITING_QA (specialised qa_claim
|
|
sets assignment without transitioning) so the spec's source-status
|
|
requirement on ``qa_pass`` still matches downstream.
|
|
"""
|
|
task = lifecycle_setup["task"]
|
|
qa_agent = lifecycle_setup["qa_agent"]
|
|
doc_agent = lifecycle_setup["doc_agent"]
|
|
|
|
# Pin the task at awaiting_qa with a PR + commit so the QA gates
|
|
# (pr exists, commits non-empty) all pass.
|
|
task.status = TaskStatus.AWAITING_QA
|
|
task.pr_number = _PR_NUMBER
|
|
task.pr_url = _PR_URL
|
|
task.commits = [
|
|
{"sha": uuid4().hex[:40], "message": "feat: /healthz", "task_id": str(task.id)}
|
|
]
|
|
task.self_verified = True
|
|
await db_session.flush()
|
|
|
|
task_service = TaskService(db_session)
|
|
c = _build_choreographer(db_session, task, task_service)
|
|
|
|
env = await c.claim_review(qa_agent.id, task.id)
|
|
assert env.error is None, f"claim_review failed: {env.message}"
|
|
after_claim = await task_service.get(task.id)
|
|
assert after_claim is not None
|
|
assert str(after_claim.status) == Status.AWAITING_QA.value
|
|
assert after_claim.assigned_to == qa_agent.id
|
|
|
|
env = await c.pass_review(qa_agent.id, task.id, notes=_QA_PASS_NOTES)
|
|
assert env.error is None, f"pass_review failed: {env.message}"
|
|
assert env.status == Status.AWAITING_DOCUMENTATION.value
|
|
|
|
final = await task_service.get(task.id)
|
|
assert final is not None
|
|
assert str(final.status) == Status.AWAITING_DOCUMENTATION.value
|
|
# pass_review reassigns to the team's documenter for handoff. Look
|
|
# up via the same path the choreographer uses (robust to other
|
|
# tests' committed BACKEND documenters).
|
|
resolved_doc = await task_service.documenter_for_team(Team.BACKEND)
|
|
assert resolved_doc is not None
|
|
assert final.assigned_to == resolved_doc.id
|
|
del doc_agent # asserted indirectly via documenter_for_team.
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# 3. QA fail path: awaiting_qa → claimed → needs_revision
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_qa_fail_path(
|
|
db_session: AsyncSession, lifecycle_setup: dict[str, Any]
|
|
) -> None:
|
|
"""awaiting_qa → claim_review → fail_review(issues) → needs_revision.
|
|
|
|
``fail_review`` reassigns to the original developer so they can
|
|
revise; that lookup walks ``quick_context``'s
|
|
``original_developer:<slug>`` marker, which the dev path stamps
|
|
on i_will_work_on. Here we set it directly so the assertion is
|
|
deterministic without re-running the dev chain.
|
|
"""
|
|
task = lifecycle_setup["task"]
|
|
qa_agent = lifecycle_setup["qa_agent"]
|
|
dev_agent = lifecycle_setup["dev_agent"]
|
|
|
|
task.status = TaskStatus.AWAITING_QA
|
|
task.pr_number = _PR_NUMBER
|
|
task.pr_url = _PR_URL
|
|
task.commits = [
|
|
{"sha": uuid4().hex[:40], "message": "feat: /healthz", "task_id": str(task.id)}
|
|
]
|
|
task.self_verified = True
|
|
# ``extract_original_developer`` parses a UUID off this line; the
|
|
# spec layer's slug-based self-review check is a separate code path
|
|
# (``_extract_original_developer`` in qa.py) which only fires when
|
|
# an actor's slug equals this value, so a UUID here doesn't trip it.
|
|
task.quick_context = f"original_developer:{dev_agent.id}"
|
|
await db_session.flush()
|
|
|
|
task_service = TaskService(db_session)
|
|
c = _build_choreographer(db_session, task, task_service)
|
|
|
|
env = await c.claim_review(qa_agent.id, task.id)
|
|
assert env.error is None, f"claim_review failed: {env.message}"
|
|
|
|
issues = ["Returns 500 on the timestamp branch", "Missing test for the second AC"]
|
|
# fail_review's notes are derived from issues; QA pass-gate also
|
|
# requires notes >= 80 chars, so we send a leading explanation as
|
|
# the issues list — the verb concatenates them and easily clears
|
|
# the threshold.
|
|
long_issues = [_QA_FAIL_NOTES_PREFIX + issues[0], issues[1]]
|
|
env = await c.fail_review(qa_agent.id, task.id, issues=long_issues)
|
|
assert env.error is None, f"fail_review failed: {env.message}"
|
|
assert env.status == Status.NEEDS_REVISION.value
|
|
|
|
final = await task_service.get(task.id)
|
|
assert final is not None
|
|
assert str(final.status) == Status.NEEDS_REVISION.value
|
|
# fail_qa reassigns to the original developer so they can revise.
|
|
assert final.assigned_to == dev_agent.id
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# 4. Doc path: awaiting_documentation → claimed → awaiting_pm_review
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_doc_path(
|
|
db_session: AsyncSession, lifecycle_setup: dict[str, Any]
|
|
) -> None:
|
|
"""awaiting_documentation → claim_doc_task → i_documented → awaiting_pm_review.
|
|
|
|
``claim_doc_task`` keeps status at AWAITING_DOCUMENTATION (doc_claim
|
|
is assignment-only, mirroring qa_claim). ``i_documented`` flips to
|
|
AWAITING_PM_REVIEW and reassigns to the cell PM for that team.
|
|
"""
|
|
task = lifecycle_setup["task"]
|
|
doc_agent = lifecycle_setup["doc_agent"]
|
|
cell_pm_agent = lifecycle_setup["cell_pm_agent"]
|
|
|
|
task.status = TaskStatus.AWAITING_DOCUMENTATION
|
|
task.pr_number = _PR_NUMBER
|
|
task.pr_url = _PR_URL
|
|
task.pr_created = True
|
|
task.qa_verified = True
|
|
task.assigned_to = None # documenter must claim from unassigned.
|
|
task.commits = [
|
|
{"sha": uuid4().hex[:40], "message": "feat: /healthz", "task_id": str(task.id)}
|
|
]
|
|
await db_session.flush()
|
|
|
|
task_service = TaskService(db_session)
|
|
c = _build_choreographer(db_session, task, task_service)
|
|
|
|
env = await c.claim_doc_task(doc_agent.id, task.id)
|
|
assert env.error is None, f"claim_doc_task failed: {env.message}"
|
|
after_claim = await task_service.get(task.id)
|
|
assert after_claim is not None
|
|
assert str(after_claim.status) == Status.AWAITING_DOCUMENTATION.value
|
|
assert after_claim.assigned_to == doc_agent.id
|
|
|
|
env = await c.i_documented(
|
|
doc_agent.id,
|
|
task.id,
|
|
notes="Documented /healthz behaviour in docs/api/health.md",
|
|
files=["docs/api/health.md"],
|
|
)
|
|
assert env.error is None, f"i_documented failed: {env.message}"
|
|
assert env.status == Status.AWAITING_PM_REVIEW.value
|
|
|
|
final = await task_service.get(task.id)
|
|
assert final is not None
|
|
assert str(final.status) == Status.AWAITING_PM_REVIEW.value
|
|
# i_documented hands off to the cell PM for the team. Resolve via
|
|
# the same lookup the choreographer uses (robust to other tests'
|
|
# committed BACKEND cell PMs).
|
|
resolved_pm = await task_service.cell_pm_for_team(Team.BACKEND)
|
|
assert resolved_pm is not None
|
|
assert final.assigned_to == resolved_pm.id
|
|
del cell_pm_agent # asserted indirectly via cell_pm_for_team.
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# 5. PM complete (Cell PM, simple task): awaiting_pm_review → completed
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_pm_complete_simple_task(
|
|
db_session: AsyncSession, lifecycle_setup: dict[str, Any]
|
|
) -> None:
|
|
"""awaiting_pm_review → cell_pm complete → completed.
|
|
|
|
A cell PM completing a non-root awaiting_pm_review task transitions
|
|
straight to COMPLETED — there is no cell→main escalation in
|
|
``complete`` (the cell→main hand-off, when intended, uses
|
|
``submit_up``).
|
|
"""
|
|
task = lifecycle_setup["task"]
|
|
cell_pm_agent = lifecycle_setup["cell_pm_agent"]
|
|
|
|
task.status = TaskStatus.AWAITING_PM_REVIEW
|
|
task.pr_number = _PR_NUMBER
|
|
task.pr_url = _PR_URL
|
|
task.pr_created = True
|
|
task.qa_verified = True
|
|
task.docs_complete = True
|
|
task.assigned_to = cell_pm_agent.id
|
|
task.commits = [
|
|
{"sha": uuid4().hex[:40], "message": "feat: /healthz", "task_id": str(task.id)}
|
|
]
|
|
await db_session.flush()
|
|
|
|
task_service = TaskService(db_session)
|
|
c = _build_choreographer(db_session, task, task_service)
|
|
|
|
env = await c.complete(
|
|
cell_pm_agent.id, task.id, notes="LGTM — merging the leaf PR."
|
|
)
|
|
assert env.error is None, f"complete failed: {env.message}"
|
|
assert env.status == Status.COMPLETED.value
|
|
|
|
final = await task_service.get(task.id)
|
|
assert final is not None
|
|
assert str(final.status) == Status.COMPLETED.value
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# 6. PM escalate: awaiting_pm_review → awaiting_ceo_approval
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_pm_escalate_to_ceo_path(
|
|
db_session: AsyncSession, lifecycle_setup: dict[str, Any]
|
|
) -> None:
|
|
"""awaiting_pm_review → main_pm complete on a root task → awaiting_ceo_approval.
|
|
|
|
Main PM completing a root task (no parent) opens the master PR if
|
|
needed and calls ``task.escalate_to_ceo``, leaving the task at
|
|
AWAITING_CEO_APPROVAL with ``assigned_to=None``. CEO approval is
|
|
human-in-the-loop (UI-driven), so the test stops there.
|
|
"""
|
|
task = lifecycle_setup["task"]
|
|
project = lifecycle_setup["project"]
|
|
system_agent = lifecycle_setup["system_agent"]
|
|
|
|
main_pm_agent = AgentTable(
|
|
id=uuid4(),
|
|
name="Main PM",
|
|
slug="main-pm",
|
|
role=AgentRole.MAIN_PM,
|
|
team=None,
|
|
status=AgentStatus.ACTIVE,
|
|
model_config={},
|
|
system_prompt="main_pm",
|
|
capabilities=["coord"],
|
|
permissions={},
|
|
metrics={},
|
|
)
|
|
db_session.add(main_pm_agent)
|
|
await db_session.flush()
|
|
del project, system_agent # only needed for fixture wiring above.
|
|
|
|
task.status = TaskStatus.AWAITING_PM_REVIEW
|
|
task.pr_number = _PR_NUMBER
|
|
task.pr_url = _PR_URL
|
|
task.pr_created = True
|
|
task.qa_verified = True
|
|
task.docs_complete = True
|
|
task.parent_task_id = None # explicit — escalate_to_ceo refuses subtasks.
|
|
task.assigned_to = main_pm_agent.id
|
|
task.commits = [
|
|
{"sha": uuid4().hex[:40], "message": "feat: /healthz", "task_id": str(task.id)}
|
|
]
|
|
await db_session.flush()
|
|
|
|
task_service = TaskService(db_session)
|
|
c = _build_choreographer(db_session, task, task_service)
|
|
|
|
# SQLAlchemy column-typed `id` attributes need an explicit UUID
|
|
# cast for mypy under the project's strict config — the values are
|
|
# already real ``uuid.UUID`` at runtime.
|
|
env = await c.complete(
|
|
UUID(str(main_pm_agent.id)),
|
|
UUID(str(task.id)),
|
|
notes="Root task ready for CEO approval — escalating.",
|
|
)
|
|
assert env.error is None, f"complete failed: {env.message}"
|
|
assert env.status == Status.AWAITING_CEO_APPROVAL.value
|
|
|
|
final = await task_service.get(task.id)
|
|
assert final is not None
|
|
assert str(final.status) == Status.AWAITING_CEO_APPROVAL.value
|
|
# main_pm_complete clears assigned_to so the orchestrator does not
|
|
# respawn an agent while the task waits on the human CEO.
|
|
assert final.assigned_to is None
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# 7. Block + unblock + restore: in_progress → blocked → in_progress
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_block_then_unblock_restore(
|
|
db_session: AsyncSession, lifecycle_setup: dict[str, Any]
|
|
) -> None:
|
|
"""in_progress → i_am_blocked → unblock(restore=True) → in_progress.
|
|
|
|
``i_am_blocked`` runs the spec's ``block`` action which delegates
|
|
to ``task_service.escalate``: the task is reassigned to the dev's
|
|
escalation target (``be-pm`` per ``ESCALATION_CHAIN``) and marked
|
|
BLOCKED with the original dev stashed in ``blocker_raised_by``.
|
|
PM ``unblock(restore=True)`` then falls through to the legacy
|
|
``unblock`` path (no ``pre_block_state`` snapshot exists for chain
|
|
escalations), restoring assignment to the dev and flipping back
|
|
to IN_PROGRESS.
|
|
"""
|
|
task = lifecycle_setup["task"]
|
|
dev_agent = lifecycle_setup["dev_agent"]
|
|
cell_pm_agent = lifecycle_setup["cell_pm_agent"]
|
|
|
|
task_service = TaskService(db_session)
|
|
c = _build_choreographer(db_session, task, task_service)
|
|
|
|
# Drive into in_progress via the real claim+start sequence.
|
|
env = await c.i_will_work_on(
|
|
dev_agent.id,
|
|
task.id,
|
|
plan=_GOOD_PLAN,
|
|
steps=_STEPS,
|
|
technical_considerations=_GOOD_TC,
|
|
risks=_GOOD_RISKS,
|
|
)
|
|
assert env.error is None, f"i_will_work_on failed: {env.message}"
|
|
assert env.status == Status.IN_PROGRESS.value
|
|
|
|
env = await c.i_am_blocked(
|
|
dev_agent.id,
|
|
task.id,
|
|
reason="external dependency on auth library upgrade",
|
|
)
|
|
assert env.error is None, f"i_am_blocked failed: {env.message}"
|
|
assert env.status == Status.BLOCKED.value
|
|
|
|
blocked = await task_service.get(task.id)
|
|
assert blocked is not None
|
|
assert str(blocked.status) == Status.BLOCKED.value
|
|
# Escalation reassigns to the cell PM (be-pm) and stashes the dev
|
|
# as blocker_raised_by so unblock can hand the task back.
|
|
assert blocked.assigned_to == cell_pm_agent.id
|
|
assert blocked.blocker_raised_by == dev_agent.id
|
|
|
|
env = await c.unblock(cell_pm_agent.id, task.id, restore=True)
|
|
assert env.error is None, f"unblock failed: {env.message}"
|
|
assert env.status == Status.IN_PROGRESS.value
|
|
|
|
restored = await task_service.get(task.id)
|
|
assert restored is not None
|
|
assert str(restored.status) == Status.IN_PROGRESS.value
|
|
# legacy unblock restores assigned_to from blocker_raised_by — the
|
|
# original dev gets the task back so the orchestrator respawns them.
|
|
assert restored.assigned_to == dev_agent.id
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# 8. Pause + resume: in_progress → paused → in_progress
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_pause_then_resume(
|
|
db_session: AsyncSession, lifecycle_setup: dict[str, Any]
|
|
) -> None:
|
|
"""in_progress → i_am_idle (auto-pause) → resume → in_progress.
|
|
|
|
There is no agent-driven ``pause`` verb. ``i_am_idle`` auto-pauses
|
|
every in_progress task the agent owns so the closure dispatcher can
|
|
wake them on respawn. ``resume`` (composes=("resume",)) then flips
|
|
the same task back to IN_PROGRESS for the same assignee.
|
|
"""
|
|
task = lifecycle_setup["task"]
|
|
dev_agent = lifecycle_setup["dev_agent"]
|
|
|
|
task_service = TaskService(db_session)
|
|
c = _build_choreographer(db_session, task, task_service)
|
|
|
|
env = await c.i_will_work_on(
|
|
dev_agent.id,
|
|
task.id,
|
|
plan=_GOOD_PLAN,
|
|
steps=_STEPS,
|
|
technical_considerations=_GOOD_TC,
|
|
risks=_GOOD_RISKS,
|
|
)
|
|
assert env.error is None, f"i_will_work_on failed: {env.message}"
|
|
assert env.status == Status.IN_PROGRESS.value
|
|
|
|
env = await c.i_am_idle(dev_agent.id)
|
|
assert env.error is None, f"i_am_idle failed: {env.message}"
|
|
assert env.status == "idle"
|
|
|
|
paused = await task_service.get(task.id)
|
|
assert paused is not None
|
|
assert str(paused.status) == Status.PAUSED.value
|
|
# Auto-pause keeps assigned_to so resume can find the same claimant.
|
|
assert paused.assigned_to == dev_agent.id
|
|
|
|
env = await c.resume(dev_agent.id, task.id)
|
|
assert env.error is None, f"resume failed: {env.message}"
|
|
assert env.status == Status.IN_PROGRESS.value
|
|
|
|
resumed = await task_service.get(task.id)
|
|
assert resumed is not None
|
|
assert str(resumed.status) == Status.IN_PROGRESS.value
|
|
assert resumed.assigned_to == dev_agent.id
|