Leak fixes, gate green again, uv/CI hardening, e2e lifecycle smoke harness (#294)

* test: align phase1 smoke mock with the armed team-match gate

The 8e5f84c4 sweep fixed 13 test files' inconsistent-team mocks but ran
only the gateway/foundation/runtime subsets; the full gate caught this
integration mock whose parent task carried an auto-generated MagicMock
team and died on not_authorized before the incomplete_input assertion.

* fix(runtime): attribute every agent.spawned audit to its dispatcher

A rogue spawner could not be identified live (2026-07-02): agent.spawned
rows carry container/model but not which dispatch loop launched them.
spawn_agent now takes spawned_by, stamps it into the spawned/spawn_failed
audit details, every call site passes its loop name, and an AST sweep
test holds future callers to it.

* fix(api): admin-complete refuses when the task's PR is still open

PATCH status=completed on a task with an OPEN PR stranded its commits
unmerged (bit the CEO twice live 2026-07-02). The override now refuses
with the PR number/URL and the consequence before the generic hatch
text; force:true stays the deliberate, audited escape.

* fix(panel): awaiting_ceo_approval offers the working ceo-approve path

The header's only approve action was Approve & Merge (POST
/approve-and-merge, no notes) which 400s NO_PR on a branchless MegaTask
umbrella — the CEO's approve button just failed. Primary action is now
Approve & Complete via the CeoApproveDialog (POST /ceo-approve, notes
>=20 chars, proven live); Approve & Merge stays for PR-bearing tasks.

* test: stop leaking self-heal + rate-limit state into live Redis

Two test files wrote real keys into a developer's localhost Redis:
self-heal originate tests left self_heal:notified:* (2h TTL) and the
i_am_blocked rate-limited tests left a NO-TTL 'anthropic rate-limited'
tracker blob — order/state-dependent poison for anything reading the
real tracker, and the prime suspect class for the one-off
test_self_heal_engine full-run failure (not reproduced in 5x dir runs,
adversarial orders, and a green full gate). Both files now point the
computed redis_url at an unreachable port; the engines' fail-open paths
keep every assertion intact. Leaked keys scrubbed live.

* docs: changelog + map delta for the leak-fix batch; mypy-clean attribution test

The attribution test's direct method assignments tripped the full gate's
mypy (method-assign) — switched to the house monkeypatch idiom, no
suppressions.

* fix(gate): clear the ten xenon C-ranks; isolate all tests from live Redis

Master CI has been red at the phase1 smoke test, so neither CI nor a
local full gate had reached the xenon step since the team-match sweep —
whose inline 'agent_team=str(agent.team) if ...' kwarg pushed nine verb
bodies from B(10) to C(11-12) unseen. A shared actor_context_fields()
(_protocol.py) computes (actor_slug, agent_team) once per verb, restoring
all nine to B with zero behavior change; the new admin-complete override
helper extraction does the same for routes/tasks.py.

tests/conftest.py gains an autouse fixture pointing the computed
redis_url at an unreachable port for every test — the root fix for the
three families caught writing live-Redis keys (self-heal dedupe,
rate-limit tracker, notification purpose-dedupe); no test uses a real
Redis, and every production path is fail-open by design.

* refactor(runtime): delete the never-wired dispatch-time spawn cooldown

_safe_spawn / gateway_pre_spawn_check / trigger_filter had no caller in
the repo's entire history (87ef42bf only flipped the flag). Its five
rules are superseded: provider parking runs inside spawn_agent, claim
freshness is the guards+reaper, runaway respawns are the progress-aware
breaker + notification cooldown; the per-task cooldown rule would
queue-stall every normal stage handoff if wired today. gateway_triggers
table kept inert. Ratified by the CEO over wiring it.

* build: serialize uv — gate recipes never implicitly sync the venv

Every uv run re-syncs implicitly, so a background make quality plus any
foreground uv run raced two writers on one .venv and tore site-packages
apart (the recurring rich/pip/bandit ImportError corruption; bit twice
today, four times on 2026-07-02's first session). UV_NO_SYNC=1 is now
exported Makefile-wide and quality/quality-fast/gate depend on one
explicit up-front sync step.

* fix(git): PR/merge/branch REST calls honor github_api_base_url

Fifteen sites hardcoded https://api.github.com while the CI-run and
open-PR-list calls already read settings.github_api_base_url — a GHE or
test override silently applied to half the surface. One _api_base()
helper keeps them uniform; default behavior unchanged.

* ci: split the monolith — backend CI, Panel CI, E2E Smoke

ci.yml keeps its file name and the backend quality job only (self-heal /
ci-watch / release-readiness default to the ci.yml workflow); the panel
job moves to panel-ci.yml scoped to panel/**, and the new scripted-agent
lifecycle smoke gets e2e-smoke.yml + a make e2e-smoke target (env-gated
out of the default pytest run). Trade: a panel-only red now lands on
Panel CI, which the ci.yml-pinned watch engines don't see.

* feat(tests): e2e lifecycle smoke harness — scripted agents, real gates

tests/e2e_smoke stands up the real API (flow/do routers + middleware on
uvicorn) over the ephemeral test Postgres, a local bare origin standing
in for GitHub, and a fake GitHub REST layer whose merges are real git
merges. A deterministic driver reloads the real MCP flow/do modules per
agent and walks claim (real clone + worktree) -> tracing-gap -> note ->
plan gate -> commit -> PR -> the full i_am_done ladder -> QA verdicts ->
documenter -> awaiting_pm_review in ~5s. Runs via make e2e-smoke + its
own CI workflow; skipped (env-gated) in the default suite. The
freeze-lift condition's first half: scenario 1 green.

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
This commit is contained in:
Renzo F
2026-07-02 18:28:07 +02:00
committed by GitHub
co-authored by Renn F
parent fe67a630ac
commit 1c87a4e4e4
35 changed files with 1661 additions and 835 deletions
-44
View File
@@ -16,7 +16,6 @@ on:
- 'pyproject.toml' - 'pyproject.toml'
- 'uv.lock' - 'uv.lock'
- 'alembic.ini' - 'alembic.ini'
- 'panel/**'
- '.github/workflows/ci.yml' - '.github/workflows/ci.yml'
pull_request: pull_request:
branches: branches:
@@ -33,7 +32,6 @@ on:
- 'pyproject.toml' - 'pyproject.toml'
- 'uv.lock' - 'uv.lock'
- 'alembic.ini' - 'alembic.ini'
- 'panel/**'
- '.github/workflows/ci.yml' - '.github/workflows/ci.yml'
workflow_dispatch: workflow_dispatch:
@@ -120,45 +118,3 @@ jobs:
- name: Run quality gate - name: Run quality gate
run: make quality run: make quality
panel:
name: Panel (Next.js)
runs-on: ubuntu-latest
defaults:
run:
working-directory: panel
steps:
- name: Checkout code
uses: actions/checkout@v7
with:
# The release-readiness smoke test calls ``git describe --tags`` to
# find the most recent release tag. ``actions/checkout``'s default
# shallow + no-tags clone makes that return empty, which made
# ``test_gather_snapshot_reads_the_real_repo`` fail with
# ``last_tag is None`` even though master had a tagged v0.13.0.
# ``fetch-depth: 0`` clones full history; the default ``fetch-tags``
# would still skip tags on shallow clones, so we also pin it true.
fetch-depth: 0
fetch-tags: true
- name: Set up Node
uses: actions/setup-node@v6
with:
node-version: '20'
- name: Enable corepack (pnpm)
run: corepack enable
- name: Install dependencies
run: pnpm install --frozen-lockfile
- name: Lint
run: pnpm lint
- name: Type-check
run: pnpm exec tsc --noEmit
- name: Test (vitest + coverage)
run: pnpm test
+82
View File
@@ -0,0 +1,82 @@
name: E2E Smoke
on:
push:
branches:
- master
paths:
- 'roboco/**'
- 'alembic/**'
- 'tests/**'
- 'Makefile'
- 'pyproject.toml'
- 'uv.lock'
- '.github/workflows/e2e-smoke.yml'
pull_request:
branches:
- master
paths:
- 'roboco/**'
- 'alembic/**'
- 'tests/**'
- 'Makefile'
- 'pyproject.toml'
- 'uv.lock'
- '.github/workflows/e2e-smoke.yml'
workflow_dispatch:
jobs:
e2e-smoke:
name: e2e lifecycle smoke (scripted agents)
runs-on: ubuntu-latest
services:
postgres:
image: pgvector/pgvector:pg16
env:
POSTGRES_USER: roboco
POSTGRES_PASSWORD: roboco
POSTGRES_DB: roboco
ports:
- 5432:5432
options: >-
--health-cmd "pg_isready -U roboco"
--health-interval 10s
--health-timeout 5s
--health-retries 5
env:
ROBOCO_DATABASE_HOST: localhost
ROBOCO_DATABASE_PORT: '5432'
ROBOCO_DATABASE_USER: roboco
ROBOCO_DATABASE_PASSWORD: roboco
ROBOCO_DATABASE_NAME: roboco
ROBOCO_ENCRYPTION_KEY: 'yp3Awiv0zmxpRa6Gi9Y9hJbi4pZ2FXHRNr4EI6-Gx9U='
ROBOCO_TEST_DB_HOST: localhost
ROBOCO_TEST_DB_PORT: '5432'
ROBOCO_TEST_DB_USER: roboco
ROBOCO_TEST_DB_PASSWORD: roboco
ROBOCO_TEST_DB_ADMIN_DB: postgres
steps:
- name: Checkout code
uses: actions/checkout@v7
- name: Set up Python
uses: actions/setup-python@v6
with:
python-version: '3.13'
- name: Install uv
run: pip install uv
- name: Install dependencies
run: uv sync --extra dev
- name: Configure git identity for the scripted agents
run: |
git config --global user.name "roboco-e2e"
git config --global user.email "e2e@roboco.local"
- name: Run the lifecycle smoke
run: make e2e-smoke
+59
View File
@@ -0,0 +1,59 @@
name: Panel CI
on:
push:
branches:
- master
paths:
- 'panel/**'
- '.github/workflows/panel-ci.yml'
pull_request:
branches:
- master
paths:
- 'panel/**'
- '.github/workflows/panel-ci.yml'
workflow_dispatch:
jobs:
panel:
name: Panel (Next.js)
runs-on: ubuntu-latest
defaults:
run:
working-directory: panel
steps:
- name: Checkout code
uses: actions/checkout@v7
with:
# The release-readiness smoke test calls ``git describe --tags`` to
# find the most recent release tag. ``actions/checkout``'s default
# shallow + no-tags clone makes that return empty, which made
# ``test_gather_snapshot_reads_the_real_repo`` fail with
# ``last_tag is None`` even though master had a tagged v0.13.0.
# ``fetch-depth: 0`` clones full history; the default ``fetch-tags``
# would still skip tags on shallow clones, so we also pin it true.
fetch-depth: 0
fetch-tags: true
- name: Set up Node
uses: actions/setup-node@v6
with:
node-version: '20'
- name: Enable corepack (pnpm)
run: corepack enable
- name: Install dependencies
run: pnpm install --frozen-lockfile
- name: Lint
run: pnpm lint
- name: Type-check
run: pnpm exec tsc --noEmit
- name: Test (vitest + coverage)
run: pnpm test
+10
View File
@@ -8,6 +8,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
### Added ### Added
- **The e2e lifecycle smoke harness — scripted agents drive the REAL gates, no LLM anywhere.** `tests/e2e_smoke/` stands up the real API (v1 flow/do routers + middleware served by uvicorn) over the ephemeral test Postgres, a local bare git origin standing in for GitHub, and a fake GitHub REST layer whose PR merges are REAL git merges (squash included) on that origin. A deterministic driver reloads the REAL MCP `flow_server`/`do_server` modules per agent (role-scoped manifests from the real `role_config`) and walks the lifecycle through every gate: claim (real clone + per-task worktree) → tracing-gap → claim-note → plan gate → commit → PR → the full `i_am_done` ladder (during-work journal, handoff section, reflect, per-AC artifacts) → QA (learning note + per-criterion verdicts) → documenter → `awaiting_pm_review`, in ~5s. Seam bugs — tool↔gate schema drift, squash merges, stale refs, worktree routing — now die in CI (`make e2e-smoke`, its own workflow) instead of in a live run. Scenario 1 covers the leaf-dev arc; PM-merge/CEO chains extend the same harness.
- **HTTP security hardening — a fastapi-guard layer for public/cloud exposure, default-off and calibrated for RoboCo's own traffic.** RoboCo can now front its API with a [fastapi-guard](https://pypi.org/project/fastapi-guard/) `SecurityMiddleware` + per-route decorator layer (`roboco/security.py`), gated behind `ROBOCO_GUARD_ENABLED` (default **off**) — when off, `create_app` never mounts the middleware and the request path is byte-for-byte unchanged, so the decorators are harmless no-ops. Armed, it adds IP/rate controls, a signature WAF, security headers, cloud-provider/honeypot checks, and an emergency-lockdown kill switch (`ROBOCO_GUARD_EMERGENCY`), plus three RoboCo-specific custom validators the stock WAF cannot cover: **prompt-injection**, **secret-exfil**, and **internal-SSRF** scanning on the prompt-facing and agent-content surfaces. Nine distinct decorators are applied thoughtfully per-surface across ingress and sensitive routes (rate-limit, size caps, content-type, behavior analysis, cloud blocking, honeypot form-traps, usage monitoring, suspicious detection, custom validation). Exposure is env-driven — `enforce_https` follows `ROBOCO_ENVIRONMENT`, so a personal NAS deploy stays relaxed while a cloud host enforces TLS — and telemetry to a guard-core platform is separately gated (`ROBOCO_GUARD_TELEMETRY_ENABLED`, no data leaves the box while off). - **HTTP security hardening — a fastapi-guard layer for public/cloud exposure, default-off and calibrated for RoboCo's own traffic.** RoboCo can now front its API with a [fastapi-guard](https://pypi.org/project/fastapi-guard/) `SecurityMiddleware` + per-route decorator layer (`roboco/security.py`), gated behind `ROBOCO_GUARD_ENABLED` (default **off**) — when off, `create_app` never mounts the middleware and the request path is byte-for-byte unchanged, so the decorators are harmless no-ops. Armed, it adds IP/rate controls, a signature WAF, security headers, cloud-provider/honeypot checks, and an emergency-lockdown kill switch (`ROBOCO_GUARD_EMERGENCY`), plus three RoboCo-specific custom validators the stock WAF cannot cover: **prompt-injection**, **secret-exfil**, and **internal-SSRF** scanning on the prompt-facing and agent-content surfaces. Nine distinct decorators are applied thoughtfully per-surface across ingress and sensitive routes (rate-limit, size caps, content-type, behavior analysis, cloud blocking, honeypot form-traps, usage monitoring, suspicious detection, custom validation). Exposure is env-driven — `enforce_https` follows `ROBOCO_ENVIRONMENT`, so a personal NAS deploy stays relaxed while a cloud host enforces TLS — and telemetry to a guard-core platform is separately gated (`ROBOCO_GUARD_TELEMETRY_ENABLED`, no data leaves the box while off).
- **Scanner honeytrap & auto-ban (Surface N) — two layers, matched to where traffic actually lands.** Behind nginx only `/api`, `/ws`, `/health`, `/ready` reach the orchestrator, so guard can only see (and ban) scanner probes on those paths — the classic root probes (`/.env`, `/wp-login.php`, `/phpmyadmin`, `/.git/config`) hit the panel. So Surface N is split: (1) the guard `threat_ban_config` now carries `recon` / `sensitive_file` / `cms_probing` categories, turning repeated scanner probes on `/api` paths into an adaptive per-IP auto-ban (redis-backed, 24h) once enforcement is active; and (2) nginx drops the classic root scanner paths at the edge with `444` (connection closed, no response) before they reach the panel, anchored to known scanner fingerprints so `/.well-known` and every real route are untouched. The auto-ban only fires in active mode (passive logs the recon hit) and needs redis; the nginx edge-drop is always on. - **Scanner honeytrap & auto-ban (Surface N) — two layers, matched to where traffic actually lands.** Behind nginx only `/api`, `/ws`, `/health`, `/ready` reach the orchestrator, so guard can only see (and ban) scanner probes on those paths — the classic root probes (`/.env`, `/wp-login.php`, `/phpmyadmin`, `/.git/config`) hit the panel. So Surface N is split: (1) the guard `threat_ban_config` now carries `recon` / `sensitive_file` / `cms_probing` categories, turning repeated scanner probes on `/api` paths into an adaptive per-IP auto-ban (redis-backed, 24h) once enforcement is active; and (2) nginx drops the classic root scanner paths at the edge with `444` (connection closed, no response) before they reach the panel, anchored to known scanner fingerprints so `/.well-known` and every real route are untouched. The auto-ban only fires in active mode (passive logs the recon hit) and needs redis; the nginx edge-drop is always on.
- **Guard WAF calibration — active enforcement no longer false-positives on RoboCo's own traffic.** The first end-to-end run of the guard surfaced that active enforcement would block ~50% of legitimate agent traffic: RoboCo's request bodies *are* code, SQL, unified diffs, file paths, HTML, and URLs (task specs, agent notes/commits, RAG queries, git bodies, chat), and the stock signature WAF (SQLi/XSS/path-traversal/URL detectors) flagged them as attacks. `build_security_config` now excludes RoboCo's free-text top-level body fields from WAF scanning (`excluded_detection_body_fields`, derived from the real request models — including the free-form container fields whose nested prose is stringified and scanned), which drops the false-positive rate to zero while keeping the WAF active on every structured (id/enum/slug/branch) field and leaving the custom prompt-injection / secret-exfil / SSRF validators — which run independently of the exclusion — fully in force. A new end-to-end integration test (`tests/unit/test_security_middleware.py`) mounts the real middleware, drives guard's lifespan, and fires real requests to prove: passive mode is genuinely log-only (never blocks), active mode does not false-positive on realistic agent payloads, threats are still blocked even inside excluded fields, and the WAF still fires on non-excluded fields. The NAS composes arm the guard in **passive/log-only** mode (`ROBOCO_GUARD_PASSIVE_MODE=true`, `ROBOCO_GUARD_FAIL_SECURE=false`) so a deploy calibrates against real traffic before any flip to active enforcement. - **Guard WAF calibration — active enforcement no longer false-positives on RoboCo's own traffic.** The first end-to-end run of the guard surfaced that active enforcement would block ~50% of legitimate agent traffic: RoboCo's request bodies *are* code, SQL, unified diffs, file paths, HTML, and URLs (task specs, agent notes/commits, RAG queries, git bodies, chat), and the stock signature WAF (SQLi/XSS/path-traversal/URL detectors) flagged them as attacks. `build_security_config` now excludes RoboCo's free-text top-level body fields from WAF scanning (`excluded_detection_body_fields`, derived from the real request models — including the free-form container fields whose nested prose is stringified and scanned), which drops the false-positive rate to zero while keeping the WAF active on every structured (id/enum/slug/branch) field and leaving the custom prompt-injection / secret-exfil / SSRF validators — which run independently of the exclusion — fully in force. A new end-to-end integration test (`tests/unit/test_security_middleware.py`) mounts the real middleware, drives guard's lifespan, and fires real requests to prove: passive mode is genuinely log-only (never blocks), active mode does not false-positive on realistic agent payloads, threats are still blocked even inside excluded fields, and the WAF still fires on non-excluded fields. The NAS composes arm the guard in **passive/log-only** mode (`ROBOCO_GUARD_PASSIVE_MODE=true`, `ROBOCO_GUARD_FAIL_SECURE=false`) so a deploy calibrates against real traffic before any flip to active enforcement.
@@ -30,6 +31,11 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
### Fixed ### Fixed
- **Every `agent.spawned` audit row names its dispatcher.** A rogue spawner could not be identified live — the audit row carried container/model but not which of the ~27 dispatch loops launched it. `spawn_agent` now takes `spawned_by`, stamps it into the `agent.spawned` / `agent.spawn_failed` details (`"unspecified"` when absent so audit queries never miss the field), every call site passes its loop name, and a whole-package AST sweep test fails any future caller that omits it.
- **Admin-complete refuses while the task's PR is still open.** `PATCH status=completed` on a task whose work session records an OPEN PR stranded its commits unmerged (bit the CEO twice live). The override now refuses with the PR number/URL and the concrete consequence — merge first, or approve via `POST /tasks/{id}/ceo-approve` — checked before the generic hatch text; `force: true` stays the deliberate, audited escape, and a merged/closed PR changes nothing.
- **The panel's CEO-approve button works on every gated task.** On `awaiting_ceo_approval` the task header offered only "Approve & Merge" (`POST /approve-and-merge`, no notes), which 400s `NO_PR` on a branchless MegaTask umbrella — the CEO's approve just failed. The primary action is now "Approve & Complete" through the `CeoApproveDialog` (`POST /ceo-approve`, notes ≥ 20 chars); "Approve & Merge" remains, but only when the task actually carries a PR.
- **Tests no longer leak state into a developer's live Redis.** The self-heal originate tests wrote `self_heal:notified:*` dedupe keys (2h TTL), the `i_am_blocked` rate-limited tests wrote a NO-TTL "anthropic rate-limited" tracker blob, and notification tests left short-TTL purpose-dedupe keys in whatever Redis listens on localhost — order/state-dependent poison for any test (or local orchestrator run) reading the real instance. A root-level autouse fixture now points the computed `redis_url` at an unreachable port for every test (no test uses a real Redis; every production Redis path is fail-open by design), with explicit per-file guards kept at the two proven writers.
- **The quality gate is fully green again — ten latent xenon C-ranks cleared.** Master CI has been red at a smoke test whose mock predated the armed team-match gate, so neither CI nor a local full gate had reached the xenon step — hiding that the team-match sweep's inline `agent_team=str(agent.team) if …` kwarg had pushed nine gateway verb bodies (`i_am_done`, `resume`, `unclaim`, `submit_up`, `submit_root`, `complete`, `escalate_up`, `escalate_to_ceo`, `fail_review`) from B to C unseen. A shared `actor_context_fields()` helper computes `(actor_slug, agent_team)` once per verb — zero behavior change — and the new admin-complete open-PR check is likewise extracted to a helper so the override function stays under the threshold.
- **Declared dependencies become real edges (MegaTask + delegate).** The live S6 out-of-order break, both halves. Batch intake: each draft's `depends_on` (the CEO's declared "Depends on" list, batch indices) is now wired verbatim into the sequencing DAG — `SequencingService.analyze` unions declared edges with the derived collision rules (self/out-of-range references rejected, cycles caught by the existing toposort); previously only analyzer-derived file-overlap edges were wired and a declared wave could be silently dropped. Delegate: a `code` subtask now REQUIRES a non-empty `intends_to_touch` collision surface (new `TASK_AT_DELEGATE` completeness spec) — a no-surface code sibling is "parallel to everything" by analyzer design, which is how two devs ran explicitly-sequenced work out of order on divergent branches. Non-code delegations and REST/manual creation are unchanged. The MCP `delegate` tool now actually carries `intends_to_touch` / `adds_migration` / `touches_shared` / `depends_on` and forwards them to the gateway — the gate demanded a field the tool could not send, so every code delegation was rejected `incomplete_input` with no way to comply (live fleet-wide delegation wall); a parity test locks plan-gate fields to tool parameters. - **Declared dependencies become real edges (MegaTask + delegate).** The live S6 out-of-order break, both halves. Batch intake: each draft's `depends_on` (the CEO's declared "Depends on" list, batch indices) is now wired verbatim into the sequencing DAG — `SequencingService.analyze` unions declared edges with the derived collision rules (self/out-of-range references rejected, cycles caught by the existing toposort); previously only analyzer-derived file-overlap edges were wired and a declared wave could be silently dropped. Delegate: a `code` subtask now REQUIRES a non-empty `intends_to_touch` collision surface (new `TASK_AT_DELEGATE` completeness spec) — a no-surface code sibling is "parallel to everything" by analyzer design, which is how two devs ran explicitly-sequenced work out of order on divergent branches. Non-code delegations and REST/manual creation are unchanged. The MCP `delegate` tool now actually carries `intends_to_touch` / `adds_migration` / `touches_shared` / `depends_on` and forwards them to the gateway — the gate demanded a field the tool could not send, so every code delegation was rejected `incomplete_input` with no way to comply (live fleet-wide delegation wall); a parity test locks plan-gate fields to tool parameters.
- **Respawn circuit breaker now guards every task-keyed spawn path.** The progress-aware breaker (strike counting with status-advance reset, tracing-gap budget, DB durability, one-shot CEO notification) was consulted by only 3 dispatch paths; doc/QA/dev/PR-review/PR-gate/revision/board spawns ran unguarded at fixed cadence — a documenter with no valid verb respawned 26× in ~100 min on one task. The gate is now consulted at all 14 task-keyed spawn sites. It also catches status ping-pong: any status change used to fully reset the strike counter, so a `blocked``in_progress` oscillation — which changes status on every spawn while advancing nothing — never tripped the gate (8 spawns over two hours, live). A never-seen status still fully resets; a REVISITED status gets a bounded reset budget (`pm_respawn_max_revisit_resets`, default 2), after which strikes accrue and the gate fires. - **Respawn circuit breaker now guards every task-keyed spawn path.** The progress-aware breaker (strike counting with status-advance reset, tracing-gap budget, DB durability, one-shot CEO notification) was consulted by only 3 dispatch paths; doc/QA/dev/PR-review/PR-gate/revision/board spawns ran unguarded at fixed cadence — a documenter with no valid verb respawned 26× in ~100 min on one task. The gate is now consulted at all 14 task-keyed spawn sites. It also catches status ping-pong: any status change used to fully reset the strike counter, so a `blocked``in_progress` oscillation — which changes status on every spawn while advancing nothing — never tripped the gate (8 spawns over two hours, live). A never-seen status still fully resets; a REVISITED status gets a bounded reset budget (`pm_respawn_max_revisit_resets`, default 2), after which strikes accrue and the gate fires.
- **Assembled-PR freshness + integrity at submit_up / submit_root.** Freshness: the assembled cell/root branch is auto-rebased onto its base when behind (children are terminal at submit time; master is never written); a rebase conflict is a clean rejection naming the files — ends the needs_revision ↔ awaiting_pr_review ping-pong of re-reviewing a stale head. Integrity: every completed child's commits must be patch-present (`git cherry`, rebase-safe) in the assembled branch before review — a completed revert whose merge was lost re-spawned the exact violation it fixed. The guard now also recognizes **squash-merged** children: `git cherry` can't patch-match N child commits against the one squashed commit, but every commit carries the `[taskid8]` prefix, so a parent commit bearing the child's marker proves the child landed (three squash-merged children read as "work missing" and every legitimate `submit_up` was refused, live). Markerless children stay flagged — the original incident the guard exists for. - **Assembled-PR freshness + integrity at submit_up / submit_root.** Freshness: the assembled cell/root branch is auto-rebased onto its base when behind (children are terminal at submit time; master is never written); a rebase conflict is a clean rejection naming the files — ends the needs_revision ↔ awaiting_pr_review ping-pong of re-reviewing a stale head. Integrity: every completed child's commits must be patch-present (`git cherry`, rebase-safe) in the assembled branch before review — a completed revert whose merge was lost re-spawned the exact violation it fixed. The guard now also recognizes **squash-merged** children: `git cherry` can't patch-match N child commits against the one squashed commit, but every commit carries the `[taskid8]` prefix, so a parent commit bearing the child's marker proves the child landed (three squash-merged children read as "work missing" and every legitimate `submit_up` was refused, live). Markerless children stay flagged — the original incident the guard exists for.
@@ -43,6 +49,10 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
- **Admin status override now reconciles claim ownership.** Forcing a `blocked` task into a review/queue state (`needs_revision`, `awaiting_qa`, `awaiting_documentation`, `awaiting_pr_review`, `awaiting_pm_review`) previously left the stale escalation claim in place, so the next claimant was handed the task by `give_me_work`/`triage` while its `note()` writes bounced `not_authorized "you do not hold the claim"` — it re-blocked immediately. The override now clears the claim (`claimed_by`/`claimed_at`/`active_claimant_id`) and consumes the pre-block snapshot for review-state targets; the pending/in_progress owner-restore path additionally syncs `active_claimant_id` so the restored owner's content writes don't bounce either. A REST PATCH that unassigns a task (`assigned_to: null`) now releases the claim with it. - **Admin status override now reconciles claim ownership.** Forcing a `blocked` task into a review/queue state (`needs_revision`, `awaiting_qa`, `awaiting_documentation`, `awaiting_pr_review`, `awaiting_pm_review`) previously left the stale escalation claim in place, so the next claimant was handed the task by `give_me_work`/`triage` while its `note()` writes bounced `not_authorized "you do not hold the claim"` — it re-blocked immediately. The override now clears the claim (`claimed_by`/`claimed_at`/`active_claimant_id`) and consumes the pre-block snapshot for review-state targets; the pending/in_progress owner-restore path additionally syncs `active_claimant_id` so the restored owner's content writes don't bounce either. A REST PATCH that unassigns a task (`assigned_to: null`) now releases the claim with it.
### Removed
- **The never-wired dispatch-time spawn-cooldown path.** `_safe_spawn` / `gateway_pre_spawn_check` / `trigger_filter.decide_spawn` had no caller anywhere in the repo's history — the "enable gateway cooldown logic in production" commit only flipped its flag, and the protections it promised have since shipped better elsewhere: provider parking lives inside `spawn_agent` itself, claim freshness is enforced by the claim guards + reaper, and runaway respawns are bounded by the progress-aware circuit breaker (all 14 task-keyed sites) plus the notification-spawn cooldown. Wiring it now would have re-introduced a per-task cooldown that queue-stalls every normal stage handoff (dev→QA→doc→PM spawn the same task within one window). Deleted: the orchestrator block, `trigger_filter.py`, its tests, and the dead `spawn_cooldown_seconds` / `role_spawn_rate_per_minute` settings. The `gateway_triggers` table is kept (inert; dropping it is a migration decision).
## [0.15.0] - 2026-07-01 ## [0.15.0] - 2026-07-01
### Added ### Added
+25 -3
View File
@@ -2,6 +2,19 @@
PYTHON_VERSIONS = 3.10 3.11 3.12 3.13 3.14 PYTHON_VERSIONS = 3.10 3.11 3.12 3.13 3.14
DEFAULT_PYTHON = 3.10 DEFAULT_PYTHON = 3.10
# Every `uv run` implicitly re-syncs the venv (rebuilding the roboco package
# after any source edit). Two uv processes doing that concurrently — a
# background `make quality` plus any foreground `uv run` — race on one
# .venv and tear site-packages apart (recurring rich/pip/bandit ImportError
# corruption). Recipes therefore never sync implicitly; targets that need a
# fresh env depend on the explicit `sync` below, which runs once, up front.
export UV_NO_SYNC := 1
.PHONY: sync
sync:
@echo "==> uv sync --extra dev"
@uv sync --extra dev
# Install dependencies # Install dependencies
.PHONY: install .PHONY: install
install: install:
@@ -245,7 +258,7 @@ security: bandit pip-audit
# Run every quality gate. Fails on any red. Use this as the merge gate. # Run every quality gate. Fails on any red. Use this as the merge gate.
.PHONY: quality .PHONY: quality
quality: quality: sync
@echo "==> ruff format --check" @echo "==> ruff format --check"
@uv run ruff format --check . @uv run ruff format --check .
@echo "==> ruff check" @echo "==> ruff check"
@@ -282,8 +295,17 @@ quality:
@echo "" @echo ""
@echo "All quality gates passed." @echo "All quality gates passed."
# Scripted-agent lifecycle smoke: the REAL MCP flow/do tools driven through
# the REAL gateway/gates against an in-process API + ephemeral test Postgres +
# local git origin (gh shimmed). No LLM. Excluded from `quality` (env-gated);
# CI runs it as its own job.
.PHONY: e2e-smoke
e2e-smoke: sync
@echo "==> e2e lifecycle smoke (scripted agents, real gates)"
@ROBOCO_E2E_SMOKE=1 uv run pytest tests/e2e_smoke -q --no-cov
.PHONY: quality-fast .PHONY: quality-fast
quality-fast: quality-fast: sync
@uv run ruff format --check . @uv run ruff format --check .
@uv run ruff check . @uv run ruff check .
@uv run mypy roboco/ tests/ @uv run mypy roboco/ tests/
@@ -294,7 +316,7 @@ quality-fast:
# pre-submit gate (run at i_am_done) executes it in the dev's workspace and # pre-submit gate (run at i_am_done) executes it in the dev's workspace and
# catches lint/type/complexity at the desk. The test suite stays on CI. # catches lint/type/complexity at the desk. The test suite stays on CI.
.PHONY: gate .PHONY: gate
gate: gate: sync
@uv run ruff format --check . @uv run ruff format --check .
@uv run ruff check . @uv run ruff check .
@uv run mypy roboco/ tests/ @uv run mypy roboco/ tests/
+28
View File
@@ -550,3 +550,31 @@ Nine live-run fixes, all merged to `master` the same day (commits `81f448bb`, `0
9. **Team-match enforcement armed** (`fe9e5589`, `8e5f84c4`) — `roboco/foundation/policy/lifecycle.py`: `resume`/`unblock`/`activate` flip `needs_team_match=True`; new `_ORG_WIDE_ROLES` exemption (main_pm, CEO, PO, head_marketing, auditor, pr_reviewer) in `_check_team_match(…, role)`. `8e5f84c4` threads `agent_team` through all 27 gateway `Context` construction sites (`choreographer/_impl.py`, `doc.py`, `pr_gate.py`, `pr_review.py`, `qa.py`) so the gate actually receives the team — it had sat in its permissive fallback since shipping. 9. **Team-match enforcement armed** (`fe9e5589`, `8e5f84c4`) — `roboco/foundation/policy/lifecycle.py`: `resume`/`unblock`/`activate` flip `needs_team_match=True`; new `_ORG_WIDE_ROLES` exemption (main_pm, CEO, PO, head_marketing, auditor, pr_reviewer) in `_check_team_match(…, role)`. `8e5f84c4` threads `agent_team` through all 27 gateway `Context` construction sites (`choreographer/_impl.py`, `doc.py`, `pr_gate.py`, `pr_review.py`, `qa.py`) so the gate actually receives the team — it had sat in its permissive fallback since shipping.
Slices touched: worksession-git (2, 3), choreographer/gateway-support (9), orchestrator (6, 7, 8), foundation-lifecycle (7, 9), mcp-servers (1), api-routes-schemas (4), panel (5). `docs/map/_complete_map.md` is the pre-delta concatenation — not regenerated. Slices touched: worksession-git (2, 3), choreographer/gateway-support (9), orchestrator (6, 7, 8), foundation-lifecycle (7, 9), mcp-servers (1), api-routes-schemas (4), panel (5). `docs/map/_complete_map.md` is the pre-delta concatenation — not regenerated.
---
## Delta 2026-07-02 (evening) — leak-fix batch (branch `fix/leak-fix-batch`)
The remaining live-run leak fixes from the S6/fb836f80 postmortems, TDD'd per item:
1. **Spawner attribution**`roboco/runtime/orchestrator.py` `spawn_agent(spawned_by=)``_launch_spawn(spawned_by=)` stamps the dispatching loop's name into `agent.spawned` / `agent.spawn_failed` audit details (`"unspecified"` when absent). All ~29 call sites pass their loop name (orchestrator dispatchers by method name; `bootstrap`, `api.orchestrator.spawn`, `event.auditor_spawn`); `OrchestratorAccessProtocol.spawn_agent` (roboco/models/events.py) gains the kwarg. A whole-package AST sweep test (`tests/unit/runtime/test_spawn_attribution.py`) fails any future caller omitting it. `pyproject.toml` adds PLR0913 to the orchestrator's per-file ignores (spawn contract > 5 params, bundle refused). NOTE found in passing: `_safe_spawn` + `gateway_pre_spawn_check` (orchestrator ~L741/L1932, the trigger_filter dispatch-time cooldown) have NEVER had a caller — flagged to CEO, untouched under the freeze.
2. **Admin-complete merge-or-refuse**`roboco/api/routes/tasks.py` `_apply_forced_status_override`: `status=completed` without `force` on a task whose work session records `pr_status == "open"` now refuses naming PR number/URL + the stranding consequence (checked before the generic hatch text). New `TaskService.open_pr_ref` (`roboco/services/task.py`, beside `_assert_pr_merged_for_complete`) is the lookup.
3. **Panel CEO-approve**`panel/src/components/tasks/task-detail/task-header.tsx` `AWAITING_CEO_APPROVAL`: primary action now `ceo-approve` (CeoApproveDialog → `POST /ceo-approve`, notes ≥ 20); `approve-and-merge` only when `task.pr_number` (umbrellas 400'd `NO_PR`).
4. **Live-Redis test leakage killed**`tests/unit/services/test_self_heal_originate_db.py` wrote `self_heal:notified:*` (2h TTL) and `tests/unit/gateway/test_i_am_blocked_rate_limited.py` wrote a NO-TTL `roboco:rate_limit:anthropic:state` "rate_limited" blob into live localhost Redis on every run (verified live, keys scrubbed). Both now patch `cfg.redis_host/redis_port` (redis_url is computed) to an unreachable port; engines' fail-open paths keep assertions intact. The one-off `test_self_heal_engine` full-run failure was NOT reproduced (adversarial orders, 5× dir loops, green full gate) — this leakage class + the 2026-07-02 corrupted-venv day remain the suspects.
5. **Phase1 smoke mock team**`tests/integration/test_foundation_phase1_smoke.py` parent mock gains `team="backend"`; the armed team-match gate (delta above) rejected the auto-generated MagicMock team before the asserted `incomplete_input` (the sweep fixed 13 files; this integration file was outside the gateway/foundation/runtime subsets it ran).
6. **Gate green again: ten xenon C-ranks + global test-Redis isolation** — master CI red at the smoke test meant no gate (CI or local) had reached xenon since the team-match sweep; its inline `agent_team=…` kwarg had pushed 9 verb bodies to C(11-12) unseen (`_impl.py` i_am_done/resume/unclaim/submit_up/submit_root/complete/escalate_up/escalate_to_ceo, `qa.py` fail_review). New shared `actor_context_fields()` (`choreographer/_protocol.py`) computes `(actor_slug, agent_team)` once per verb; the admin-complete open-PR check is extracted to `_refuse_unforced_complete_with_open_pr` (routes/tasks.py). `tests/conftest.py` gains the autouse `_no_live_redis` fixture (root fix for the live-Redis leak class; notif_dedup purpose-dedupe keys were a third, self-expiring family).
7. **Dead spawn-cooldown path deleted (CEO-ratified)**`_safe_spawn` + `gateway_pre_spawn_check` (orchestrator) + `roboco/services/gateway/trigger_filter.py` + both test files + the dead `spawn_cooldown_seconds`/`role_spawn_rate_per_minute` settings. Never called in repo history; superseded by provider parking (in `spawn_agent`), claim guards/reaper, the respawn circuit breaker, and the notification-spawn cooldown. Rule 4 (per-task cooldown) would have queue-stalled every stage handoff if wired. `GatewayTriggerTable` kept inert (drop = future migration decision).
Slices touched: orchestrator (1, 7), api-routes-schemas + taskservice (2), panel (3), choreographer/gateway-support (6, 7), tests (4, 5, 6, 7). `docs/map/_complete_map.md` still not regenerated.
---
## Delta 2026-07-02 (late) — CEO-ratified follow-ups (branch `fix/leak-fix-batch`)
1. **Dead spawn-cooldown path deleted** (see item 7 above; commit `78fa0f6f`).
2. **uv serialization in the Makefile** (`cf5043fe`) — `export UV_NO_SYNC := 1` + `sync` prerequisite on quality/quality-fast/gate: gate recipes never implicitly re-sync the venv (the recurring rich/pip/bandit corruption came from two uv writers racing one `.venv`).
3. **`_api_base()` in git.py** (`13b9c5d4`) — 15 hardcoded `https://api.github.com` PR/merge/branch sites now honor `settings.github_api_base_url` like the CI/open-PR sites already did (GHE/test override fix; enables the smoke harness's fake GitHub).
4. **CI split** (`db6b3088`) — ci.yml (backend `quality` only; FILE name kept — self-heal/ci-watch/release default to it), `panel-ci.yml` (panel job, `panel/**` paths), `e2e-smoke.yml` (new job). Panel-only reds now land on Panel CI, unwatched by the ci.yml-pinned engines.
5. **e2e smoke harness**`tests/e2e_smoke/{conftest,harness,test_dev_lifecycle}.py` + `make e2e-smoke` (env-gated `ROBOCO_E2E_SMOKE=1`, skipped in the default suite). Harness: real routers/middleware on uvicorn over the ephemeral test DB (`settings.database_*` patched + `_DbHolder` reset), bare origin at `<tmp>/github.com/e2e-smoke/proj.git` (satisfies `_parse_git_url`, clones tokenless), fake GitHub REST router doing real squash-merges via an admin clone, `ScriptedAgent` reloading the real MCP modules per role. Scenario 1 (leaf dev arc → awaiting_pm_review) GREEN in ~5s. Learned seams scripted: post-claim tracing gap keeps the claim; note scopes are decision/learning/note/reflect/struggle (no 'progress'); i_am_done demands during-work+handoff+reflect+per-AC artifacts; pass_review demands learning note + ac_verdicts; A2A resolves roles from the STATIC agents_config registry (seed canonical slugs: be-dev-1/be-qa/be-doc/be-pm/main-pm).
Slices touched: worksession-git (3), orchestrator (1), deployment-tooling (2, 4), tests (5).
@@ -0,0 +1,105 @@
import { describe, it, expect, vi } from "vitest";
import { render } from "@testing-library/react";
import React from "react";
import { TaskStatus, Team, TaskType, type Task } from "@/types";
// awaiting_ceo_approval must offer the PROVEN approval path: the
// CeoApproveDialog -> POST /tasks/{id}/ceo-approve (notes >= 20 chars).
// The old wiring offered only approve-and-merge, which 400s NO_PR on a
// branchless MegaTask umbrella — the CEO's approve button just failed.
vi.mock("next/navigation", () => ({
useRouter: () => ({ push: vi.fn(), back: vi.fn() }),
}));
vi.mock("@/hooks/use-tasks", () => ({
useUpdateTask: () => ({ mutateAsync: vi.fn(), isPending: false }),
useDeleteTask: () => ({ mutateAsync: vi.fn(), isPending: false }),
useTaskValidTransitions: () => ({ data: [], isLoading: false }),
}));
vi.mock("sonner", () => ({
toast: { success: vi.fn(), error: vi.fn() },
}));
vi.mock("@/components/ui/select", () => ({
Select: ({ children }: { children: React.ReactNode }) => <div>{children}</div>,
SelectTrigger: ({ children }: { children: React.ReactNode }) => (
<div>{children}</div>
),
SelectValue: () => null,
SelectContent: ({ children }: { children: React.ReactNode }) => (
<div>{children}</div>
),
SelectItem: ({ children }: { children: React.ReactNode }) => (
<div>{children}</div>
),
}));
// Render the dropdown inline so the action items are clickable without
// Radix's portal/pointer machinery.
vi.mock("@/components/ui/dropdown-menu", () => ({
DropdownMenu: ({ children }: { children: React.ReactNode }) => (
<div>{children}</div>
),
DropdownMenuTrigger: ({ children }: { children: React.ReactNode }) => (
<div>{children}</div>
),
DropdownMenuContent: ({ children }: { children: React.ReactNode }) => (
<div>{children}</div>
),
DropdownMenuItem: ({
children,
onClick,
}: {
children: React.ReactNode;
onClick?: () => void;
}) => (
<button type="button" onClick={onClick}>
{children}
</button>
),
DropdownMenuSeparator: () => null,
}));
import { TaskHeader } from "../task-header";
function buildTask(overrides: Partial<Task> = {}): Task {
return {
id: "t1",
title: "Umbrella awaiting CEO",
description: "d",
status: TaskStatus.AWAITING_CEO_APPROVAL,
team: Team.BACKEND,
task_type: TaskType.CODE,
acceptance_criteria: [],
...overrides,
} as unknown as Task;
}
describe("TaskHeader awaiting_ceo_approval actions", () => {
it("offers Approve & Complete wired to the ceo-approve dialog", () => {
const onAction = vi.fn();
const { getByText } = render(
<TaskHeader task={buildTask()} onAction={onAction} />,
);
getByText("Approve & Complete").click();
expect(onAction).toHaveBeenCalledWith("ceo-approve");
});
it("hides Approve & Merge when the task has no PR (umbrella)", () => {
const { queryByText } = render(
<TaskHeader task={buildTask()} onAction={vi.fn()} />,
);
expect(queryByText("Approve & Merge")).toBeNull();
});
it("still offers Approve & Merge for a PR-bearing task", () => {
const onAction = vi.fn();
const { getByText } = render(
<TaskHeader task={buildTask({ pr_number: 42 })} onAction={onAction} />,
);
getByText("Approve & Merge").click();
expect(onAction).toHaveBeenCalledWith("approve-and-merge");
});
});
@@ -389,11 +389,23 @@ export function TaskHeader({ task, onAction }: TaskHeaderProps) {
}); });
break; break;
case TaskStatus.AWAITING_CEO_APPROVAL: case TaskStatus.AWAITING_CEO_APPROVAL:
// The proven approval path: CeoApproveDialog -> POST /ceo-approve
// (notes >= 20 chars). Works for every gated task, including the
// branchless MegaTask umbrella, which has no PR to merge.
actions.push({ actions.push({
label: "Approve & Merge", label: "Approve & Complete",
action: "approve-and-merge", action: "ceo-approve",
icon: <ThumbsUp className="h-4 w-4 mr-2" />, icon: <ThumbsUp className="h-4 w-4 mr-2" />,
}); });
// One-click merge+complete only makes sense when a PR exists;
// on an umbrella it 400s NO_PR — the CEO's approve just failed.
if (task.pr_number) {
actions.push({
label: "Approve & Merge",
action: "approve-and-merge",
icon: <GitMerge className="h-4 w-4 mr-2" />,
});
}
actions.push({ actions.push({
label: "Request Changes", label: "Request Changes",
action: "ceo-reject", action: "ceo-reject",
+9 -1
View File
@@ -165,10 +165,18 @@ select = [
# to avoid import cycles with the modules it wires (same rationale as above). # to avoid import cycles with the modules it wires (same rationale as above).
"roboco/api/deps.py" = ["PLC0415"] "roboco/api/deps.py" = ["PLC0415"]
"roboco/runtime/*.py" = ["PLC0415"] "roboco/runtime/*.py" = ["PLC0415"]
# The e2e smoke harness defers every roboco import until the stack fixture
# runs, so the default (skipped) suite never pays the app-surface import
# cost; ARG001 covers FastAPI path params the fake-GitHub handlers must
# name but not read.
"tests/e2e_smoke/*.py" = ["PLC0415", "ARG001"]
# PTH119: _grok_usage_json sanitizes the agent id with os.path.basename — the # PTH119: _grok_usage_json sanitizes the agent id with os.path.basename — the
# path-injection sanitizer CodeQL's query models; the pathlib equivalent # path-injection sanitizer CodeQL's query models; the pathlib equivalent
# (Path(...).name) is not recognized by that query, so we keep os.path here. # (Path(...).name) is not recognized by that query, so we keep os.path here.
"roboco/runtime/orchestrator.py" = ["PTH119"] # PLR0913: spawn_agent / _launch_spawn carry the spawn contract (task, model,
# git context, spawner attribution) — a bundle dataclass would just relocate
# the same six fields behind one hop at the fleet's hottest call surface.
"roboco/runtime/orchestrator.py" = ["PTH119", "PLR0913"]
# The intake driver/entrypoint lazily import the heavy `claude-agent-sdk` (and # The intake driver/entrypoint lazily import the heavy `claude-agent-sdk` (and
# uvicorn) so the modules import without those installed and don't pay the cost # uvicorn) so the modules import without those installed and don't pay the cost
# until a live container runs them — same rationale as the dirs above. # until a live container runs them — same rationale as the dirs above.
+1
View File
@@ -213,6 +213,7 @@ async def spawn_agent(
initial_prompt=data.initial_prompt if data else None, initial_prompt=data.initial_prompt if data else None,
task_id=data.task_id if data else None, task_id=data.task_id if data else None,
model=data.model if data else None, model=data.model if data else None,
spawned_by="api.orchestrator.spawn",
) )
except FileNotFoundError as e: except FileNotFoundError as e:
raise HTTPException( raise HTTPException(
+25
View File
@@ -126,6 +126,30 @@ class _StatusOverride:
agent: AgentContext agent: AgentContext
async def _refuse_unforced_complete_with_open_pr(req: _StatusOverride) -> None:
"""Admin-complete must merge-or-refuse.
Completing a task whose PR is still OPEN strands its commits unmerged
(bit the CEO twice live, 2026-07-02). Checked before the generic hatch
text so the refusal names the PR and the consequence instead of a vague
gate message; ``force`` stays the deliberate, audited escape.
"""
if req.new_status != TaskStatus.COMPLETED or req.force:
return
open_ws = await req.service.open_pr_ref(req.task)
if open_ws is not None:
raise HTTPException(
status_code=status.HTTP_400_BAD_REQUEST,
detail=(
f"Task still has OPEN PR #{open_ws.pr_number}"
f" ({open_ws.pr_url}); completing it now would strand"
" those commits unmerged. Merge the PR first (or approve"
" via POST /api/tasks/{id}/ceo-approve), or pass"
' "force": true to strand it deliberately.'
),
)
async def _apply_forced_status_override(req: _StatusOverride) -> TaskTable: async def _apply_forced_status_override(req: _StatusOverride) -> TaskTable:
"""Apply an audited admin status override, gating the lifecycle bypass. """Apply an audited admin status override, gating the lifecycle bypass.
@@ -141,6 +165,7 @@ async def _apply_forced_status_override(req: _StatusOverride) -> TaskTable:
status_code=status.HTTP_403_FORBIDDEN, status_code=status.HTTP_403_FORBIDDEN,
detail="Only privileged roles may override task status.", detail="Only privileged roles may override task status.",
) )
await _refuse_unforced_complete_with_open_pr(req)
if req.new_status in _HATCH_OVERRIDE_STATES and not req.force: if req.new_status in _HATCH_OVERRIDE_STATES and not req.force:
raise HTTPException( raise HTTPException(
status_code=status.HTTP_400_BAD_REQUEST, status_code=status.HTTP_400_BAD_REQUEST,
+1
View File
@@ -138,6 +138,7 @@ async def main(
await orchestrator.spawn_agent( await orchestrator.spawn_agent(
agent_id=agent_id, agent_id=agent_id,
initial_prompt=startup_prompt, initial_prompt=startup_prompt,
spawned_by="bootstrap",
) )
except Exception as e: except Exception as e:
logger.error("Failed to spawn agent", agent_id=agent_id, error=str(e)) logger.error("Failed to spawn agent", agent_id=agent_id, error=str(e))
+7 -21
View File
@@ -950,14 +950,10 @@ class Settings(BaseSettings):
) )
# Gateway coordination thresholds # Gateway coordination thresholds
# Single source of truth for "claim heartbeat is stale": consumed both by # Single source of truth for "claim heartbeat is stale", consumed via
# `trigger_filter` (deciding whether to QUEUE a fresh spawn) and by # `claimant_lock.is_stale` wherever a claim's freshness gates an action
# `_reap_stale_claims` (deciding whether to RELEASE the claim back to # (e.g. `_reap_stale_claims` deciding whether to RELEASE the claim back
# pending). Keeping them on one field guarantees both layers agree on # to pending). One field keeps every consumer on the same tick.
# the same tick — the reaper runs first, releases the row, and the
# queued spawn finds an unclaimed task. Splitting them into two fields
# opens a window where trigger_filter queues duplicate spawns against a
# claim the reaper hasn't yet released — pure dispatcher churn.
claim_stale_seconds: int = Field( claim_stale_seconds: int = Field(
default=180, default=180,
ge=60, ge=60,
@@ -986,9 +982,9 @@ class Settings(BaseSettings):
# retrying — LLM inference + retry loops routinely exceed 3 min # retrying — LLM inference + retry loops routinely exceed 3 min
# between verb successes. 600s is large enough to accommodate that # between verb successes. 600s is large enough to accommodate that
# without letting a genuinely-stuck container linger. # without letting a genuinely-stuck container linger.
# Distinct from claim_stale_seconds (which drives trigger_filter # Distinct from claim_stale_seconds (the general claim-freshness
# spawn queueing); keeping them separate avoids a window where a # threshold); keeping them separate lets the reaper run on a longer
# higher reap threshold would also delay spawn-queue decisions. # window than other claim-staleness consumers.
stale_claim_reap_seconds: int = Field( stale_claim_reap_seconds: int = Field(
default=600, default=600,
ge=60, ge=60,
@@ -1086,16 +1082,6 @@ class Settings(BaseSettings):
"gating verbs; override via ROBOCO_PM_DECISION_WINDOW_SECONDS" "gating verbs; override via ROBOCO_PM_DECISION_WINDOW_SECONDS"
), ),
) )
spawn_cooldown_seconds: int = Field(
default=60,
ge=1,
description="Per-task spawn rate cooldown (seconds)",
)
role_spawn_rate_per_minute: int = Field(
default=6,
ge=1,
description="Per-role spawn rate limit (per minute)",
)
# Tracing-gate thresholds # Tracing-gate thresholds
qa_notes_min_chars: int = Field( qa_notes_min_chars: int = Field(
+3 -1
View File
@@ -352,7 +352,9 @@ async def handle_auditor_spawn(event: Event) -> None:
) )
try: try:
await _context.orchestrator.spawn_agent(agent_id="auditor") await _context.orchestrator.spawn_agent(
agent_id="auditor", spawned_by="event.auditor_spawn"
)
except Exception as exc: except Exception as exc:
# Auditor spawn failure must NOT block the underlying event from # Auditor spawn failure must NOT block the underlying event from
# being processed. Log the warning and return cleanly. # being processed. Log the warning and return cleanly.
+2
View File
@@ -182,6 +182,8 @@ class OrchestratorAccessProtocol(Protocol):
self, self,
agent_id: str, agent_id: str,
initial_prompt: str | None = None, initial_prompt: str | None = None,
*,
spawned_by: str | None = None,
) -> Any: ... ) -> Any: ...
+43 -245
View File
@@ -669,184 +669,6 @@ def _build_manifest_for_agent(
return write_path return write_path
# =============================================================================
# GATEWAY PRE-SPAWN CHECK (trigger_filter spawn cooldown)
# =============================================================================
async def _count_recent_spawns_for_task(
db_session: Any,
task_id: Any,
cutoff: datetime,
) -> int:
"""Count recent SPAWN decisions for ``task_id`` since ``cutoff``."""
from sqlalchemy import select
from roboco.db.tables import GatewayTriggerTable
result = await db_session.execute(
select(GatewayTriggerTable).where(
GatewayTriggerTable.task_id == task_id,
GatewayTriggerTable.created_at >= cutoff,
GatewayTriggerTable.decision == "spawn",
)
)
return len(result.scalars().all())
async def _count_recent_spawns_for_role(
db_session: Any,
target_role: str,
cutoff: datetime,
) -> int:
"""Count recent SPAWN decisions for ``target_role`` since ``cutoff``."""
from sqlalchemy import select
from roboco.db.tables import GatewayTriggerTable
result = await db_session.execute(
select(GatewayTriggerTable).where(
GatewayTriggerTable.target_role == target_role,
GatewayTriggerTable.created_at >= cutoff,
GatewayTriggerTable.decision == "spawn",
)
)
return len(result.scalars().all())
async def _record_trigger_decision(
db_session: Any,
task_id: Any,
trigger_kind: str,
target_role: str,
decision: Any,
) -> None:
"""Persist a gateway trigger decision row."""
from uuid import uuid4 as _uuid4
from roboco.db.tables import GatewayTriggerTable
row = GatewayTriggerTable(
id=_uuid4(),
trigger_kind=trigger_kind,
task_id=task_id,
target_role=target_role,
decision=decision.outcome.value,
decision_reason=decision.reason,
)
db_session.add(row)
await db_session.flush()
async def gateway_pre_spawn_check(
*,
task_id: str | None,
trigger_kind: str,
target_role: str,
provider: str | None = None,
) -> tuple[str, str]:
"""Consult trigger_filter before spawning a container.
Returns a ``(outcome, reason)`` tuple where ``outcome`` is one of
``"spawn"``, ``"queue"``, or ``"drop"``.
The trigger_filter spawn cooldown runs unconditionally for every spawn.
Args:
provider: Optional provider name (e.g. ``"anthropic"``) for the
agent about to be spawned. When given, the
``RateLimitStateTracker`` is consulted and a QUEUE decision is
returned when that provider is currently rate-limited.
"""
from roboco.db.base import get_session_factory
from roboco.services.gateway.trigger_filter import (
Decision,
SpawnConfig,
SpawnDecision,
TriggerContext,
TriggerKind,
decide_spawn,
)
cutoff = datetime.now(tz=UTC) - timedelta(seconds=settings.spawn_cooldown_seconds)
role_cutoff = datetime.now(tz=UTC) - timedelta(seconds=60)
# When no task_id we cannot query counts; allow (no-task spawns like idle PMs).
if task_id is None:
return SpawnDecision.SPAWN, "no task_id — no-task spawn, skip gate"
try:
from sqlalchemy import select as _select
from roboco.db.tables import TaskTable as _TaskTable
factory = get_session_factory()
async with factory() as db:
recent_for_task = await _count_recent_spawns_for_task(db, task_id, cutoff)
recent_for_role = await _count_recent_spawns_for_role(
db, target_role, role_cutoff
)
# Load the lightweight task proxy needed by is_stale / decide_spawn.
task_result = await db.execute(
_select(_TaskTable).where(_TaskTable.id == task_id)
)
task_row = task_result.scalars().first()
if task_row is None:
return SpawnDecision.SPAWN, "task not found in DB — allow by default"
# Check provider rate-limit status when a provider is known.
# Failure is non-fatal — degrade to False (allow spawn) so Redis
# unavailability never permanently blocks the dispatcher.
provider_rate_limited = False
if provider is not None:
try:
from roboco.services.gateway.rate_limit_tracker import (
RateLimitStateTracker,
)
provider_rate_limited = await RateLimitStateTracker(
provider
).is_rate_limited()
except Exception:
provider_rate_limited = False
trigger = TriggerContext(
kind=TriggerKind(trigger_kind),
skill=None,
recent_spawns_for_task=recent_for_task,
recent_spawns_for_role=recent_for_role,
provider=provider,
provider_rate_limited=provider_rate_limited,
)
config = SpawnConfig(
cooldown_seconds=settings.spawn_cooldown_seconds,
role_rate_per_minute=settings.role_spawn_rate_per_minute,
claim_stale_seconds=settings.claim_stale_seconds,
)
decision: Decision = decide_spawn(
task=task_row, trigger=trigger, config=config
)
await _record_trigger_decision(
db, task_id, trigger_kind, target_role, decision
)
await db.commit()
return decision.outcome.value, decision.reason
except Exception as exc:
# Gateway errors must never block a spawn — degrade gracefully.
logger.warning(
"Gateway pre-spawn check failed; defaulting to spawn",
task_id=task_id,
trigger_kind=trigger_kind,
error=str(exc),
)
return "spawn", f"gateway error (degraded): {exc}"
class AgentReadinessError(Exception): class AgentReadinessError(Exception):
"""Raised when spawn_agent refuses to spawn because the task isn't ready. """Raised when spawn_agent refuses to spawn because the task isn't ready.
@@ -1006,9 +828,8 @@ class AgentOrchestrator:
self._board_review_ceo_notified: set[str] = set() self._board_review_ceo_notified: set[str] = set()
# Stale-claim reaper config, sourced from # Stale-claim reaper config, sourced from
# stale_claim_reap_seconds (default 600) rather than # stale_claim_reap_seconds (default 600) rather than
# claim_stale_seconds (default 180). The two settings are now # claim_stale_seconds (default 180) — the reaper gets the longer
# distinct: claim_stale_seconds drives trigger_filter (spawn # window of the two claim-staleness thresholds.
# queueing); stale_claim_reap_seconds drives the reaper.
# Smoke run 3 showed agents reaped at 180s while actively retrying # Smoke run 3 showed agents reaped at 180s while actively retrying
# rejected verbs — LLM inference routinely exceeds that window. # rejected verbs — LLM inference routinely exceeds that window.
# Tests bypass `__init__` via `__new__` and set _claim_heartbeat_ttl # Tests bypass `__init__` via `__new__` and set _claim_heartbeat_ttl
@@ -1929,69 +1750,6 @@ class AgentOrchestrator:
) )
return None return None
async def _safe_spawn(
self,
*,
agent_id: str,
task_id: str | None = None,
initial_prompt: str | None = None,
git_context: SpawnGitContext | None = None,
context_label: str = "dispatcher",
) -> AgentInstance | None:
"""Spawn an agent, absorbing errors so one bad spawn doesn't abort the
rest of the dispatcher's loop.
Each dispatcher iterates many tasks; if `spawn_agent` raised, the
remaining tasks were skipped until the next tick. This wrapper logs
and returns None on failure so siblings still get dispatched.
The gateway pre-spawn check runs first; a QUEUE or DROP outcome skips
the container launch.
"""
# Gateway pre-spawn cooldown gate.
target_role = get_agent_role(agent_id) or "unknown"
# Map context_label to one of the TriggerKind string values; unknown
# labels fall back to "scan" which is the least-specific kind.
trigger_kind_map = {
"a2a": "a2a",
"escalation": "escalation",
"notification": "notification",
}
trigger_kind = trigger_kind_map.get(context_label, "scan")
outcome, reason = await gateway_pre_spawn_check(
task_id=task_id,
trigger_kind=trigger_kind,
target_role=target_role,
provider=self.get_provider_for_agent(agent_id),
)
if outcome != "spawn":
logger.info(
"Gateway pre-spawn check suppressed spawn",
agent_id=agent_id,
task_id=task_id,
outcome=outcome,
reason=reason,
)
return None
try:
return await self.spawn_agent(
agent_id=agent_id,
task_id=task_id,
initial_prompt=initial_prompt,
git_context=git_context,
)
except Exception as e:
logger.error(
"Spawn failed during dispatch; continuing with next task",
context=context_label,
agent_id=agent_id,
task_id=task_id,
error=str(e),
)
return None
async def _resolve_spawn_git_context( async def _resolve_spawn_git_context(
self, self,
git_context: SpawnGitContext | None, git_context: SpawnGitContext | None,
@@ -2211,11 +1969,17 @@ class AgentOrchestrator:
instance: AgentInstance, instance: AgentInstance,
initial_prompt: str | None, initial_prompt: str | None,
agent_settings_path: Path | None, agent_settings_path: Path | None,
*,
spawned_by: str | None = None,
) -> AgentInstance: ) -> AgentInstance:
"""Launch the container and emit spawn audit events. """Launch the container and emit spawn audit events.
`agent_id` was dropped as a redundant parameter `config.agent_id` `agent_id` was dropped as a redundant parameter `config.agent_id`
is the same value and was always the caller's source. is the same value and was always the caller's source.
``spawned_by`` names the dispatch loop that requested the spawn; it is
stamped into the spawned/spawn_failed audit details so a rogue spawner
is identifiable from the audit log alone.
""" """
agent_slug = config.agent_id agent_slug = config.agent_id
try: try:
@@ -2242,6 +2006,7 @@ class AgentOrchestrator:
details={ details={
"container_id": container_id[:12], "container_id": container_id[:12],
"model": config.model, "model": config.model,
"spawned_by": spawned_by or "unspecified",
}, },
) )
@@ -2264,7 +2029,10 @@ class AgentOrchestrator:
event_type="agent.spawn_failed", event_type="agent.spawn_failed",
agent_slug=agent_slug, agent_slug=agent_slug,
task_id=task_id, task_id=task_id,
details={"error": str(e)}, details={
"error": str(e),
"spawned_by": spawned_by or "unspecified",
},
severity="error", severity="error",
) )
raise raise
@@ -2339,6 +2107,8 @@ class AgentOrchestrator:
task_id: str | None = None, task_id: str | None = None,
model: str | None = None, model: str | None = None,
git_context: SpawnGitContext | None = None, git_context: SpawnGitContext | None = None,
*,
spawned_by: str | None = None,
) -> AgentInstance: ) -> AgentInstance:
""" """
Spawn a Claude Code container for an agent. Spawn a Claude Code container for an agent.
@@ -2349,6 +2119,8 @@ class AgentOrchestrator:
task_id: Optional task ID being worked on task_id: Optional task ID being worked on
model: Override model selection model: Override model selection
git_context: Optional git context (project_slug, branch_name) git_context: Optional git context (project_slug, branch_name)
spawned_by: Name of the dispatch loop / entry point requesting
the spawn stamped into the agent.spawned audit details
Returns: Returns:
AgentInstance handle AgentInstance handle
@@ -2456,6 +2228,7 @@ class AgentOrchestrator:
instance, instance,
initial_prompt, initial_prompt,
agent_settings_path, agent_settings_path,
spawned_by=spawned_by,
) )
def _resolve_host_paths( def _resolve_host_paths(
@@ -3798,6 +3571,7 @@ class AgentOrchestrator:
task_id=task["id"], task_id=task["id"],
initial_prompt=self._build_dev_prompt(task), initial_prompt=self._build_dev_prompt(task),
git_context=self._task_git_context(task), git_context=self._task_git_context(task),
spawned_by="_respawn_dev_for_pr_half",
) )
# ========================================================================= # =========================================================================
@@ -6319,6 +6093,7 @@ class AgentOrchestrator:
initial_prompt=resume_prompt, initial_prompt=resume_prompt,
task_id=record.task_id, task_id=record.task_id,
git_context=prior_git_context, git_context=prior_git_context,
spawned_by="resolve_wait",
) )
except Exception: except Exception:
# Spawn failed (e.g. readiness refused → task auto-blocked). Tear # Spawn failed (e.g. readiness refused → task auto-blocked). Tear
@@ -6946,6 +6721,7 @@ Start by:
agent_id=agent_id, agent_id=agent_id,
task_id=instance.current_task_id, task_id=instance.current_task_id,
git_context=(instance.config.git_context if instance.config else None), git_context=(instance.config.git_context if instance.config else None),
spawned_by="_crash_retry_or_escalate",
) )
elif instance.error_count == max_retries: elif instance.error_count == max_retries:
# Exactly at the threshold — escalate once to humans so a # Exactly at the threshold — escalate once to humans so a
@@ -10069,6 +9845,7 @@ Start now: evidence(task_id="{task_id}")
task_id=task["id"], task_id=task["id"],
initial_prompt=pm_prompt, initial_prompt=pm_prompt,
git_context=self._task_git_context(task), git_context=self._task_git_context(task),
spawned_by="_handle_pm_assigned_task",
) )
async def _handle_board_assigned_task( async def _handle_board_assigned_task(
@@ -10131,6 +9908,7 @@ Start now: evidence(task_id="{task_id}")
task_id=task["id"], task_id=task["id"],
initial_prompt=self._build_board_prompt(task), initial_prompt=self._build_board_prompt(task),
git_context=self._task_git_context(task), git_context=self._task_git_context(task),
spawned_by="_dispatch_board_reviewer",
) )
def _board_review_complete(self, task_id: str) -> bool: def _board_review_complete(self, task_id: str) -> bool:
@@ -10312,6 +10090,7 @@ Start now: evidence(task_id="{task_id}")
task_id=task["id"], task_id=task["id"],
initial_prompt=prompt, initial_prompt=prompt,
git_context=self._task_git_context(task), git_context=self._task_git_context(task),
spawned_by="_route_unassigned_pm_task",
) )
async def _dispatch_pm_work(self, client: httpx.AsyncClient) -> None: async def _dispatch_pm_work(self, client: httpx.AsyncClient) -> None:
@@ -10391,6 +10170,7 @@ Start now: evidence(task_id="{task_id}")
task_id=task["id"], task_id=task["id"],
initial_prompt=self._get_prompt_for_agent(agent_slug, task), initial_prompt=self._get_prompt_for_agent(agent_slug, task),
git_context=self._task_git_context(task), git_context=self._task_git_context(task),
spawned_by="_dispatch_revision_coordination_roots",
) )
@staticmethod @staticmethod
@@ -10521,6 +10301,7 @@ Start now: evidence(task_id="{task_id}")
task_id=task_id, task_id=task_id,
initial_prompt=prompt, initial_prompt=prompt,
git_context=self._task_git_context(task), git_context=self._task_git_context(task),
spawned_by="_maybe_spawn_pm_closure",
) )
async def _dispatch_pm_closure_work(self, client: httpx.AsyncClient) -> None: async def _dispatch_pm_closure_work(self, client: httpx.AsyncClient) -> None:
@@ -10766,6 +10547,7 @@ Never `commit`, never write code, never run `git`. PMs coordinate.
task_id=task["id"], task_id=task["id"],
initial_prompt=self._build_dev_prompt(task), initial_prompt=self._build_dev_prompt(task),
git_context=self._task_git_context(task), git_context=self._task_git_context(task),
spawned_by="_respawn_dev_if_inactive",
) )
async def _spawn_pending_dev( async def _spawn_pending_dev(
@@ -10801,6 +10583,7 @@ Never `commit`, never write code, never run `git`. PMs coordinate.
task_id=task["id"], task_id=task["id"],
initial_prompt=self._get_prompt_for_agent(agent_slug, task), initial_prompt=self._get_prompt_for_agent(agent_slug, task),
git_context=self._task_git_context(task), git_context=self._task_git_context(task),
spawned_by="_spawn_pending_dev",
) )
@staticmethod @staticmethod
@@ -10936,6 +10719,7 @@ Never `commit`, never write code, never run `git`. PMs coordinate.
task_id=task["id"], task_id=task["id"],
initial_prompt=self._build_qa_prompt(task), initial_prompt=self._build_qa_prompt(task),
git_context=self._task_git_context(task), git_context=self._task_git_context(task),
spawned_by="_spawn_assigned_qa",
) )
return True return True
@@ -10983,6 +10767,7 @@ Never `commit`, never write code, never run `git`. PMs coordinate.
task_id=task["id"], task_id=task["id"],
initial_prompt=self._build_qa_prompt(task), initial_prompt=self._build_qa_prompt(task),
git_context=self._task_git_context(task), git_context=self._task_git_context(task),
spawned_by="_dispatch_qa_work",
) )
# Only spawn one QA at a time per cell # Only spawn one QA at a time per cell
break break
@@ -11015,6 +10800,7 @@ Never `commit`, never write code, never run `git`. PMs coordinate.
task_id=task["id"], task_id=task["id"],
initial_prompt=self._build_pr_review_prompt(task), initial_prompt=self._build_pr_review_prompt(task),
git_context=self._task_git_context(task), git_context=self._task_git_context(task),
spawned_by="_dispatch_pr_review_work",
) )
break break
@@ -11050,6 +10836,7 @@ Never `commit`, never write code, never run `git`. PMs coordinate.
task_id=task["id"], task_id=task["id"],
initial_prompt=self._build_pr_gate_prompt(task), initial_prompt=self._build_pr_gate_prompt(task),
git_context=self._task_git_context(task), git_context=self._task_git_context(task),
spawned_by="_dispatch_pr_gate_work",
) )
async def _dispatch_doc_work(self, client: httpx.AsyncClient) -> None: async def _dispatch_doc_work(self, client: httpx.AsyncClient) -> None:
@@ -11112,6 +10899,7 @@ Never `commit`, never write code, never run `git`. PMs coordinate.
task_id=task["id"], task_id=task["id"],
initial_prompt=self._build_doc_prompt(task), initial_prompt=self._build_doc_prompt(task),
git_context=self._task_git_context(task), git_context=self._task_git_context(task),
spawned_by="_auto_assign_doc",
) )
async def _doc_dispatch_one( async def _doc_dispatch_one(
@@ -11172,6 +10960,7 @@ Never `commit`, never write code, never run `git`. PMs coordinate.
task_id=task["id"], task_id=task["id"],
initial_prompt=self._build_doc_prompt(task), initial_prompt=self._build_doc_prompt(task),
git_context=self._task_git_context(task), git_context=self._task_git_context(task),
spawned_by="_respawn_doc_if_assigned",
) )
return True return True
@@ -11337,6 +11126,7 @@ Never `commit`, never write code, never run `git`. PMs coordinate.
task_id=task["id"], task_id=task["id"],
initial_prompt=self._build_pm_review_prompt(task), initial_prompt=self._build_pm_review_prompt(task),
git_context=self._task_git_context(task), git_context=self._task_git_context(task),
spawned_by="_dispatch_pm_review_work",
) )
continue continue
@@ -11365,6 +11155,7 @@ Never `commit`, never write code, never run `git`. PMs coordinate.
task_id=task["id"], task_id=task["id"],
initial_prompt=self._build_pm_review_prompt(task), initial_prompt=self._build_pm_review_prompt(task),
git_context=self._task_git_context(task), git_context=self._task_git_context(task),
spawned_by="_dispatch_pm_review_work",
) )
break break
@@ -11391,6 +11182,7 @@ Never `commit`, never write code, never run `git`. PMs coordinate.
task_id=task["id"], task_id=task["id"],
initial_prompt=self._build_marketing_prompt(task), initial_prompt=self._build_marketing_prompt(task),
git_context=self._task_git_context(task), git_context=self._task_git_context(task),
spawned_by="_dispatch_marketing_work",
) )
break break
@@ -11470,6 +11262,7 @@ Never `commit`, never write code, never run `git`. PMs coordinate.
task_id=task["id"], task_id=task["id"],
initial_prompt=self._build_pm_blocker_prompt(task), initial_prompt=self._build_pm_blocker_prompt(task),
git_context=self._task_git_context(task), git_context=self._task_git_context(task),
spawned_by="_dispatch_blocker_work",
) )
break break
@@ -11550,6 +11343,7 @@ Never `commit`, never write code, never run `git`. PMs coordinate.
task_id=str(task_id), task_id=str(task_id),
initial_prompt=self._get_prompt_for_agent(agent_slug, task), initial_prompt=self._get_prompt_for_agent(agent_slug, task),
git_context=self._task_git_context(task), git_context=self._task_git_context(task),
spawned_by="_dispatch_claimed_without_agent",
) )
break break
@@ -11617,6 +11411,7 @@ Never `commit`, never write code, never run `git`. PMs coordinate.
await self.spawn_agent( await self.spawn_agent(
agent_id=agent_slug, agent_id=agent_slug,
initial_prompt=self._build_escalation_prompt(notif), initial_prompt=self._build_escalation_prompt(notif),
spawned_by="_dispatch_escalation_work",
) )
break break
@@ -11647,6 +11442,7 @@ Never `commit`, never write code, never run `git`. PMs coordinate.
await self.spawn_agent( await self.spawn_agent(
agent_id=agent_slug, agent_id=agent_slug,
initial_prompt=self._build_approval_prompt(notif), initial_prompt=self._build_approval_prompt(notif),
spawned_by="_dispatch_approval_work",
) )
break break
@@ -11671,6 +11467,7 @@ Never `commit`, never write code, never run `git`. PMs coordinate.
await self.spawn_agent( await self.spawn_agent(
agent_id="auditor", agent_id="auditor",
initial_prompt=self._build_audit_prompt(alert), initial_prompt=self._build_audit_prompt(alert),
spawned_by="_dispatch_audit_work",
) )
return return
@@ -11951,6 +11748,7 @@ Never `commit`, never write code, never run `git`. PMs coordinate.
await self.spawn_agent( await self.spawn_agent(
agent_id=agent_slug, agent_id=agent_slug,
initial_prompt=self._build_a2a_prompt(notif), initial_prompt=self._build_a2a_prompt(notif),
spawned_by="_dispatch_a2a_work",
) )
break break
-1
View File
@@ -20,5 +20,4 @@ __all__ = [
"remediation", "remediation",
"role_config", "role_config",
"tracing_gate", "tracing_gate",
"trigger_filter",
] ]
+25 -16
View File
@@ -24,6 +24,7 @@ from roboco.foundation.policy import lifecycle as spec_module
from roboco.foundation.policy.batch import is_batch_umbrella from roboco.foundation.policy.batch import is_batch_umbrella
from roboco.foundation.policy.content import markers from roboco.foundation.policy.content import markers
from roboco.foundation.policy.content.validators import reject_trivial from roboco.foundation.policy.content.validators import reject_trivial
from roboco.services.gateway.choreographer._protocol import actor_context_fields
from roboco.services.gateway.choreographer._verb_runner import VerbRunner from roboco.services.gateway.choreographer._verb_runner import VerbRunner
from roboco.services.gateway.claim_guards import ( from roboco.services.gateway.claim_guards import (
already_active_guard, already_active_guard,
@@ -1877,10 +1878,11 @@ class Choreographer:
context_briefing=briefing, context_briefing=briefing,
), ),
) )
actor_slug, agent_team = actor_context_fields(agent)
spec_ctx = spec_module.Context( spec_ctx = spec_module.Context(
actor_id=agent_id, actor_id=agent_id,
actor_slug=getattr(agent, "slug", None) if agent is not None else None, actor_slug=actor_slug,
agent_team=str(agent.team) if agent is not None and agent.team else None, agent_team=agent_team,
original_developer_slug=_extract_original_developer(t), original_developer_slug=_extract_original_developer(t),
notes=notes, notes=notes,
) )
@@ -3389,10 +3391,11 @@ class Choreographer:
task_id=task_id, task_id=task_id,
verb="unclaim", verb="unclaim",
) )
actor_slug, agent_team = actor_context_fields(agent)
spec_ctx = spec_module.Context( spec_ctx = spec_module.Context(
actor_id=agent_id, actor_id=agent_id,
actor_slug=getattr(agent, "slug", None) if agent is not None else None, actor_slug=actor_slug,
agent_team=str(agent.team) if agent is not None and agent.team else None, agent_team=agent_team,
original_developer_slug=_extract_original_developer(t), original_developer_slug=_extract_original_developer(t),
) )
decision = spec_module.can_invoke_intent(role, "unclaim", t, spec_ctx) decision = spec_module.can_invoke_intent(role, "unclaim", t, spec_ctx)
@@ -3633,10 +3636,11 @@ class Choreographer:
task_id=task_id, task_id=task_id,
verb="resume", verb="resume",
) )
actor_slug, agent_team = actor_context_fields(agent)
spec_ctx = spec_module.Context( spec_ctx = spec_module.Context(
actor_id=agent_id, actor_id=agent_id,
actor_slug=getattr(agent, "slug", None) if agent is not None else None, actor_slug=actor_slug,
agent_team=str(agent.team) if agent is not None and agent.team else None, agent_team=agent_team,
original_developer_slug=_extract_original_developer(t), original_developer_slug=_extract_original_developer(t),
) )
decision = spec_module.can_invoke_intent(role, "resume", t, spec_ctx) decision = spec_module.can_invoke_intent(role, "resume", t, spec_ctx)
@@ -5594,10 +5598,11 @@ class Choreographer:
task_id=task_id, task_id=task_id,
verb="submit_up", verb="submit_up",
) )
actor_slug, agent_team = actor_context_fields(agent)
spec_ctx = spec_module.Context( spec_ctx = spec_module.Context(
actor_id=pm_agent_id, actor_id=pm_agent_id,
actor_slug=getattr(agent, "slug", None) if agent is not None else None, actor_slug=actor_slug,
agent_team=str(agent.team) if agent is not None and agent.team else None, agent_team=agent_team,
original_developer_slug=_extract_original_developer(t), original_developer_slug=_extract_original_developer(t),
notes=notes, notes=notes,
) )
@@ -6530,10 +6535,11 @@ class Choreographer:
verb="submit_root", verb="submit_root",
) )
role = spec_module.Role(role_str) role = spec_module.Role(role_str)
actor_slug, agent_team = actor_context_fields(agent)
spec_ctx = spec_module.Context( spec_ctx = spec_module.Context(
actor_id=main_pm_agent_id, actor_id=main_pm_agent_id,
actor_slug=getattr(agent, "slug", None) if agent is not None else None, actor_slug=actor_slug,
agent_team=str(agent.team) if agent is not None and agent.team else None, agent_team=agent_team,
notes=notes, notes=notes,
) )
soup = self._free_text_soup(checks=(("notes", notes, 10),)) soup = self._free_text_soup(checks=(("notes", notes, 10),))
@@ -6882,10 +6888,11 @@ class Choreographer:
task_id=task_id, task_id=task_id,
verb="complete", verb="complete",
) )
actor_slug, agent_team = actor_context_fields(agent)
spec_ctx = spec_module.Context( spec_ctx = spec_module.Context(
actor_id=agent_id, actor_id=agent_id,
actor_slug=getattr(agent, "slug", None) if agent is not None else None, actor_slug=actor_slug,
agent_team=str(agent.team) if agent is not None and agent.team else None, agent_team=agent_team,
original_developer_slug=_extract_original_developer(t), original_developer_slug=_extract_original_developer(t),
) )
if soup := await self._guard_free_text( if soup := await self._guard_free_text(
@@ -7092,10 +7099,11 @@ class Choreographer:
task_id=task_id, task_id=task_id,
verb="escalate_up", verb="escalate_up",
) )
actor_slug, agent_team = actor_context_fields(me)
spec_ctx = spec_module.Context( spec_ctx = spec_module.Context(
actor_id=pm_agent_id, actor_id=pm_agent_id,
actor_slug=getattr(me, "slug", None) if me is not None else None, actor_slug=actor_slug,
agent_team=str(me.team) if me is not None and me.team else None, agent_team=agent_team,
original_developer_slug=_extract_original_developer(t), original_developer_slug=_extract_original_developer(t),
notes=reason, notes=reason,
) )
@@ -7263,10 +7271,11 @@ class Choreographer:
task_id=task_id, task_id=task_id,
verb="escalate_to_ceo", verb="escalate_to_ceo",
) )
actor_slug, agent_team = actor_context_fields(me)
spec_ctx = spec_module.Context( spec_ctx = spec_module.Context(
actor_id=agent_id, actor_id=agent_id,
actor_slug=getattr(me, "slug", None) if me is not None else None, actor_slug=actor_slug,
agent_team=str(me.team) if me is not None and me.team else None, agent_team=agent_team,
original_developer_slug=_extract_original_developer(t), original_developer_slug=_extract_original_developer(t),
notes=reason, notes=reason,
) )
@@ -132,3 +132,16 @@ class ChoreographerHelpers:
task: Any | None = None, task: Any | None = None,
) -> Envelope: ) -> Envelope:
raise NotImplementedError raise NotImplementedError
def actor_context_fields(agent: Any) -> tuple[str | None, str | None]:
"""``(actor_slug, agent_team)`` for a spec ``Context``, None-agent safe.
Every verb builds its policy Context with the same two conditional
kwargs; inlining them pushed nine verbs over the xenon B threshold when
the team-match sweep added ``agent_team`` to all 27 sites.
"""
if agent is None:
return None, None
team = getattr(agent, "team", None)
return getattr(agent, "slug", None), str(team) if team else None
+4 -2
View File
@@ -44,6 +44,7 @@ from roboco.foundation.policy import lifecycle as spec_module
from roboco.foundation.policy import tracing as _tr from roboco.foundation.policy import tracing as _tr
from roboco.foundation.policy.content import ContentValidationError, markers from roboco.foundation.policy.content import ContentValidationError, markers
from roboco.services.content_notes import apply_structured_note from roboco.services.content_notes import apply_structured_note
from roboco.services.gateway.choreographer._protocol import actor_context_fields
from roboco.services.gateway.envelope import Envelope from roboco.services.gateway.envelope import Envelope
from roboco.services.gateway.evidence_builder import build_evidence_for_task from roboco.services.gateway.evidence_builder import build_evidence_for_task
@@ -648,10 +649,11 @@ class QAMixin(_Base):
return gate_rejection return gate_rejection
briefing = await self._briefing_for(qa_agent_id, task_id) briefing = await self._briefing_for(qa_agent_id, task_id)
actor_slug, agent_team = actor_context_fields(agent)
spec_ctx = spec_module.Context( spec_ctx = spec_module.Context(
actor_id=qa_agent_id, actor_id=qa_agent_id,
actor_slug=getattr(agent, "slug", None) if agent is not None else None, actor_slug=actor_slug,
agent_team=str(agent.team) if agent is not None and agent.team else None, agent_team=agent_team,
original_developer_slug=_extract_original_developer(t), original_developer_slug=_extract_original_developer(t),
notes=notes, notes=notes,
issues=tuple(issues), issues=tuple(issues),
-129
View File
@@ -1,129 +0,0 @@
"""Stale-trigger cleanup + cooldown decisions.
Decides whether to spawn an agent for a (task, trigger) pair. Reads counts
from caller (recent spawns within window). Pure function caller queries
the gateway_triggers table and persists the resulting decision.
"""
from __future__ import annotations
from dataclasses import dataclass
from enum import StrEnum
from typing import Any
from roboco.services.gateway.claimant_lock import is_stale
class TriggerKind(StrEnum):
A2A = "a2a"
NOTIFICATION = "notification"
SCAN = "scan"
ESCALATION = "escalation"
class SpawnDecision(StrEnum):
SPAWN = "spawn"
QUEUE = "queue"
DROP = "drop"
@dataclass(frozen=True)
class Decision:
outcome: SpawnDecision
reason: str
@dataclass(frozen=True)
class SpawnConfig:
"""Numeric tunables for spawn-gating decisions."""
cooldown_seconds: int
role_rate_per_minute: int
claim_stale_seconds: int
@dataclass(frozen=True)
class TriggerContext:
"""Trigger identity and recent-spawn counts passed by the caller."""
kind: TriggerKind
skill: str | None
recent_spawns_for_task: int
recent_spawns_for_role: int
# Provider rate-limit fields. Optional — callers that don't know the
# provider (e.g. no-task spawns) leave these at their defaults so the
# gate is a no-op.
provider: str | None = None
provider_rate_limited: bool = False
_TERMINAL_STATUSES: frozenset[str] = frozenset({"completed", "cancelled"})
# A2A code_review only relevant when task is in awaiting_qa or earlier review states
_A2A_CODE_REVIEW_RELEVANT_STATES: frozenset[str] = frozenset(
{"awaiting_qa", "claimed", "in_progress", "verifying"}
)
def _stale_trigger_decision(task: Any, trigger: TriggerContext) -> Decision | None:
"""DROP decision for a trigger that no longer applies to the task, else None."""
if task.status in _TERMINAL_STATUSES:
return Decision(SpawnDecision.DROP, "task in terminal state — trigger stale")
if (
trigger.kind is TriggerKind.A2A
and trigger.skill == "code_review"
and task.status not in _A2A_CODE_REVIEW_RELEVANT_STATES
):
return Decision(
SpawnDecision.DROP,
f"a2a code_review for task in {task.status} — stale",
)
return None
def decide_spawn(
*,
task: Any,
trigger: TriggerContext,
config: SpawnConfig,
) -> Decision:
"""Apply five rules in order.
stale > provider-rate-limit > claimant-lock > task-cooldown > role-rate
"""
# 1. Stale-trigger cleanup
stale = _stale_trigger_decision(task, trigger)
if stale is not None:
return stale
# 2. Provider rate-limit gate
if trigger.provider_rate_limited:
return Decision(
SpawnDecision.QUEUE,
f"provider {trigger.provider or 'unknown'} rate-limited",
)
# 3. Single-claimant invariant
if task.active_claimant_id is not None and not is_stale(
task, threshold_seconds=config.claim_stale_seconds
):
return Decision(
SpawnDecision.QUEUE,
"task has active claimant with fresh heartbeat",
)
# 4. Per-task spawn cooldown
if trigger.recent_spawns_for_task >= 1:
return Decision(
SpawnDecision.QUEUE,
f"per-task spawn cooldown ({config.cooldown_seconds}s) active",
)
# 5. Per-role rate limit
if trigger.recent_spawns_for_role >= config.role_rate_per_minute:
return Decision(
SpawnDecision.QUEUE,
f"role spawn rate limit ({config.role_rate_per_minute}/min) reached",
)
return Decision(SpawnDecision.SPAWN, "all gates clear")
+25 -16
View File
@@ -249,6 +249,16 @@ def _select_ci_head_run(runs: list[dict[str, Any]]) -> dict[str, Any]:
return max(same_head, key=lambda r: int(r.get("run_attempt") or 0)) return max(same_head, key=lambda r: int(r.get("run_attempt") or 0))
def _api_base() -> str:
"""GitHub REST base URL — honors ``settings.github_api_base_url``.
Five call sites already read the setting (CI runs, open-PR list); the
PR create/merge/branch sites hardcoded the public host, which broke any
GitHub Enterprise or test override. One helper keeps them uniform.
"""
return settings.github_api_base_url.rstrip("/")
@dataclass(frozen=True) @dataclass(frozen=True)
class _CiRunQuery: class _CiRunQuery:
"""Bundle of per-project inputs to a CI-run fetch (owner/repo, branch, token, """Bundle of per-project inputs to a CI-run fetch (owner/repo, branch, token,
@@ -1856,7 +1866,7 @@ class GitService(BaseService):
"""Return the first open PR for head→base, or None.""" """Return the first open PR for head→base, or None."""
async with httpx.AsyncClient(timeout=_default_git_timeout()) as client: async with httpx.AsyncClient(timeout=_default_git_timeout()) as client:
existing = await client.get( existing = await client.get(
f"https://api.github.com/repos/{owner}/{repo}/pulls", f"{_api_base()}/repos/{owner}/{repo}/pulls",
headers={ headers={
"Authorization": f"Bearer {git_token}", "Authorization": f"Bearer {git_token}",
"Accept": "application/vnd.github+json", "Accept": "application/vnd.github+json",
@@ -2107,7 +2117,7 @@ class GitService(BaseService):
try: try:
async with httpx.AsyncClient(timeout=_default_git_timeout()) as client: async with httpx.AsyncClient(timeout=_default_git_timeout()) as client:
return await client.post( return await client.post(
f"https://api.github.com/repos/{owner}/{repo}/pulls", f"{_api_base()}/repos/{owner}/{repo}/pulls",
headers={ headers={
"Authorization": f"Bearer {git_token}", "Authorization": f"Bearer {git_token}",
"Accept": "application/vnd.github+json", "Accept": "application/vnd.github+json",
@@ -2304,7 +2314,7 @@ class GitService(BaseService):
try: try:
async with httpx.AsyncClient(timeout=_default_git_timeout()) as client: async with httpx.AsyncClient(timeout=_default_git_timeout()) as client:
resp = await client.patch( resp = await client.patch(
f"https://api.github.com/repos/{owner}/{repo}/pulls/{pr_number}", f"{_api_base()}/repos/{owner}/{repo}/pulls/{pr_number}",
headers={ headers={
"Authorization": f"Bearer {git_token}", "Authorization": f"Bearer {git_token}",
"Accept": "application/vnd.github+json", "Accept": "application/vnd.github+json",
@@ -2345,7 +2355,7 @@ class GitService(BaseService):
try: try:
async with httpx.AsyncClient(timeout=_default_git_timeout()) as client: async with httpx.AsyncClient(timeout=_default_git_timeout()) as client:
resp = await client.post( resp = await client.post(
f"https://api.github.com/repos/{owner}/{repo}/pulls/" f"{_api_base()}/repos/{owner}/{repo}/pulls/"
f"{pr_number}/requested_reviewers", f"{pr_number}/requested_reviewers",
headers={ headers={
"Authorization": f"Bearer {git_token}", "Authorization": f"Bearer {git_token}",
@@ -2722,8 +2732,7 @@ class GitService(BaseService):
try: try:
async with httpx.AsyncClient(timeout=_default_git_timeout()) as client: async with httpx.AsyncClient(timeout=_default_git_timeout()) as client:
return await client.put( return await client.put(
f"https://api.github.com/repos/{owner}/{repo}/pulls/" f"{_api_base()}/repos/{owner}/{repo}/pulls/{pr_number}/merge",
f"{pr_number}/merge",
headers={ headers={
"Authorization": f"Bearer {git_token}", "Authorization": f"Bearer {git_token}",
"Accept": "application/vnd.github+json", "Accept": "application/vnd.github+json",
@@ -2820,7 +2829,7 @@ class GitService(BaseService):
try: try:
async with httpx.AsyncClient(timeout=10.0) as client: async with httpx.AsyncClient(timeout=10.0) as client:
resp = await client.get( resp = await client.get(
f"https://api.github.com/repos/{owner}/{repo}/pulls", f"{_api_base()}/repos/{owner}/{repo}/pulls",
params={"base": branch, "state": "open", "per_page": 1}, params={"base": branch, "state": "open", "per_page": 1},
headers={ headers={
"Authorization": f"Bearer {git_token}", "Authorization": f"Bearer {git_token}",
@@ -2857,7 +2866,7 @@ class GitService(BaseService):
try: try:
async with httpx.AsyncClient(timeout=10.0) as client: async with httpx.AsyncClient(timeout=10.0) as client:
await client.delete( await client.delete(
f"https://api.github.com/repos/{owner}/{repo}/git/refs/heads/{branch}", f"{_api_base()}/repos/{owner}/{repo}/git/refs/heads/{branch}",
headers={ headers={
"Authorization": f"Bearer {git_token}", "Authorization": f"Bearer {git_token}",
"Accept": "application/vnd.github+json", "Accept": "application/vnd.github+json",
@@ -2877,7 +2886,7 @@ class GitService(BaseService):
try: try:
async with httpx.AsyncClient(timeout=10.0) as client: async with httpx.AsyncClient(timeout=10.0) as client:
pr_resp = await client.get( pr_resp = await client.get(
f"https://api.github.com/repos/{owner}/{repo}/pulls/{pr_number}", f"{_api_base()}/repos/{owner}/{repo}/pulls/{pr_number}",
headers={ headers={
"Authorization": f"Bearer {git_token}", "Authorization": f"Bearer {git_token}",
"Accept": "application/vnd.github+json", "Accept": "application/vnd.github+json",
@@ -2934,7 +2943,7 @@ class GitService(BaseService):
try: try:
async with httpx.AsyncClient(timeout=_default_git_timeout()) as client: async with httpx.AsyncClient(timeout=_default_git_timeout()) as client:
resp = await client.get( resp = await client.get(
f"https://api.github.com/repos/{owner}/{repo}", f"{_api_base()}/repos/{owner}/{repo}",
headers={ headers={
"Authorization": f"Bearer {git_token}", "Authorization": f"Bearer {git_token}",
"Accept": "application/vnd.github+json", "Accept": "application/vnd.github+json",
@@ -3677,7 +3686,7 @@ class GitService(BaseService):
try: try:
async with httpx.AsyncClient(timeout=_default_git_timeout()) as client: async with httpx.AsyncClient(timeout=_default_git_timeout()) as client:
resp = await client.get( resp = await client.get(
f"https://api.github.com/repos/{owner}/{repo}/pulls/{pr_number}", f"{_api_base()}/repos/{owner}/{repo}/pulls/{pr_number}",
headers={ headers={
"Authorization": f"Bearer {git_token}", "Authorization": f"Bearer {git_token}",
"Accept": "application/vnd.github+json", "Accept": "application/vnd.github+json",
@@ -3789,7 +3798,7 @@ class GitService(BaseService):
try: try:
async with httpx.AsyncClient(timeout=_default_git_timeout()) as client: async with httpx.AsyncClient(timeout=_default_git_timeout()) as client:
resp = await client.get( resp = await client.get(
f"https://api.github.com/repos/{owner}/{repo}/pulls/{pr_number}", f"{_api_base()}/repos/{owner}/{repo}/pulls/{pr_number}",
headers={ headers={
"Authorization": f"Bearer {git_token}", "Authorization": f"Bearer {git_token}",
"Accept": "application/vnd.github+json", "Accept": "application/vnd.github+json",
@@ -4166,7 +4175,7 @@ class GitService(BaseService):
} }
async with httpx.AsyncClient(timeout=_default_git_timeout()) as client: async with httpx.AsyncClient(timeout=_default_git_timeout()) as client:
existing = await client.get( existing = await client.get(
f"https://api.github.com/repos/{owner}/{repo}/pulls/{pr_number}", f"{_api_base()}/repos/{owner}/{repo}/pulls/{pr_number}",
headers=headers, headers=headers,
) )
already_closed = ( already_closed = (
@@ -4175,13 +4184,13 @@ class GitService(BaseService):
if not already_closed: if not already_closed:
if comment: if comment:
await client.post( await client.post(
f"https://api.github.com/repos/{owner}/{repo}/issues/" f"{_api_base()}/repos/{owner}/{repo}/issues/"
f"{pr_number}/comments", f"{pr_number}/comments",
headers=headers, headers=headers,
json={"body": comment}, json={"body": comment},
) )
resp = await client.patch( resp = await client.patch(
f"https://api.github.com/repos/{owner}/{repo}/pulls/{pr_number}", f"{_api_base()}/repos/{owner}/{repo}/pulls/{pr_number}",
headers=headers, headers=headers,
json={"state": "closed"}, json={"state": "closed"},
) )
@@ -4239,7 +4248,7 @@ class GitService(BaseService):
try: try:
async with httpx.AsyncClient(timeout=_default_git_timeout()) as client: async with httpx.AsyncClient(timeout=_default_git_timeout()) as client:
resp = await client.get( resp = await client.get(
f"https://api.github.com/repos/{owner}/{repo}/pulls/{pr_number}", f"{_api_base()}/repos/{owner}/{repo}/pulls/{pr_number}",
headers={ headers={
"Authorization": f"Bearer {git_token}", "Authorization": f"Bearer {git_token}",
"Accept": "application/vnd.github+json", "Accept": "application/vnd.github+json",
+20 -2
View File
@@ -2488,7 +2488,7 @@ class TaskService(BaseService):
# tight loop hammering the orchestrator. # tight loop hammering the orchestrator.
task.last_heartbeat_at = now task.last_heartbeat_at = now
# Single-claimant invariant (alembic 006): claimant_lock.try_acquire # Single-claimant invariant (alembic 006): claimant_lock.try_acquire
# and trigger_filter.decide_spawn both branch on this column. Was # branches on this column. Was
# declared but never written; now wired so the # declared but never written; now wired so the
# invariant is functional. # invariant is functional.
task.active_claimant_id = cast("Any", agent_id) task.active_claimant_id = cast("Any", agent_id)
@@ -5014,6 +5014,24 @@ class TaskService(BaseService):
return await self.escalate_to_ceo(task_id, "main_pm") return await self.escalate_to_ceo(task_id, "main_pm")
return None return None
async def open_pr_ref(self, task: TaskTable) -> WorkSessionTable | None:
"""The task's work session iff its PR is recorded still open.
Backs the admin-override open-PR refusal: completing a task whose PR
is open strands its commits unmerged. Only ``pr_status == "open"``
counts merged is safe, closed was a deliberate discard, and a
missing session/status is unknowable here.
"""
if not task.work_session_id:
return None
result = await self.session.execute(
select(WorkSessionTable).where(WorkSessionTable.id == task.work_session_id)
)
ws = result.scalar_one_or_none()
if ws is not None and ws.pr_status == "open":
return ws
return None
async def _assert_pr_merged_for_complete(self, task: TaskTable) -> bool: async def _assert_pr_merged_for_complete(self, task: TaskTable) -> bool:
"""True if the task's PR is merged (or no PR gate applies). """True if the task's PR is merged (or no PR gate applies).
@@ -8290,7 +8308,7 @@ class TaskService(BaseService):
# tasks will misclassify the live claim as abandoned. # tasks will misclassify the live claim as abandoned.
task.last_heartbeat_at = now task.last_heartbeat_at = now
# Single-claimant invariant — see _finalize_claim. Same column # Single-claimant invariant — see _finalize_claim. Same column
# used by claimant_lock + trigger_filter. Cleared by QA pass/fail # used by claimant_lock. Cleared by QA pass/fail
# and doc-complete when the review hand-off finishes. # and doc-complete when the review hand-off finishes.
task.active_claimant_id = cast("Any", agent_id) task.active_claimant_id = cast("Any", agent_id)
await self.session.flush() await self.session.flush()
+23
View File
@@ -22,6 +22,16 @@ Behaviour:
collection time when no Postgres is reachable on `localhost:5432`. Set collection time when no Postgres is reachable on `localhost:5432`. Set
`ROBOCO_TEST_DB_HOST`, `ROBOCO_TEST_DB_PORT`, `ROBOCO_TEST_DB_USER`, or `ROBOCO_TEST_DB_HOST`, `ROBOCO_TEST_DB_PORT`, `ROBOCO_TEST_DB_USER`, or
`ROBOCO_TEST_DB_PASSWORD` to override. `ROBOCO_TEST_DB_PASSWORD` to override.
Redis isolation:
No test uses a real Redis, so `_no_live_redis` (autouse) points the
computed `settings.redis_url` at an unreachable port for every test.
Three test families were caught writing real keys into whatever Redis
listens on localhost:6379 (self-heal notify-dedupe, the rate-limit
tracker's NO-TTL "provider rate-limited" blob, notification
purpose-dedupe) order/state poison for anything reading the real
instance. Every production Redis path is fail-open by design, so an
unreachable port keeps behavior identical to "no redis available".
""" """
from __future__ import annotations from __future__ import annotations
@@ -36,6 +46,7 @@ from uuid import UUID, uuid4
import asyncpg import asyncpg
import pytest import pytest
import pytest_asyncio import pytest_asyncio
from roboco.config import settings as _settings
from roboco.db import tables as roboco_tables from roboco.db import tables as roboco_tables
from roboco.db.base import Base from roboco.db.base import Base
from roboco.db.tables import ( from roboco.db.tables import (
@@ -67,6 +78,18 @@ if TYPE_CHECKING:
from collections.abc import AsyncIterator from collections.abc import AsyncIterator
@pytest.fixture(autouse=True)
def _no_live_redis(monkeypatch: pytest.MonkeyPatch) -> None:
"""Keep every test off the real localhost Redis (see module docstring).
``settings.redis_url`` is a computed property, so its inputs are patched.
Port 1 refuses instantly fail-open code paths behave exactly as with no
Redis, and nothing can read or write live keys.
"""
monkeypatch.setattr(_settings, "redis_host", "127.0.0.1")
monkeypatch.setattr(_settings, "redis_port", 1)
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
# Test DB endpoint discovery — env-overridable. # Test DB endpoint discovery — env-overridable.
# #
+45
View File
@@ -0,0 +1,45 @@
"""e2e lifecycle smoke harness — collection gate + the stack fixture.
Scripted-agent smoke: an in-process RoboCo API (real routers, real
middleware, real gateway/choreographer/services) over the ephemeral test
Postgres, a local bare git origin standing in for GitHub, and a fake
GitHub REST layer whose merges are REAL git merges on that origin. A
deterministic driver calls the REAL MCP flow/do tool functions no LLM
anywhere so seam bugs (toolgate schema drift, squash merges, stale
refs, workspace routing) die here instead of in a live run.
Gating: excluded from the default suite (`make quality`); runs via
`make e2e-smoke` (sets ROBOCO_E2E_SMOKE=1). Needs the test Postgres
reachable and git on PATH, nothing else.
"""
from __future__ import annotations
import os
from typing import TYPE_CHECKING
import pytest
from tests.e2e_smoke.harness import build_e2e_stack
if TYPE_CHECKING:
from collections.abc import Iterator
from tests.e2e_smoke.harness import E2EStack
def pytest_collection_modifyitems(
config: pytest.Config, items: list[pytest.Item]
) -> None:
if os.environ.get("ROBOCO_E2E_SMOKE") == "1":
return
skip = pytest.mark.skip(reason="e2e smoke runs via `make e2e-smoke` only")
for item in items:
if "tests/e2e_smoke" in str(item.path):
item.add_marker(skip)
@pytest.fixture(scope="session")
def e2e_stack(
_test_database_url: str, tmp_path_factory: pytest.TempPathFactory
) -> Iterator[E2EStack]:
yield from build_e2e_stack(_test_database_url, tmp_path_factory)
+472
View File
@@ -0,0 +1,472 @@
"""e2e smoke harness — in-process RoboCo stack + scripted-agent driver.
Pieces (all REAL except GitHub and the LLM):
- The API: the real v1 flow/do routers + real middleware/exception handlers,
served by uvicorn in a thread, over the ephemeral test Postgres (the app's
own lazy engine is pointed at it by patching ``settings.database_*`` and
resetting ``_DbHolder``).
- Git: a local bare origin whose path CONTAINS ``github.com/<owner>/<repo>``
``_parse_git_url`` extracts owner/repo from it while clone/fetch/push
run tokenless over the local protocol.
- GitHub REST: a fake ``/_github`` router mounted on the same app
(``settings.github_api_base_url`` points at it). PR state lives in memory;
merges perform REAL git merges (squash included) on the bare origin, so
downstream git logic (cherry checks, freshness, branch sync) sees reality.
- Agents: ``ScriptedAgent`` reloads the REAL ``roboco.mcp.flow_server`` /
``do_server`` modules with that agent's env (id, role, role-scoped
manifest built from the real ``role_config``) and calls the REAL tool
functions, which POST to the in-process API over loopback HTTP.
"""
from __future__ import annotations
import asyncio
import importlib
import json
import os
import socket
import subprocess
import threading
import time
from contextlib import suppress
from dataclasses import dataclass, field
from typing import TYPE_CHECKING, Any
import pytest
import uvicorn
from cryptography.fernet import Fernet
from fastapi import APIRouter, FastAPI, Request
from fastapi.responses import JSONResponse
from sqlalchemy.engine.url import make_url
from sqlalchemy.ext.asyncio import async_sessionmaker, create_async_engine
if TYPE_CHECKING:
from collections.abc import Iterator
from pathlib import Path
from types import ModuleType
from uuid import UUID
_OWNER = "e2e-smoke"
_REPO = "proj"
def _git(cwd: Path, *args: str) -> str:
res = subprocess.run(
["git", "-C", str(cwd), *args],
capture_output=True,
text=True,
check=True,
)
return res.stdout.strip()
# ---------------------------------------------------------------------------
# Fake GitHub REST — PR state in memory, merges as REAL git ops on the origin
# ---------------------------------------------------------------------------
@dataclass
class _FakeGitHub:
origin: Path
admin_clone: Path
prs: dict[int, dict[str, Any]] = field(default_factory=dict)
comments: list[dict[str, Any]] = field(default_factory=list)
next_number: int = 1
def create_pr(self, title: str, body: str, head: str, base: str) -> dict[str, Any]:
number = self.next_number
self.next_number += 1
pr = {
"number": number,
"html_url": f"https://github.com/{_OWNER}/{_REPO}/pull/{number}",
"title": title,
"body": body,
"state": "open",
"merged": False,
"head": {
"ref": head,
"sha": self._sha_of(head),
"repo": {"full_name": f"{_OWNER}/{_REPO}"},
},
"base": {"ref": base},
"user": {"login": "e2e-bot"},
"author_association": "MEMBER",
}
self.prs[number] = pr
return pr
def _sha_of(self, branch: str) -> str:
try:
return _git(self.origin, "rev-parse", branch)
except subprocess.CalledProcessError:
return "0" * 40
def merge_pr(self, number: int, merge_method: str) -> dict[str, Any]:
pr = self.prs[number]
head, base = pr["head"]["ref"], pr["base"]["ref"]
admin = self.admin_clone
_git(admin, "fetch", "origin", "--prune")
_git(admin, "checkout", "-B", base, f"origin/{base}")
if merge_method == "squash":
_git(admin, "merge", "--squash", f"origin/{head}")
_git(admin, "commit", "-m", f"{pr['title']} (#{number})")
else:
_git(
admin,
"merge",
"--no-ff",
"-m",
f"Merge pull request #{number} from {head}",
f"origin/{head}",
)
_git(admin, "push", "origin", base)
sha = _git(admin, "rev-parse", "HEAD")
pr["merged"] = True
pr["state"] = "closed"
return {
"merged": True,
"sha": sha,
"message": "Pull Request successfully merged",
}
def open_prs(self, head: str | None, base: str | None) -> list[dict[str, Any]]:
out = []
for pr in self.prs.values():
if pr["state"] != "open":
continue
if head and pr["head"]["ref"] != head.split(":", 1)[-1]:
continue
if base and pr["base"]["ref"] != base:
continue
out.append(pr)
return out
def _fake_github_router(gh: _FakeGitHub) -> APIRouter:
r = APIRouter(prefix="/_github")
@r.get("/repos/{owner}/{repo}")
async def repo_caps(owner: str, repo: str) -> dict[str, Any]:
return {
"allow_squash_merge": True,
"allow_merge_commit": True,
"allow_rebase_merge": False,
}
@r.get("/repos/{owner}/{repo}/pulls/{number}")
async def get_pr(owner: str, repo: str, number: int) -> JSONResponse:
pr = gh.prs.get(number)
if pr is None:
return JSONResponse({"message": "Not Found"}, status_code=404)
return JSONResponse(pr)
@r.get("/repos/{owner}/{repo}/pulls")
async def list_prs(
owner: str,
repo: str,
head: str | None = None,
base: str | None = None,
state: str = "open",
) -> list[dict[str, Any]]:
return gh.open_prs(head, base)
@r.post("/repos/{owner}/{repo}/pulls", status_code=201)
async def create_pr(owner: str, repo: str, request: Request) -> dict[str, Any]:
body = await request.json()
return gh.create_pr(
body["title"], body.get("body", ""), body["head"], body["base"]
)
@r.patch("/repos/{owner}/{repo}/pulls/{number}")
async def patch_pr(
owner: str, repo: str, number: int, request: Request
) -> JSONResponse:
pr = gh.prs.get(number)
if pr is None:
return JSONResponse({"message": "Not Found"}, status_code=404)
body = await request.json()
for key in ("title", "body", "state"):
if key in body:
pr[key] = body[key]
return JSONResponse(pr)
@r.put("/repos/{owner}/{repo}/pulls/{number}/merge")
async def merge_pr(
owner: str, repo: str, number: int, request: Request
) -> JSONResponse:
if number not in gh.prs:
return JSONResponse({"message": "Not Found"}, status_code=404)
body = await request.json()
try:
result = gh.merge_pr(number, body.get("merge_method", "merge"))
except subprocess.CalledProcessError as exc:
return JSONResponse(
{"message": f"Merge conflict: {exc.stderr}"}, status_code=409
)
return JSONResponse(result)
@r.post("/repos/{owner}/{repo}/pulls/{number}/requested_reviewers", status_code=201)
async def request_reviewers(
owner: str, repo: str, number: int, request: Request
) -> dict[str, Any]:
return gh.prs.get(number, {})
@r.post("/repos/{owner}/{repo}/issues/{number}/comments", status_code=201)
async def comment(
owner: str, repo: str, number: int, request: Request
) -> dict[str, Any]:
gh.comments.append({"number": number, "body": (await request.json())})
return {"id": len(gh.comments)}
@r.delete("/repos/{owner}/{repo}/git/refs/heads/{branch:path}", status_code=204)
async def delete_branch(owner: str, repo: str, branch: str) -> None:
with suppress(subprocess.CalledProcessError):
_git(gh.origin, "branch", "-D", branch)
return r
# ---------------------------------------------------------------------------
# Stack: settings patches + origin + app + uvicorn thread
# ---------------------------------------------------------------------------
@dataclass
class E2EStack:
base_url: str
root: Path
origin: Path
workspaces_root: Path
db_url: str
github: _FakeGitHub
def workspace_of(self, project_slug: str, team: str, agent_slug: str) -> Path:
return self.workspaces_root / project_slug / team / agent_slug
def run_db(self, coro_fn: Any) -> Any:
"""Run ``coro_fn(session)`` against a fresh engine/session and return."""
async def _run() -> Any:
engine = create_async_engine(self.db_url)
factory = async_sessionmaker(engine, expire_on_commit=False)
try:
async with factory() as session:
result = await coro_fn(session)
await session.commit()
return result
finally:
await engine.dispose()
return asyncio.run(_run())
def _free_port() -> int:
with socket.socket() as s:
s.bind(("127.0.0.1", 0))
return int(s.getsockname()[1])
def _seed_origin(root: Path) -> Path:
"""Bare origin at a path _parse_git_url can read owner/repo from."""
origin = root / "github.com" / _OWNER / f"{_REPO}.git"
origin.parent.mkdir(parents=True)
subprocess.run(
["git", "init", "--bare", "--initial-branch=master", str(origin)],
check=True,
capture_output=True,
)
seed = root / "seed-clone"
subprocess.run(
["git", "clone", str(origin), str(seed)], check=True, capture_output=True
)
_git(seed, "config", "user.name", "roboco-e2e")
_git(seed, "config", "user.email", "e2e@roboco.local")
(seed / "README.md").write_text("# e2e smoke project\n")
_git(seed, "add", "README.md")
_git(seed, "commit", "-m", "Initial commit")
_git(seed, "push", "origin", "master")
return origin
def _make_admin_clone(root: Path, origin: Path) -> Path:
admin = root / "gh-admin-clone"
subprocess.run(
["git", "clone", str(origin), str(admin)], check=True, capture_output=True
)
_git(admin, "config", "user.name", "fake-github")
_git(admin, "config", "user.email", "merge@github.local")
return admin
def _build_app(gh: _FakeGitHub) -> FastAPI:
from roboco.api.middleware import setup_middleware
from roboco.api.routes.health import router as health_router
from roboco.api.routes.v1 import do as do_module
from roboco.api.routes.v1 import flow_auditor as fa
from roboco.api.routes.v1 import flow_board as fb
from roboco.api.routes.v1 import flow_cell_pm as fcp
from roboco.api.routes.v1 import flow_dev as fd
from roboco.api.routes.v1 import flow_doc as fdoc
from roboco.api.routes.v1 import flow_main_pm as fmp
from roboco.api.routes.v1 import flow_pr_reviewer as fpr
from roboco.api.routes.v1 import flow_qa as fq
app = FastAPI(title="roboco-e2e-smoke")
setup_middleware(app)
app.include_router(health_router)
for module in (fd, fq, fdoc, fcp, fmp, fb, fa, fpr):
app.include_router(module.router)
app.include_router(do_module.router)
app.include_router(_fake_github_router(gh))
return app
def build_e2e_stack(
_test_database_url: str, tmp_path_factory: pytest.TempPathFactory
) -> Iterator[E2EStack]:
"""Generator behind the ``e2e_stack`` fixture (defined in conftest)."""
from roboco.config import settings
from roboco.db import base as db_base
mp = pytest.MonkeyPatch()
root = tmp_path_factory.mktemp("e2e")
origin = _seed_origin(root)
admin = _make_admin_clone(root, origin)
gh = _FakeGitHub(origin=origin, admin_clone=admin)
workspaces = root / "workspaces"
workspaces.mkdir()
url = make_url(_test_database_url)
mp.setattr(settings, "database_host", url.host or "localhost")
mp.setattr(settings, "database_port", url.port or 5432)
mp.setattr(settings, "database_user", url.username or "")
mp.setattr(settings, "database_password", url.password or "")
mp.setattr(settings, "database_name", url.database or "")
mp.setattr(settings, "workspaces_root", str(workspaces))
mp.setattr(settings, "workspace_auto_clone", True)
mp.setattr(settings, "encryption_key", Fernet.generate_key().decode())
# The app's lazy engine must bind to the patched settings, not a leftover.
db_base._DbHolder.engine = None
db_base._DbHolder.session_factory = None
port = _free_port()
base_url = f"http://127.0.0.1:{port}"
mp.setattr(settings, "github_api_base_url", f"{base_url}/_github")
app = _build_app(gh)
server = uvicorn.Server(
uvicorn.Config(app, host="127.0.0.1", port=port, log_level="warning")
)
thread = threading.Thread(target=server.run, daemon=True)
thread.start()
import httpx
deadline = time.time() + 30
while time.time() < deadline:
try:
# Any HTTP response at all means the server thread is up.
httpx.get(f"{base_url}/health", timeout=1)
break
except httpx.HTTPError:
time.sleep(0.1)
else:
raise RuntimeError("e2e app server did not become ready")
try:
yield E2EStack(
base_url=base_url,
root=root,
origin=origin,
workspaces_root=workspaces,
db_url=_test_database_url,
github=gh,
)
finally:
server.should_exit = True
thread.join(timeout=10)
db_base._DbHolder.engine = None
db_base._DbHolder.session_factory = None
mp.undo()
# ---------------------------------------------------------------------------
# Scripted agents — the REAL MCP tool functions, per-agent module reloads
# ---------------------------------------------------------------------------
class ScriptedAgent:
"""Drives the real flow/do MCP tool functions as one seeded agent."""
def __init__(self, stack: E2EStack, agent_id: UUID, slug: str, role: str) -> None:
self.stack = stack
self.agent_id = agent_id
self.slug = slug
self.role = role
self._manifest_path = stack.root / f"manifest-{slug}.json"
self._manifest_path.write_text(json.dumps(self._manifest()))
def _manifest(self) -> dict[str, Any]:
from roboco.services.gateway.role_config import get_role_config
cfg = get_role_config(self.role)
return {
"agent_id": str(self.agent_id),
"role": self.role,
"team": "backend",
"workspace_path": str(self.stack.workspaces_root),
"flow_tools": list(cfg.flow_tools),
"do_tools": list(cfg.do_tools),
"read_tools": ["Read", "Glob", "Grep"],
"write_tools": ["Edit", "Write"] if cfg.allows_write else [],
"bash_allowed": True,
"subagent_allowed": False,
"subagent_model": None,
"env": {},
}
def _module(self, name: str) -> ModuleType:
os.environ["ROBOCO_AGENT_ID"] = str(self.agent_id)
os.environ["ROBOCO_AGENT_ROLE"] = self.role
os.environ["ROBOCO_ORCHESTRATOR_URL"] = self.stack.base_url
os.environ["ROBOCO_TOOL_MANIFEST_PATH"] = str(self._manifest_path)
module = importlib.import_module(name)
if getattr(module, "AGENT_ID", None) != str(self.agent_id):
module = importlib.reload(module)
return module
def flow(self, verb: str, /, **kwargs: Any) -> dict[str, Any]:
result: dict[str, Any] = getattr(self._module("roboco.mcp.flow_server"), verb)(
**kwargs
)
return result
def do(self, tool: str, /, **kwargs: Any) -> dict[str, Any]:
result: dict[str, Any] = getattr(self._module("roboco.mcp.do_server"), tool)(
**kwargs
)
return result
def expect_error(env: dict[str, Any], kind: str, context: str) -> dict[str, Any]:
"""Assert an envelope is the EXPECTED rejection kind."""
assert env.get("error") == kind, (
f"{context}: expected rejection {kind!r}, got error={env.get('error')!r}\n"
f" full: {json.dumps(env, default=str, indent=2)[:4000]}"
)
return env
def expect_ok(env: dict[str, Any], context: str) -> dict[str, Any]:
"""Assert an envelope is a success; on failure show the whole envelope."""
assert isinstance(env, dict), f"{context}: non-dict envelope: {env!r}"
assert not env.get("error"), (
f"{context}: rejected with error={env.get('error')!r}\n"
f" message : {env.get('message')}\n"
f" remediate: {env.get('remediate')}\n"
f" missing : {env.get('missing')}\n"
f" full : {json.dumps(env, default=str, indent=2)[:4000]}"
)
return env
+369
View File
@@ -0,0 +1,369 @@
"""Scenario 1: a leaf dev task walks claim → work → PR → QA → docs → PM queue.
Every hop goes through the REAL MCP tool functions real HTTP real
gateway gates real services real git against the local origin, with a
fake GitHub REST layer whose merges are real git merges. No LLM: this file
IS the agent script, and every rejection envelope is printed verbatim so a
seam regression names itself.
"""
from __future__ import annotations
from typing import TYPE_CHECKING, Any
from uuid import uuid4
import pytest
from tests.e2e_smoke.harness import (
E2EStack,
ScriptedAgent,
expect_error,
expect_ok,
)
if TYPE_CHECKING:
from sqlalchemy.ext.asyncio import AsyncSession
pytestmark = pytest.mark.usefixtures("e2e_stack")
_PROJECT_SLUG = "e2e-proj"
class _Company:
dev_id: Any
qa_id: Any
doc_id: Any
cell_pm_id: Any
project_id: Any
task_id: Any
def _seed(stack: E2EStack) -> _Company:
from roboco.db.tables import AgentTable, ProjectTable, TaskTable
from roboco.models import AgentRole, AgentStatus, Team
from roboco.models.base import Complexity, TaskNature, TaskStatus, TaskType
from roboco.utils.crypto import encrypt_token
out = _Company()
async def _run(session: AsyncSession) -> None:
def agent(slug: str, role: AgentRole) -> AgentTable:
row = AgentTable(
id=uuid4(),
name=slug,
slug=slug,
role=role,
team=Team.BACKEND,
status=AgentStatus.ACTIVE,
model_config={},
system_prompt=slug,
capabilities=[],
permissions={},
metrics={},
)
session.add(row)
return row
dev = agent("be-dev-1", AgentRole.DEVELOPER)
qa = agent("be-qa", AgentRole.QA)
doc = agent("be-doc", AgentRole.DOCUMENTER)
pm = agent("be-pm", AgentRole.CELL_PM)
await session.flush()
project = ProjectTable(
id=uuid4(),
name="E2E Project",
slug=_PROJECT_SLUG,
git_url=str(stack.origin),
default_branch="master",
protected_branches=["master"],
assigned_cell=Team.BACKEND,
created_by=pm.id,
is_active=True,
git_token_encrypted=encrypt_token("e2e-dummy-token"),
)
session.add(project)
await session.flush()
task = TaskTable(
id=uuid4(),
title="Add the greeting module",
description=(
"Create greeting.txt with a friendly greeting so the smoke "
"harness has a real file change to commit, push, and merge."
),
acceptance_criteria=[
"greeting.txt exists at the repo root",
"its content greets the reader",
],
status=TaskStatus.PENDING,
priority=2,
task_type=TaskType.CODE,
nature=TaskNature.TECHNICAL,
estimated_complexity=Complexity.LOW,
project_id=project.id,
created_by=pm.id,
team=Team.BACKEND,
confirmed_by_human=True,
# The pool→agent routing lane is the orchestrator dispatcher's
# job (not under test here); a dev container is always spawned
# with its task already routed, which give_me_work serves via
# the pre-assigned-pending lane.
assigned_to=dev.id,
)
session.add(task)
await session.flush()
out.dev_id = dev.id
out.qa_id = qa.id
out.doc_id = doc.id
out.cell_pm_id = pm.id
out.project_id = project.id
out.task_id = task.id
stack.run_db(_run)
return out
def _task_state(stack: E2EStack, task_id: Any) -> dict[str, Any]:
from roboco.db.tables import TaskTable
from sqlalchemy import select
async def _run(session: AsyncSession) -> dict[str, Any]:
row = (
await session.execute(select(TaskTable).where(TaskTable.id == task_id))
).scalar_one()
return {
"status": str(row.status),
"branch_name": row.branch_name,
"pr_number": row.pr_number,
"docs_complete": row.docs_complete,
"assigned_to": row.assigned_to,
}
state: dict[str, Any] = stack.run_db(_run)
return state
def test_leaf_dev_task_reaches_pm_review(e2e_stack: E2EStack) -> None:
stack = e2e_stack
ids = _seed(stack)
task_id = str(ids.task_id)
# --- developer: discover, claim, work, PR, submit -----------------------
dev = ScriptedAgent(stack, ids.dev_id, "be-dev-1", "developer")
env = expect_ok(dev.flow("give_me_work"), "dev give_me_work")
assert env.get("task_id") == task_id, f"expected our task, got: {env}"
def _claim() -> dict:
return dev.flow(
"i_will_work_on",
task_id=task_id,
plan=(
"Create greeting.txt at the repository root containing a "
"friendly greeting, commit it on the task branch with the "
"task-prefixed message, push the branch to origin, open the "
"pull request against master, and self-verify both acceptance "
"criteria by re-reading the committed file content."
),
steps=[
{
"title": "Write greeting.txt",
"description": (
"Create greeting.txt at the repo root containing a "
"friendly greeting for the reader."
),
},
{
"title": "Commit and push",
"description": (
"Commit the new file on the task branch with a "
"task-prefixed message and push it to origin."
),
},
{
"title": "Open PR and self-verify",
"description": (
"Open the pull request against master and re-read the "
"file to confirm both acceptance criteria hold."
),
},
],
technical_considerations=["Plain text file; no build impact."],
risks=[
{
"risk": "None of substance — purely additive file.",
"mitigation": "Self-verify the file content before submit.",
}
],
open_questions=[],
)
# The composed claim succeeds and STAYS; the post-claim tracing gate
# then demands the claim-time journal note — the real agent choreography
# is claim → tracing_gap → note (now claim-held) → retry short-circuits.
expect_error(_claim(), "tracing_gap", "dev first i_will_work_on")
expect_ok(
dev.do(
"note",
scope="note",
task_id=task_id,
text=(
"Initial assessment: a single additive text file at the repo "
"root satisfies both acceptance criteria; no existing code is "
"touched, so risk is minimal and the plan is a three-step "
"write/commit/PR sequence."
),
),
"dev note at claim",
)
expect_ok(_claim(), "dev i_will_work_on retry")
state = _task_state(stack, ids.task_id)
assert state["status"] in ("claimed", "in_progress"), state
assert state["branch_name"], f"claim did not set a branch: {state}"
workspace = stack.workspace_of(_PROJECT_SLUG, "backend", "be-dev-1")
assert workspace.is_dir(), f"workspace clone missing at {workspace}"
# F123: the agent works in the per-task worktree, not the clone root.
workdir = workspace / ".worktrees" / task_id[:8]
assert workdir.is_dir(), f"per-task worktree missing at {workdir}"
(workdir / "greeting.txt").write_text("Hello from the e2e smoke agent!\n")
expect_ok(
dev.do(
"commit",
message="Add greeting.txt with a friendly greeting",
files=["greeting.txt"],
),
"dev commit",
)
expect_ok(
dev.do(
"note",
scope="note",
task_id=task_id,
text=(
"greeting.txt written and committed on the task branch; "
"opening the PR next, then self-verifying the acceptance "
"criteria before submit."
),
),
"dev progress note",
)
env = expect_ok(dev.flow("open_pr", task_id=task_id), "dev open_pr")
state = _task_state(stack, ids.task_id)
assert state["pr_number"], f"open_pr did not record a PR: {state} / {env}"
# The i_am_done tracing gate demands: a during-work journal entry, the
# dev_notes handoff section, a reflect entry, and an artifact referencing
# every acceptance criterion (quoted verbatim in the decision note).
expect_ok(
dev.do(
"note",
scope="decision",
task_id=task_id,
text=(
"Verified both acceptance criteria on the branch: "
'"greeting.txt exists at the repo root" holds (file committed '
'at the root), and "its content greets the reader" holds '
"(content is a friendly hello). Decision: no README change "
"needed; the greeting file is self-contained."
),
),
"dev during-work decision note",
)
expect_ok(
dev.do(
"note",
text="Handoff summary below (section carries the content).",
scope="handoff",
task_id=task_id,
section={
"summary": (
"Built the greeting module: greeting.txt added at the "
"repo root with a friendly greeting. Key change is one "
"additive file on the task branch; PR is open against "
"master; no risks beyond trivial content review."
)
},
),
"dev handoff section",
)
expect_ok(
dev.do(
"note",
scope="reflect",
task_id=task_id,
text=(
"Reflection: implemented the greeting task exactly per plan — "
"wrote the file, committed on the task branch, opened the PR, "
"and self-verified both acceptance criteria against the "
"committed content."
),
),
"dev reflect note",
)
expect_ok(dev.flow("i_am_done", task_id=task_id), "dev i_am_done")
assert _task_state(stack, ids.task_id)["status"] == "awaiting_qa"
# --- QA: claim the review, inspect, pass --------------------------------
qa = ScriptedAgent(stack, ids.qa_id, "be-qa", "qa")
expect_ok(qa.flow("claim_review", task_id=task_id), "qa claim_review")
expect_ok(
qa.do(
"note",
scope="learning",
task_id=task_id,
text=(
"Review learning: the greeting change is a single additive "
"file; diff inspection on the PR confirms both acceptance "
"criteria with no side effects on existing files."
),
),
"qa learning note",
)
expect_ok(
qa.flow(
"pass_review",
task_id=task_id,
notes=(
"Verified the PR diff on the fake origin: greeting.txt exists "
"at the repo root and greets the reader. Both acceptance "
"criteria hold; no regressions in the diff, and the branch "
"contains exactly the one additive commit described."
),
ac_verdicts=[
(
"greeting.txt exists at the repo root — verified in the "
"PR diff: the file is added at the repository root."
),
(
"its content greets the reader — verified: the committed "
"content is a friendly hello message."
),
],
),
"qa pass_review",
)
assert _task_state(stack, ids.task_id)["status"] == "awaiting_documentation"
# --- documenter: claim, document -----------------------------------------
doc = ScriptedAgent(stack, ids.doc_id, "be-doc", "documenter")
expect_ok(doc.flow("claim_doc_task", task_id=task_id), "doc claim_doc_task")
expect_ok(
doc.flow(
"i_documented",
task_id=task_id,
files=["greeting.txt"],
notes=(
"Documented the greeting module: greeting.txt carries the "
"user-facing greeting; no API surface changed, README "
"untouched by design."
),
),
"doc i_documented",
)
final = _task_state(stack, ids.task_id)
assert final["status"] == "awaiting_pm_review", final
assert final["docs_complete"] is True, final
@@ -76,6 +76,7 @@ async def test_skeleton_task_path_returns_incomplete_input() -> None:
status="in_progress", status="in_progress",
assigned_to=pm_id, assigned_to=pm_id,
priority=2, priority=2,
team="backend",
) )
task_svc = AsyncMock() task_svc = AsyncMock()
task_svc.get.return_value = parent task_svc.get.return_value = parent
+88 -1
View File
@@ -22,7 +22,7 @@ from roboco.api.routes.tasks import (
from roboco.api.routes.tasks import ( from roboco.api.routes.tasks import (
router as tasks_router, router as tasks_router,
) )
from roboco.db.tables import AgentTable, ProjectTable, TaskTable from roboco.db.tables import AgentTable, ProjectTable, TaskTable, WorkSessionTable
from roboco.exceptions import GitError, TaskLifecycleError from roboco.exceptions import GitError, TaskLifecycleError
from roboco.foundation.policy.lifecycle import STATUS_GRAPH from roboco.foundation.policy.lifecycle import STATUS_GRAPH
from roboco.foundation.policy.lifecycle import Status as LifecycleStatus from roboco.foundation.policy.lifecycle import Status as LifecycleStatus
@@ -389,6 +389,93 @@ async def test_update_task_resurrect_terminal_requires_force(task_client: dict)
assert with_force.json()["status"] == "in_progress" assert with_force.json()["status"] == "in_progress"
async def _seed_open_pr_session(setup: dict, task: TaskTable, pr_status: str) -> None:
"""Attach a work session with the given PR state to ``task``."""
ws = WorkSessionTable(
id=uuid4(),
project_id=setup["project"].id,
task_id=task.id,
agent_id=setup["agent"].id,
branch_name="feature/backend/ABC12345",
base_branch="master",
target_branch="master",
pr_number=123,
pr_url="https://example.com/r/pull/123",
pr_status=pr_status,
)
setup["db"].add(ws)
await setup["db"].flush()
task.work_session_id = ws.id
task.pr_number = 123
task.pr_url = ws.pr_url
await setup["db"].flush()
@pytest.mark.asyncio
async def test_admin_complete_with_open_pr_names_the_pr(task_client: dict) -> None:
"""Admin status→completed on a task whose PR is still OPEN strands its
commits (bit the CEO twice live, 2026-07-02). The refusal must name the
PR and the stranding not just the generic lifecycle-gate text."""
client = task_client["client"]
task = _seed_task(task_client, status=TaskStatus.AWAITING_CEO_APPROVAL)
await task_client["db"].flush()
await _seed_open_pr_session(task_client, task, "open")
response = await client.patch(
f"/api/tasks/{task.id}",
json={"status": "completed"},
headers=_HDR,
)
assert response.status_code == HTTPStatus.BAD_REQUEST
detail = response.json()["detail"]
assert "#123" in detail
assert "open" in detail.lower()
assert "force" in detail
@pytest.mark.asyncio
async def test_admin_complete_with_open_pr_force_still_escapes(
task_client: dict,
) -> None:
"""``force`` remains the deliberate, audited escape — an operator who
KNOWS the PR should be stranded can still complete."""
client = task_client["client"]
task = _seed_task(task_client, status=TaskStatus.AWAITING_CEO_APPROVAL)
await task_client["db"].flush()
await _seed_open_pr_session(task_client, task, "open")
response = await client.patch(
f"/api/tasks/{task.id}",
json={"status": "completed", "force": True},
headers=_HDR,
)
assert response.status_code == HTTPStatus.OK
assert response.json()["status"] == "completed"
@pytest.mark.asyncio
async def test_admin_complete_with_merged_pr_gets_generic_gate_only(
task_client: dict,
) -> None:
"""A merged PR strands nothing — the refusal stays the generic hatch
text (no PR callout), and force completes as before."""
client = task_client["client"]
task = _seed_task(task_client, status=TaskStatus.AWAITING_CEO_APPROVAL)
await task_client["db"].flush()
await _seed_open_pr_session(task_client, task, "merged")
no_force = await client.patch(
f"/api/tasks/{task.id}",
json={"status": "completed"},
headers=_HDR,
)
assert no_force.status_code == HTTPStatus.BAD_REQUEST
assert "#123" not in no_force.json()["detail"]
with_force = await client.patch(
f"/api/tasks/{task.id}",
json={"status": "completed", "force": True},
headers=_HDR,
)
assert with_force.status_code == HTTPStatus.OK
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_delete_task(task_client: dict) -> None: async def test_delete_task(task_client: dict) -> None:
client = task_client["client"] client = task_client["client"]
@@ -18,10 +18,28 @@ from __future__ import annotations
from unittest.mock import AsyncMock, MagicMock, patch from unittest.mock import AsyncMock, MagicMock, patch
from uuid import uuid4 from uuid import uuid4
import pytest
from roboco.config import settings as cfg
from roboco.models.events import EventType from roboco.models.events import EventType
from roboco.services.gateway.choreographer import Choreographer, ChoreographerDeps from roboco.services.gateway.choreographer import Choreographer, ChoreographerDeps
from structlog.testing import capture_logs from structlog.testing import capture_logs
@pytest.fixture(autouse=True)
def _unreachable_redis(monkeypatch: pytest.MonkeyPatch) -> None:
"""Point the tracker at an unreachable Redis for every test here.
The real RateLimitStateTracker otherwise wrote a NO-TTL "anthropic
rate-limited" state blob into a developer's live localhost Redis on
every run order/state-dependent poison for anything reading the real
tracker. activate() failing is fine: the parking handler catches and
logs it, and these tests assert orchestrator parking, not the write.
(redis_url is a computed property patch its inputs.)
"""
monkeypatch.setattr(cfg, "redis_host", "127.0.0.1")
monkeypatch.setattr(cfg, "redis_port", 1)
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
# Helpers # Helpers
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
-252
View File
@@ -1,252 +0,0 @@
"""Tests for stale-trigger cleanup + cooldown decisions."""
from __future__ import annotations
from datetime import UTC, datetime, timedelta
from unittest.mock import MagicMock
from uuid import UUID, uuid4
from roboco.services.gateway.trigger_filter import (
SpawnConfig,
SpawnDecision,
TriggerContext,
TriggerKind,
decide_spawn,
)
_DEFAULT_CONFIG = SpawnConfig(
cooldown_seconds=60,
role_rate_per_minute=6,
claim_stale_seconds=180,
)
def _task(
status: str,
active_claimant_id: UUID | None = None,
last_heartbeat_at: datetime | None = None,
) -> MagicMock:
t = MagicMock()
t.id = uuid4()
t.status = status
t.active_claimant_id = active_claimant_id
t.last_heartbeat_at = last_heartbeat_at
return t
def _trigger( # noqa: PLR0913
kind: TriggerKind,
skill: str | None = None,
recent_spawns_for_task: int = 0,
recent_spawns_for_role: int = 0,
provider: str | None = None,
provider_rate_limited: bool = False,
) -> TriggerContext:
return TriggerContext(
kind=kind,
skill=skill,
recent_spawns_for_task=recent_spawns_for_task,
recent_spawns_for_role=recent_spawns_for_role,
provider=provider,
provider_rate_limited=provider_rate_limited,
)
class TestStaleTriggerCleanup:
def test_a2a_code_review_for_completed_task_dropped(self) -> None:
t = _task(status="completed")
decision = decide_spawn(
task=t,
trigger=_trigger(TriggerKind.A2A, skill="code_review"),
config=_DEFAULT_CONFIG,
)
assert decision.outcome == SpawnDecision.DROP
assert "stale" in decision.reason.lower()
def test_a2a_code_review_for_paused_task_dropped(self) -> None:
"""Line 79-82: non-relevant non-terminal status (paused) → drop."""
t = _task(status="paused")
decision = decide_spawn(
task=t,
trigger=_trigger(TriggerKind.A2A, skill="code_review"),
config=_DEFAULT_CONFIG,
)
assert decision.outcome == SpawnDecision.DROP
assert "code_review" in decision.reason
def test_a2a_code_review_for_awaiting_qa_spawns(self) -> None:
t = _task(status="awaiting_qa")
decision = decide_spawn(
task=t,
trigger=_trigger(TriggerKind.A2A, skill="code_review"),
config=_DEFAULT_CONFIG,
)
assert decision.outcome == SpawnDecision.SPAWN
def test_notification_for_terminal_task_dropped(self) -> None:
t = _task(status="cancelled")
decision = decide_spawn(
task=t,
trigger=_trigger(TriggerKind.NOTIFICATION),
config=_DEFAULT_CONFIG,
)
assert decision.outcome == SpawnDecision.DROP
class TestSingleClaimantQueue:
def test_active_fresh_claimant_queues(self) -> None:
recent = datetime.now(tz=UTC)
t = _task(
status="in_progress",
active_claimant_id=uuid4(),
last_heartbeat_at=recent,
)
decision = decide_spawn(
task=t,
trigger=_trigger(TriggerKind.NOTIFICATION),
config=_DEFAULT_CONFIG,
)
assert decision.outcome == SpawnDecision.QUEUE
assert "claimant" in decision.reason.lower()
def test_stale_claimant_does_not_queue(self) -> None:
old = datetime.now(tz=UTC) - timedelta(seconds=600)
t = _task(
status="awaiting_qa",
active_claimant_id=uuid4(),
last_heartbeat_at=old,
)
decision = decide_spawn(
task=t,
trigger=_trigger(TriggerKind.A2A, skill="code_review"),
config=_DEFAULT_CONFIG,
)
assert decision.outcome == SpawnDecision.SPAWN
class TestCooldown:
def test_per_task_cooldown_queues(self) -> None:
t = _task(status="awaiting_qa")
decision = decide_spawn(
task=t,
trigger=_trigger(
TriggerKind.A2A,
skill="code_review",
recent_spawns_for_task=1,
),
config=_DEFAULT_CONFIG,
)
assert decision.outcome == SpawnDecision.QUEUE
assert "cooldown" in decision.reason.lower()
def test_role_rate_limit_queues(self) -> None:
t = _task(status="awaiting_qa")
decision = decide_spawn(
task=t,
trigger=_trigger(
TriggerKind.A2A,
skill="code_review",
recent_spawns_for_role=6,
),
config=_DEFAULT_CONFIG,
)
assert decision.outcome == SpawnDecision.QUEUE
assert "rate" in decision.reason.lower()
class TestProviderRateLimitGate:
"""Rule 2: provider rate-limit gate fires before claimant-lock/cooldown."""
def test_queues_when_provider_rate_limited(self) -> None:
"""QUEUE outcome when provider_rate_limited=True."""
t = _task(status="in_progress")
decision = decide_spawn(
task=t,
trigger=_trigger(
TriggerKind.NOTIFICATION,
provider="anthropic",
provider_rate_limited=True,
),
config=_DEFAULT_CONFIG,
)
assert decision.outcome == SpawnDecision.QUEUE
assert "provider anthropic rate-limited" in decision.reason
def test_reason_contains_provider_name(self) -> None:
"""Reason string must contain the provider name."""
t = _task(status="pending")
decision = decide_spawn(
task=t,
trigger=_trigger(
TriggerKind.SCAN,
provider="ollama_cloud",
provider_rate_limited=True,
),
config=_DEFAULT_CONFIG,
)
assert "ollama_cloud" in decision.reason
def test_reason_contains_unknown_when_no_provider_name(self) -> None:
"""When provider is None, reason still contains 'unknown'."""
t = _task(status="in_progress")
decision = decide_spawn(
task=t,
trigger=_trigger(
TriggerKind.NOTIFICATION,
provider=None,
provider_rate_limited=True,
),
config=_DEFAULT_CONFIG,
)
assert decision.outcome == SpawnDecision.QUEUE
assert "unknown" in decision.reason
def test_no_queue_injection_when_not_rate_limited(self) -> None:
"""SPAWN when provider_rate_limited=False and all other gates clear."""
t = _task(status="in_progress")
decision = decide_spawn(
task=t,
trigger=_trigger(
TriggerKind.NOTIFICATION,
provider="anthropic",
provider_rate_limited=False,
),
config=_DEFAULT_CONFIG,
)
assert decision.outcome == SpawnDecision.SPAWN
def test_stale_drop_fires_before_rate_limit_gate(self) -> None:
"""Rule 1 (stale-drop) fires before rule 2 (rate-limit gate)."""
t = _task(status="completed")
decision = decide_spawn(
task=t,
trigger=_trigger(
TriggerKind.NOTIFICATION,
provider="anthropic",
provider_rate_limited=True,
),
config=_DEFAULT_CONFIG,
)
# Rule 1 fires first — outcome must be DROP, not QUEUE
assert decision.outcome == SpawnDecision.DROP
def test_rate_limit_gate_fires_before_claimant_lock(self) -> None:
"""Rule 2 (rate-limit gate) fires before rule 3 (single-claimant invariant)."""
recent = datetime.now(tz=UTC)
t = _task(
status="in_progress",
active_claimant_id=uuid4(),
last_heartbeat_at=recent,
)
decision = decide_spawn(
task=t,
trigger=_trigger(
TriggerKind.NOTIFICATION,
provider="anthropic",
provider_rate_limited=True,
),
config=_DEFAULT_CONFIG,
)
# Both gates would QUEUE but reason must come from rate-limit (rule 2)
assert decision.outcome == SpawnDecision.QUEUE
assert "rate-limited" in decision.reason
@@ -60,7 +60,7 @@ async def test_auditor_spawns_on_task_blocked() -> None:
await handle_auditor_spawn(event) await handle_auditor_spawn(event)
spawn.assert_awaited_once_with(agent_id="auditor") spawn.assert_awaited_once_with(agent_id="auditor", spawned_by="event.auditor_spawn")
@pytest.mark.asyncio @pytest.mark.asyncio
@@ -75,7 +75,7 @@ async def test_auditor_spawns_on_task_cancelled() -> None:
await handle_auditor_spawn(event) await handle_auditor_spawn(event)
spawn.assert_awaited_once_with(agent_id="auditor") spawn.assert_awaited_once_with(agent_id="auditor", spawned_by="event.auditor_spawn")
@pytest.mark.asyncio @pytest.mark.asyncio
@@ -90,7 +90,7 @@ async def test_auditor_spawns_on_task_escalated_to_ceo() -> None:
await handle_auditor_spawn(event) await handle_auditor_spawn(event)
spawn.assert_awaited_once_with(agent_id="auditor") spawn.assert_awaited_once_with(agent_id="auditor", spawned_by="event.auditor_spawn")
@pytest.mark.asyncio @pytest.mark.asyncio
@@ -118,7 +118,7 @@ async def test_auditor_spawn_failure_does_not_propagate() -> None:
# Must not raise # Must not raise
await handle_auditor_spawn(event) await handle_auditor_spawn(event)
spawn.assert_awaited_once_with(agent_id="auditor") spawn.assert_awaited_once_with(agent_id="auditor", spawned_by="event.auditor_spawn")
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
@@ -1,95 +0,0 @@
"""Gateway spawn-cooldown gate.
`gateway_pre_spawn_check` always reaches
``roboco.services.gateway.trigger_filter.decide_spawn`` (except the no-task
carve-out), whose 4-rule cooldown machinery is the real spawn gate.
Without these assertions a regression that drops the call site would leave
the orchestrator with no server-side spawn cooldown beyond
``_pm_respawn_should_gate``.
"""
from __future__ import annotations
from unittest.mock import AsyncMock, MagicMock, patch
from uuid import uuid4
import pytest
from roboco.runtime.orchestrator import gateway_pre_spawn_check
from roboco.services.gateway.trigger_filter import Decision, SpawnDecision
@pytest.mark.asyncio
async def test_gateway_enabled_consults_decide_spawn() -> None:
"""``decide_spawn`` is invoked and its decision propagates."""
task_id = str(uuid4())
# Stub task row that decide_spawn will receive.
fake_task_row = MagicMock()
fake_task_row.status = "pending"
fake_task_row.active_claimant_id = None
fake_task_row.last_heartbeat_at = None
# Stub the async DB session: count queries return 0, task lookup returns
# our fake row. ``async with factory() as db`` -> ``db.execute(...)``.
fake_count_result = MagicMock()
fake_count_result.scalars.return_value.all.return_value = []
fake_task_result = MagicMock()
fake_task_result.scalars.return_value.first.return_value = fake_task_row
fake_db = AsyncMock()
fake_db.execute = AsyncMock(
side_effect=[fake_count_result, fake_count_result, fake_task_result]
)
fake_db.add = MagicMock()
fake_db.flush = AsyncMock()
fake_db.commit = AsyncMock()
fake_factory = MagicMock()
fake_factory.return_value.__aenter__ = AsyncMock(return_value=fake_db)
fake_factory.return_value.__aexit__ = AsyncMock(return_value=None)
expected = Decision(SpawnDecision.QUEUE, "per-task spawn cooldown active")
with (
patch("roboco.db.base.get_session_factory", return_value=fake_factory),
patch(
"roboco.services.gateway.trigger_filter.decide_spawn",
return_value=expected,
) as mock_decide_spawn,
):
outcome, reason = await gateway_pre_spawn_check(
task_id=task_id,
trigger_kind="scan",
target_role="developer",
)
mock_decide_spawn.assert_called_once()
call_kwargs = mock_decide_spawn.call_args.kwargs
assert call_kwargs["task"] is fake_task_row
assert call_kwargs["trigger"].kind.value == "scan"
assert call_kwargs["config"].cooldown_seconds > 0
assert outcome == "queue"
assert reason == "per-task spawn cooldown active"
@pytest.mark.asyncio
async def test_gateway_enabled_skips_decide_spawn_when_no_task_id() -> None:
"""No task_id -> early return; ``decide_spawn`` not called.
Documents the no-task-spawn carve-out: idle PM ticks pass through.
"""
with patch(
"roboco.services.gateway.trigger_filter.decide_spawn"
) as mock_decide_spawn:
outcome, reason = await gateway_pre_spawn_check(
task_id=None,
trigger_kind="scan",
target_role="main_pm",
)
assert outcome == "spawn"
assert "no task_id" in reason
mock_decide_spawn.assert_not_called()
@@ -0,0 +1,134 @@
"""Spawner attribution: every ``agent.spawned`` audit row names its dispatcher.
During the 2026-07-02 live run a rogue spawner could not be identified from the
audit log ``agent.spawned`` rows carry the container/model but not WHICH
dispatch loop launched them. ``spawn_agent`` now takes ``spawned_by`` and
stamps it into the ``agent.spawned`` / ``agent.spawn_failed`` details, and an
AST sweep holds every call site to passing it.
"""
from __future__ import annotations
import ast
import asyncio
from pathlib import Path
from typing import Any
from unittest.mock import AsyncMock
import pytest
from roboco.models.runtime import AgentInstance
from roboco.runtime.orchestrator import AgentConfig, AgentOrchestrator, AgentState
REPO_ROOT = Path(__file__).resolve().parents[3]
def _make_orchestrator(
monkeypatch: pytest.MonkeyPatch,
captured: list[dict[str, Any]],
container_result: Any,
) -> AgentOrchestrator:
orch = AgentOrchestrator.__new__(AgentOrchestrator)
orch._instances = {}
orch._lock = asyncio.Lock()
orch._bg_tasks = set()
orch._running = True
monkeypatch.setattr(orch, "_fire_audit", lambda **kw: captured.append(kw))
monkeypatch.setattr(orch, "_record_spawn_session", AsyncMock(return_value=None))
monkeypatch.setattr(orch, "_spawn_container", container_result)
return orch
def _config_and_instance() -> tuple[AgentConfig, AgentInstance]:
config = AgentConfig(
agent_id="be-dev-1",
blueprint_path=Path(),
model="opus",
provider_type="anthropic",
)
instance = AgentInstance(
agent_id="be-dev-1", state=AgentState.STARTING, config=config
)
return config, instance
@pytest.mark.asyncio
async def test_launch_spawn_audit_carries_spawned_by(
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""``agent.spawned`` details must name the dispatcher that launched it."""
captured: list[dict[str, Any]] = []
orch = _make_orchestrator(
monkeypatch, captured, AsyncMock(return_value="c0ffee" * 11)
)
config, instance = _config_and_instance()
await orch._launch_spawn(
"task-1", config, instance, None, None, spawned_by="_dispatch_qa_work"
)
spawned = [c for c in captured if c["event_type"] == "agent.spawned"]
assert len(spawned) == 1
assert spawned[0]["details"]["spawned_by"] == "_dispatch_qa_work"
@pytest.mark.asyncio
async def test_spawn_failed_audit_carries_spawned_by(
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""A failed launch must attribute the spawner too — a rogue dispatcher
that keeps crashing containers is exactly the live-debug case."""
captured: list[dict[str, Any]] = []
orch = _make_orchestrator(
monkeypatch, captured, AsyncMock(side_effect=RuntimeError("boom"))
)
config, instance = _config_and_instance()
with pytest.raises(RuntimeError):
await orch._launch_spawn(
"task-1", config, instance, None, None, spawned_by="_spawn_pending_dev"
)
failed = [c for c in captured if c["event_type"] == "agent.spawn_failed"]
assert len(failed) == 1
assert failed[0]["details"]["spawned_by"] == "_spawn_pending_dev"
@pytest.mark.asyncio
async def test_launch_spawn_without_attribution_stamps_unspecified(
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""The field is always present so audit queries never KeyError."""
captured: list[dict[str, Any]] = []
orch = _make_orchestrator(
monkeypatch, captured, AsyncMock(return_value="c0ffee" * 11)
)
config, instance = _config_and_instance()
await orch._launch_spawn("task-1", config, instance, None, None)
spawned = [c for c in captured if c["event_type"] == "agent.spawned"]
assert spawned[0]["details"]["spawned_by"] == "unspecified"
def _spawn_agent_calls_missing_spawned_by(path: Path) -> list[str]:
"""Return ``file:line`` for spawn_agent() calls without a spawned_by kwarg."""
tree = ast.parse(path.read_text())
missing: list[str] = []
for node in ast.walk(tree):
if not isinstance(node, ast.Call):
continue
func = node.func
if not (isinstance(func, ast.Attribute) and func.attr == "spawn_agent"):
continue
if not any(kw.arg == "spawned_by" for kw in node.keywords):
missing.append(f"{path.name}:{node.lineno}")
return missing
def test_every_spawn_agent_call_site_passes_spawned_by() -> None:
"""Sweep-guard over the whole package: a dispatcher added without
attribution fails here, not in a 3am live-debug session."""
missing: list[str] = []
for path in sorted((REPO_ROOT / "roboco").rglob("*.py")):
missing.extend(_spawn_agent_calls_missing_spawned_by(path))
assert not missing, f"spawn_agent() calls missing spawned_by=: {missing}"
@@ -121,6 +121,12 @@ def _enable(monkeypatch: pytest.MonkeyPatch, **overrides: object) -> None:
monkeypatch.setattr(cfg, key, value) monkeypatch.setattr(cfg, key, value)
# Keep notification a no-op (its own session/IO is out of scope here). # Keep notification a no-op (its own session/IO is out of scope here).
monkeypatch.setattr(NotificationService, "send_ack_notification", AsyncMock()) monkeypatch.setattr(NotificationService, "send_ack_notification", AsyncMock())
# Point the notify-dedupe at an unreachable Redis: _already_notified fails
# open and _mark_notified swallows, so these tests neither read from nor
# leak `self_heal:notified:*` keys (2h TTL) into a developer's live Redis.
# (redis_url is a computed property — patch its inputs.)
monkeypatch.setattr(cfg, "redis_host", "127.0.0.1")
monkeypatch.setattr(cfg, "redis_port", 1)
@pytest.mark.asyncio @pytest.mark.asyncio