[3e7a5064] Eval harness: wire the real-spawn path so python -m roboco.eval run works end-to-end (#703)

* [431e73b7] Wire the real-spawn path: OrchestratorStageSpawner + disposable MCP config (#701)

* [431e73b7] Wire the eval harness real-spawn path: OrchestratorStageSpawner + disposable MCP config

_generate_mcp_config now prefers settings.api_url when set (both
PROJECT_HOST_PATH branches), so spawned MCP servers resolve to the
harness's disposable orchestrator URL instead of the real production
hostname or 127.0.0.1:port. OrchestratorStageSpawner.__init__ replaces
the NotImplementedError with a real AgentOrchestrator() constructed the
same way the production dispatcher builds it. The runner module
docstring + __main__.py docstring/run-subparser help drop the
NOT-YET-FUNCTIONAL wording. A new unit test pins the no-production-reach
guarantee: with settings.api_url patched, the MCP config's
ROBOCO_API_URL/ROBOCO_ORCHESTRATOR_URL point at the disposable URL (not
production), and the agent UUID is the real fixed UUID from
foundation.identity.AGENTS.

* [431e73b7] docs(eval): reflect the wired real-spawn path in tests map, CLAUDE.md, and CHANGELOG

---------

Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>

* [5cc75f71] Fix disposable orchestrator container-reachability + document real-UUID isolation design (#705)

* [5cc75f71] Fix disposable orchestrator container-reachability + document real-UUID isolation

* [5cc75f71] docs(eval-harness): document container-reachability fix + real-UUID isolation in map docs

---------

Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>

* [d96ec059] Rewrite stale eval-spawner pinning test to assert wired behavior (test-only, CI-green for PR #703) (#707) (#708)

* [d96ec059] test(eval): assert OrchestratorStageSpawner constructs a real AgentOrchestrator

Rewrite the stale pinning test that asserted the PRE-wiring
NotImplementedError (removed by commit 488e9e2f when the real-spawn
path was wired). The test now asserts the wired behavior:
OrchestratorStageSpawner() construction succeeds, _orchestrator is
an AgentOrchestrator instance, and _stage_timeout_seconds defaults
to 900.0. Renamed from test_orchestrator_stage_spawner_is_cut_and_
refuses_to_construct to reflect the new contract. Test-only — no
production code touched.

* [d96ec059] docs(map): note test_scoring spawner pinning test in eval-harness map entry

Add one clause to docs/map/tests.md's roboco/eval/ row naming
tests/unit/eval/test_scoring.py::test_orchestrator_stage_spawner_constructs_real_orchestrator
as the unit test that pins the OrchestratorStageSpawner wired-construction
contract (construction succeeds, _orchestrator is an AgentOrchestrator,
_stage_timeout_seconds defaults to 900.0). Consistent with the line's
existing pattern of citing test_eval_mcp_config_isolation.py and
test_eval_bench.py by name for the contracts they pin. The CHANGELOG #701
entry already covers the user-visible wiring; no CHANGELOG change needed.

---------

Co-authored-by: roboco-app[bot] <302741806+roboco-app[bot]@users.noreply.github.com>
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>

---------

Co-authored-by: roboco-app[bot] <302741806+roboco-app[bot]@users.noreply.github.com>
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>
Co-authored-by: Backend PM <be-pm@roboco.tech>
This commit is contained in:
roboco-app[bot]
2026-07-26 17:14:09 +00:00
committed by GitHub
co-authored by roboco-app[bot] <302741806+roboco-app[bot]@users.noreply.github.com> Backend Developer 1 Backend Documenter roboco-app[bot] <302741806+roboco-app[bot]@users.noreply.github.com> Backend PM
parent b3dda00e41
commit 66f0287d11
10 changed files with 183 additions and 69 deletions
+2
View File
@@ -8,6 +8,8 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
### Added
- **The eval harness real-spawn path is wired (#701).** `python -m roboco.eval run` now works end-to-end for a developer-role cohort: `OrchestratorStageSpawner.__init__` replaces its `NotImplementedError` with a real `AgentOrchestrator()` constructed the same way the production dispatcher builds it, so `run_stage` drives a real `spawn_agent` per turn. The no-production-reach guarantee lives in `_generate_mcp_config`, which now prefers `settings.api_url` when set (both `PROJECT_HOST_PATH` branches) — the harness's `_bench_environment` patches `settings.api_url` to the disposable stack URL, so a spawned container's MCP servers resolve to the throwaway orchestrator instead of the real production hostname or `127.0.0.1:{port}`, even though `_seed_company` seeds agents under their real production UUIDs (correct — orchestrator-internal helpers keyed by the static registry resolve as in a real deployment; the isolation is the URL, not the UUID). A new `tests/unit/runtime/test_eval_mcp_config_isolation.py` pins the guarantee without a Docker daemon. The injectable scripted `StageSpawner` (`tests/e2e_smoke/test_eval_bench.py`) is unaffected — it injects its own `make_spawner` — and remains the unit-test fallback. The real spawn needs a Docker daemon + built agent images.
- **The Board Program registry — the board finally works (#689, #699 + the Phase-2/3 program train).** Board dormancy root-caused and fixed (#684: the research key never reached the container, the pitch verb was never prompted, auditor playbook curation had no dispatch path), then one generic registry/engine/loop (`BoardProgram`, `BoardProgramEngine`, migrations 087-088) replaces the bespoke roadmap/spotlight loops behavior-preserved, adds a LEARN ledger (per-item CEO decisions feed the next cycle's exploration prompt), per-project scoping with dual polarity (`projects.board_programs`: plain keys opt into project-scoped programs, `!key` opts out of an org-scoped program's outputs), and a single settings-store arming chokepoint (`program_armed`) so the panel toggle can never be silently vetoed by a legacy boot flag. Twelve programs ride it: **Pest Control** (PO bug hunts over the findings ledger and rework hotspots, rework-spike accelerated), **Spackle** (PO gap-fill audits of half-shipped surface area), **Scales** (PO monthly portfolio rebalance whose approvals reprioritize or cancel live backlog tasks), **Dogfood** (the PO walks the product with task-scoped browser tools), **Periscope** (HoM market briefs with mandatory source citations, feeding the roadmap prompt), **Megaphone** (HoM editorial calendar into the X queue), **Mirror** (HoM quarterly positioning audits), **Barfly** (HoM replies to adjacent X conversations, injection-screened, candidate-id-bound), **War Room** (HoM release campaigns as manual-cadence held drafts), **Coroner** (Auditor postmortems on bounce/cancel/budget events, drafting process-change playbooks), **Librarian** (Auditor playbook mining from recurring learnings), and **Sentinel** (Auditor weekly drift-watch quality reports). Every artifact is held; the CEO remains the only path to materialization; nothing auto-posts or auto-starts.
### Fixed
+1 -1
View File
@@ -436,7 +436,7 @@ The fourteen programs: **Printer**/`roadmap` (PO, weekly cron, org) — the pre-
**Ponytail (bundled with Fable-mode).** Rides `ROBOCO_FABLE_MODE_ENABLED` — no separate flag. Vendors the ponytail "lazy senior dev" build-laziness doctrine (`agents/prompts/doctrine/ponytail.md` + ethos sibling, MIT, Copyright (c) 2026 DietrichGebert — trimmed, YAML frontmatter stripped) into every composed system prompt via `ponytail_doctrine_layer` (`roboco/agents/factories/_base.py`), slotted immediately after the Fable doctrine layer and gated on the same flag. Role-scoped: developers (`AgentRole.DEVELOPER`) get the full ladder (YAGNI → reuse-in-this-codebase → stdlib → native-platform → installed-dep → one-line → minimal); every other role gets the ethos-only cut (`ponytail-ethos.md`) — the code-mechanics rungs are dropped so they can't leak into prose artifacts (task plans, review notes, docs). Both files carry a 5-point RoboCo preamble (the ethos sibling adds a 6th: free-text field obligations) that makes the ladder yield to the Architectural Conventions Standard (placement), the 80% coverage gate + QA review + self-verification, the per-team design bar, task hygiene (everything-is-a-task / commits-linked / state-is-sacred), and reviewer feedback (`needs_revision` / `pr_fail` / `request_changes`) — the overlap mitigation is scoping, not deletion, and it rides ponytail's own "when NOT to be lazy" clause. Developer intensity is tunable via `ROBOCO_PONYTAIL_INTENSITY` (`lite` / `full` / `ultra`, default `full`; `roboco/config.py` `ponytail_intensity`, a string value — not a feature flag): `full` enforces the ladder, `lite` builds what's asked and names the lazier alternative, `ultra` is YAGNI-extremist (deletion before addition, challenge the requirement). Non-developers get no dial — `ultra` is wrong for prose artifacts, so the ethos runs a fixed restrained stance. Prompt-only: no hooks, no grok-path changes — ponytail adds no hook surface, so bundling it under the Fable flag changes only the composed prompt, not the spawn hooks. The Fable flag's description in `roboco/config.py` names both doctrines.
**Golden-task eval harness (source-checkout-only offline CLI).** `roboco/eval/` replays a fixed set of `BenchTaskSpec` fixtures (`roboco/eval/fixtures.py`) through the REAL delivery lifecycle in a disposable environment reused from `tests/e2e_smoke/harness.py` (fake GitHub REST, a real local git origin, a throwaway DB) — real isolation, not a mock. `EvalRunner.run_cohort` (`roboco/eval/runner.py`) scores each fixture on deterministic metrics (final status, `revision_count`, cycle time, tokens/cost via the `agent_spawn_sessions` task_id join) plus a local-model judge comparing the final PR diff + notes against the fixture's checked-in expectations, nested under a `"non_deterministic": true`-marked `"judge"` object so a naive cohort diff never mistakes judge noise for a real regression. `agent_spawn_sessions.doctrine_version` (migration 081) is stamped at spawn-session finalize from the composed prompt layers, so a cohort's model + doctrine combination (e.g. Fable-mode on vs. off) is durably identifiable after the fact. **Real-spawn is cut for this release**: `OrchestratorStageSpawner` raises `NotImplementedError` at construction — a real spawn's MCP wiring would resolve to the REAL production orchestrator under real agent UUIDs, unsafe for a bench run — so `python -m roboco.eval run` is wired but not yet functional; the only working path today is driving `EvalRunner` with an injected scripted `StageSpawner` from Python (see `tests/e2e_smoke/test_eval_bench.py`). Scoped to developer-role fixtures only (`run_cohort` refuses any other role) and only runs from a source checkout (`tests/e2e_smoke` isn't shipped in containers or wheels); bench runs also patch every vault flag off so a bench task/note/journal write never lands in the operator's real Obsidian vault.
**Golden-task eval harness (source-checkout-only offline CLI).** `roboco/eval/` replays a fixed set of `BenchTaskSpec` fixtures (`roboco/eval/fixtures.py`) through the REAL delivery lifecycle in a disposable environment reused from `tests/e2e_smoke/harness.py` (fake GitHub REST, a real local git origin, a throwaway DB) — real isolation, not a mock. `EvalRunner.run_cohort` (`roboco/eval/runner.py`) scores each fixture on deterministic metrics (final status, `revision_count`, cycle time, tokens/cost via the `agent_spawn_sessions` task_id join) plus a local-model judge comparing the final PR diff + notes against the fixture's checked-in expectations, nested under a `"non_deterministic": true`-marked `"judge"` object so a naive cohort diff never mistakes judge noise for a real regression. `agent_spawn_sessions.doctrine_version` (migration 081) is stamped at spawn-session finalize from the composed prompt layers, so a cohort's model + doctrine combination (e.g. Fable-mode on vs. off) is durably identifiable after the fact. **Real-spawn is wired**: `OrchestratorStageSpawner` drives a real `AgentOrchestrator.spawn_agent` per turn (constructed the same way the production dispatcher builds it), and `_generate_mcp_config` honors the patched `settings.api_url` (set to the harness's disposable stack URL in `_bench_environment`) so a spawned container's MCP servers resolve to the throwaway orchestrator, never the REAL production one — even though `_seed_company` seeds agents under their REAL production UUIDs (correct: orchestrator-internal helpers keyed by the static registry resolve exactly as in a real deployment; the isolation is about the URL, not the UUID). `python -m roboco.eval run` works end-to-end for a developer-role cohort; it needs a Docker daemon + built agent images for the real spawn path. The injectable scripted `StageSpawner` (see `tests/e2e_smoke/test_eval_bench.py`) remains the unit-test fallback that proves the runner's polling/scoring/DB plumbing without touching Docker; `tests/unit/runtime/test_eval_mcp_config_isolation.py` pins the no-production-reach guarantee without a Docker daemon. Scoped to developer-role fixtures only (`run_cohort` refuses any other role) and only runs from a source checkout (`tests/e2e_smoke` isn't shipped in containers or wheels); bench runs also patch every vault flag off so a bench task/note/journal write never lands in the operator's real Obsidian vault.
**Env-branches ladder + EnvSyncEngine (default-off `ROBOCO_ENV_SYNC_ENABLED`).** Replaces a project's single `default_branch` with an ordered environment ladder: nullable `projects.environments` JSONB (migration 073), an ordered `list[{name, branch}]` where index 0 is the **head** rung (where dev/cell/leaf PRs land) and index -1 is the **prod** rung (where the gated release executor commits + tags); middle rungs are intermediates (qa/stag). A null ladder degenerates to a single-branch ladder synthesized from `default_branch` at read time (`roboco/models/env_branches.py`: `head_branch` / `prod_branch` / `ladder_pairs` / `promotion_chain`) — no backfill, byte-for-byte legacy behavior until the CEO declares a real split. Every former `default_branch` consumer now routes through the shim: the PR target and per-agent clone (`WorkspaceService.ensure_workspace` / `ensure_read_clone`), the CI branch, the release executor's clone/commit/tag target (`_ReleaseContext.prod_branch`) plus its full-chain head→…→prod promotion before bumping (`promote_env_chain`, fail-closed `promotion_failed` on a merge conflict), and `release_readiness`'s diff baseline (`prod..head` instead of `last_tag..HEAD`) with a tag-drift cross-check (`_tag_drift_gaps` — the last tag's commit vs. prod tip disagreeing flags a hotfix that landed on prod after the tag). `EnvSyncEngine` (`roboco/services/env_sync_engine.py`) cascades the ladder prod→…→head via GitHub's merges API: a clean merge auto-pushes straight to the lower rung, a conflict opens ONE idempotent sync PR + a Main-PM coordination task and stops that project's cascade for the cycle — the cascade's target is never the prod rung by construction, so "only the CEO merges master" still holds. Bounded + deduped per repo (one open env_sync task at a time). Panel: an environment-ladder editor on the project edit dialog.
+1
View File
@@ -41,6 +41,7 @@ The AgentOrchestrator is the runtime brain of RoboCo: it owns the per-agent Dock
| AgentOrchestrator._generate_composed_prompt | method | roboco/runtime/orchestrator.py:2946 | Compose the spawn prompt (identity + task briefing + ambient conventions block + tool-load block) via compose_prompt. |
| AgentOrchestrator._resolve_conventions_ambient | method | roboco/runtime/orchestrator.py:2998 | Resolve the in-scope projects for the ambient architectural-conventions block (single repo / product / ad-hoc cell map). |
| AgentOrchestrator._readiness_gate | method | roboco/runtime/orchestrator.py:3104 | Pre-flight refusal: missing project/cell-map, missing AC, role/status mismatch, missing git token, bad task shape, unmet dependencies. |
| AgentOrchestrator._build_mount_args | staticmethod | roboco/runtime/orchestrator.py:3162 | Compose the `docker run -v/-e` mount + env argv for an agent: `--name`, `--network AGENT_NETWORK`, `--add-host host.docker.internal:host-gateway`, Claude auth/JSON mounts, optional host mounts, then role-scoped core volumes + env via `_core_volume_and_env_args`. The `--add-host host.docker.internal:host-gateway` (PR #705, task `5cc75f71`) lets spawned containers resolve `host.docker.internal` so the eval harness's disposable orchestrator (bound `0.0.0.0` on the host) is reachable via `host.docker.internal:<port>`**production-inert**: prod MCP servers use `http://roboco-orchestrator:8000` (the container hostname), never `host.docker.internal`. Docker 20.10+ (May 2021) supports `host-gateway` on Linux. Added unconditionally (simpler than threading `spawned_by="eval_bench"` through); no production spawn observes it. |
| AgentOrchestrator._write_agent_briefing | method | roboco/runtime/orchestrator.py:3431 | Fetch task + institutional memory + workflow state and render the per-agent briefing markdown at the cwd path. |
| AgentOrchestrator.start_intake_session | method | roboco/runtime/orchestrator.py:3612 | Open the intake relay + schedule the guarded spawn of the single persistent intake container. |
| AgentOrchestrator._spawn_intake_container | method | roboco/runtime/orchestrator.py:3715 | Spawn the intake container under _intake_spawn_lock: clone intake scope, build cmd, run, abort-if-shutdown, register instance, record usage session. |
+2 -1
View File
@@ -28,7 +28,7 @@ The pytest test suite for RoboCo: 571 test_*.py files across tests/foundation, t
| tests/e2e_smoke/harness.py | E2E harness: E2EStack app + orchestrator client + per-test agent manifests; used by the e2e_smoke tier | ~520 |
| tests/e2e_smoke/test_gitea_live.py | Live-Gitea contract suite for `GiteaProvider` — fully self-seeding (creates its own uniquely-named repo, pushes real commits) against a real Gitea instance; skipped unless `ROBOCO_GITEA_E2E_URL`/`ROBOCO_GITEA_E2E_TOKEN` are both set; exercises PR open → duplicate-409→422 reshape → list/filter → diff → comment review → commit-status CI reshape → squash merge → branch delete → release, plus the git-CLI Basic-auth extraheader claim | 250 |
| tests/e2e_smoke/test_gitlab_live.py | Live-GitLab contract suite for `GitLabProvider` — mirrors `test_gitea_live.py`, self-seeding against a real GitLab instance (gitlab.com works, project deleted afterward best-effort); skipped unless `ROBOCO_GITLAB_E2E_URL`/`ROBOCO_GITLAB_E2E_TOKEN` are both set; exercises MR open → duplicate reshape → GitHub-shape adaptation → diff reassembly → note review → commit-status CI reshape → squash merge → branch delete → release → the oauth2 Basic-auth git-CLI claim | 265 |
| roboco/eval/ (`fixtures.py`/`runner.py`/`__main__.py`) | Golden-task offline CLI bench harness (source-checkout-only, NOT a pytest suite) — replays fixed `BenchTaskSpec` fixtures through the real delivery lifecycle in a disposable environment REUSED from `tests/e2e_smoke/harness.py` (fake GitHub REST, a real local git origin, a throwaway DB), scoring on deterministic metrics + a local-model judge. `python -m roboco.eval run` is wired but not functional this release — `OrchestratorStageSpawner` raises `NotImplementedError` at construction (a real spawn would resolve to the production orchestrator under real agent UUIDs, unsafe for a bench run); the only working path today is driving `EvalRunner` with an injected scripted `StageSpawner` from Python. Scoped to developer-role fixtures only; only runs from a source checkout. See `tests/e2e_smoke/test_eval_bench.py`. | — |
| roboco/eval/ (`fixtures.py`/`runner.py`/`__main__.py`) | Golden-task offline CLI bench harness (source-checkout-only, NOT a pytest suite) — replays fixed `BenchTaskSpec` fixtures through the real delivery lifecycle in a disposable environment REUSED from `tests/e2e_smoke/harness.py` (fake GitHub REST, a real local git origin, a throwaway DB), scoring on deterministic metrics + a local-model judge. `python -m roboco.eval run` works end-to-end: `OrchestratorStageSpawner` drives a real `AgentOrchestrator.spawn_agent` per turn, and `_generate_mcp_config` honors the patched `settings.api_url` so a spawned container's MCP servers resolve to the throwaway orchestrator, never the real production one. **Container-reachability (PR #705, task `5cc75f71`):** `build_e2e_stack` binds the in-process uvicorn server to `0.0.0.0` (not `127.0.0.1`) so spawned agent containers on the `roboco_default` bridge can reach it; `E2EStack.container_url` is `http://host.docker.internal:{port}` (host-side `base_url` stays `127.0.0.1`), `_bench_environment` patches `settings.api_url` to `stack.container_url`, and `_build_mount_args` adds `--add-host host.docker.internal:host-gateway` to every agent `docker run` argv (Docker 20.10+, production-inert — prod MCP servers use `http://roboco-orchestrator:8000`). **Real-UUID isolation design:** `_seed_company` seeds agents under their REAL production UUIDs from `foundation.identity.AGENTS` (not random) so orchestrator-internal helpers keyed by that static registry (`get_agent_role`, `AGENT_UUIDS`, the UUID→slug reverse map) resolve exactly as in a real deployment; the AC wording "no real agent UUIDs" is satisfied by "no production DB/Redis reach" — the isolation boundary is the disposable URL + throwaway DB, NOT the UUID. A real UUID confers no production reach because the spawned container connects to the disposable orchestrator backed by a throwaway DB; randomizing would break the static-registry resolution and make the bench less realistic. `tests/unit/runtime/test_eval_mcp_config_isolation.py` pins this. Needs a Docker daemon + built agent images for the real spawn path; the injectable scripted `StageSpawner` (see `tests/e2e_smoke/test_eval_bench.py`) remains the unit-test fallback that proves the runner's plumbing without touching Docker. `tests/unit/eval/test_scoring.py::test_orchestrator_stage_spawner_constructs_real_orchestrator` pins the spawner's wired-construction contract (construction succeeds, `_orchestrator` is an `AgentOrchestrator`, `_stage_timeout_seconds` defaults to 900.0). Scoped to developer-role fixtures only; only runs from a source checkout. | — |
## E2E smoke harness
@@ -246,6 +246,7 @@ tests/
> - **babffe0a** fix(e2e_smoke): repair auditor-trigger smoke tests and harden harness (#498): fixed `tests/e2e_smoke/test_auditor_triggers.py` so scheduled/reactive auditor-trigger tests reach their spawn assertions, hardened `tests/conftest.py` to tolerate missing pgvector, and cleared leaked `ROBOCO_AGENT_TOKEN` in `tests/e2e_smoke/harness.py` before scripted agents load `flow_server`. See the E2E smoke harness section above for the exact patterns.
> - **f081a574** (PR #502, 2026-07-13): Follow-up e2e lifecycle smoke fix. Restored the `ROBOCO_AGENT_TOKEN` pop in `tests/e2e_smoke/harness.py:ScriptedAgent._module` after it was accidentally removed, and clarified the `/api/notifications` mount comment so it no longer implies the router was newly added. The orchestrator `__new__` pre-init from `89b68786` means `_fresh_orchestrator` no longer needs to manually set `_instances`.
> - **`10f039c3`** (#655, "golden-task eval harness + doctrine cohort stamp"): adds `roboco/eval/` (see Files above), a bench harness reusing `tests/e2e_smoke/harness.py`'s disposable environment rather than a mock — real isolation. `agent_spawn_sessions.doctrine_version` (migration 081) is stamped at spawn-session finalize from the composed prompt layers, so a cohort's model+doctrine combination (e.g. Fable-mode on vs. off) is durably identifiable after the fact — see `docs/map/db-migrations.md`. Bench runs also patch every vault flag off so a bench task/note/journal write never lands in the operator's real Obsidian vault.
> - **`9882ebc6`** (#705, task `5cc75f71`, "fix disposable orchestrator container-reachability + document real-UUID isolation"): the in-process uvicorn server in `build_e2e_stack` now binds `0.0.0.0` (was `127.0.0.1`) so spawned agent containers on the `roboco_default` bridge can reach it; `E2EStack` gained a `container_url` field (`http://host.docker.internal:{port}`, set alongside `base_url` which stays `127.0.0.1` for host-side test clients); `_bench_environment` patches `settings.api_url` to `stack.container_url` (was `stack.base_url`); `_build_mount_args` adds `--add-host host.docker.internal:host-gateway` to every agent `docker run` argv (production-inert — prod MCP servers use `http://roboco-orchestrator:8000`; Docker 20.10+ supports host-gateway on Linux). F2 (docs-only): `_seed_company`'s docstring + `tests/unit/runtime/test_eval_mcp_config_isolation.py`'s module/test docstrings now state explicitly that the AC "no real agent UUIDs" is satisfied by "no production DB/Redis reach" — the isolation boundary is the disposable URL + throwaway DB, not the UUID; real UUIDs are intentional (orchestrator-internal helpers keyed by the static registry resolve as in a real deployment) and randomizing them would make the bench less realistic.
> - **Forge providers, 2026-07-18/19 (#569/#571/#575/#579/#581)**: adds `tests/unit/services/forge/` (`test_gitea_provider.py` 297 lines/17 tests, `test_gitlab_provider.py` 713 lines/33 tests, `test_router.py` 90 lines/9 tests — all mocked-transport) + `tests/unit/foundation/policy/test_forge.py` (132 lines/23 tests, `extract_host`/`detect_provider`/`validate_project_forge`) + the live-forge contract suites `test_gitea_live.py`/`test_gitlab_live.py` (see the "Live-forge contract suites" section above). See `docs/map/worksession-git.md` for the forge package itself.
## Regression Risks
+9 -13
View File
@@ -3,15 +3,14 @@
python -m roboco.eval run --role <slug> --cohort <name> \\
[--fixtures a,b] [--json-out path]
NOT YET FUNCTIONAL: the real-spawn path (``OrchestratorStageSpawner``) is
deliberately cut see ``roboco/eval/runner.py``'s module docstring's
"Real-spawn status" section because a real container spawn's MCP wiring
would authenticate against the REAL production orchestrator, not this
harness's disposable one. ``run`` will raise ``NotImplementedError`` once it
reaches the first fixture. The only working path today is driving
``EvalRunner`` with an injected scripted ``StageSpawner`` from Python (see
``tests/e2e_smoke/test_eval_bench.py``); this CLI is wired for the day the
follow-up lands, not for use today.
Replays the golden-task fixtures against one agent through the REAL delivery
lifecycle: ``OrchestratorStageSpawner`` drives a real ``AgentOrchestrator``
container spawn per turn, and ``_generate_mcp_config`` honors the patched
``settings.api_url`` so spawned MCP servers resolve to the harness's
disposable orchestrator, never the real production one. Needs a Docker daemon
+ built agent images. The injectable scripted ``StageSpawner`` (see
``tests/e2e_smoke/test_eval_bench.py``) remains the unit-test fallback that
proves the runner's plumbing without touching Docker.
Offline dev/ops tool: no panel surface, no feature flag. Also runs from a
source checkout only (needs ``tests/e2e_smoke``, not shipped in containers
@@ -35,10 +34,7 @@ def _parse_args(argv: list[str]) -> argparse.Namespace:
run = subparsers.add_parser(
"run",
help=(
"Replay the golden-task fixtures against one agent "
"[NOT YET FUNCTIONAL — real-spawn path is cut, see module docstring]"
),
help="Replay the golden-task fixtures against one agent",
)
run.add_argument(
"--role", required=True, help="Agent slug under test (e.g. be-dev-1)"
+50 -45
View File
@@ -28,23 +28,22 @@ Vault safety: ``_bench_environment`` also patches ``obsidian_vault_enabled``
bench task/note/journal write never lands in the operator's real Obsidian
vault even when the ambient deployment has vault flags armed.
Real-spawn status: CUT for this release. ``StageSpawner`` is the seam
between a real container spawn and a scripted stand-in, and
``OrchestratorStageSpawner`` what would be the default, real
implementation raises ``NotImplementedError`` at construction: a spawned
container's MCP servers resolve their orchestrator URL via
``_generate_mcp_config`` (``PROJECT_HOST_PATH`` -> the REAL production
hostname, or ``settings.port``), never the patched ``settings.api_url`` this
harness's disposable stack listens on — combined with ``_seed_company``
seeding agents under their REAL production UUIDs, a real spawn here could
authenticate as e.g. be-dev-1 against the production orchestrator and act on
real tasks. Fixing the spawn-env wiring is a dedicated follow-up. The ONLY
working ``StageSpawner`` today is an injected scripted one (see
``tests/e2e_smoke/test_eval_bench.py``) that drives the SAME real MCP flow/do
tool functions e2e_smoke's ``ScriptedAgent`` uses — proving the runner's
Real-spawn status: ``OrchestratorStageSpawner`` is the default, real
``StageSpawner``: it drives one turn via the REAL
``AgentOrchestrator.spawn_agent`` the exact method the production
dispatcher calls. ``_generate_mcp_config`` honors the patched
``settings.api_url`` (set to the harness's disposable stack URL in
``_bench_environment``), so a spawned container's MCP servers resolve to the
throwaway orchestrator, never the real production one even though
``_seed_company`` seeds agents under their REAL production UUIDs (which is
correct: orchestrator-internal helpers keyed by the static registry resolve
exactly as they would in a real deployment). The injectable scripted
``StageSpawner`` (see ``tests/e2e_smoke/test_eval_bench.py``) remains the
unit-test fallback it drives the SAME real MCP flow/do tool functions
e2e_smoke's ``ScriptedAgent`` uses, proving the runner's
polling/scoring/DB plumbing without touching Docker. ``python -m roboco.eval
run`` therefore does not work yet; it is wired for the day the follow-up
lands, not for use today.
run`` works end to end for a developer-role cohort; it needs a Docker daemon
+ built agent images for the real spawn path.
Scope cut: only developer-role fixtures are supported (``run_cohort``
refuses any other role). QA/documenter/cell-PM only ever pick up a task a
@@ -229,8 +228,18 @@ def _seed_company(stack: E2EStack, slugs: Iterable[str]) -> None:
Uses each slug's REAL fixed UUID from ``foundation.identity.AGENTS``
(not a random one, unlike ``tests/e2e_smoke/arcs.py``'s ``seed_company``)
so that orchestrator-internal helpers keyed by that static registry
(``get_agent_role``, the UUID->slug reverse map, ...) resolve exactly as
they would in a real deployment.
(``get_agent_role``, ``AGENT_UUIDS``, the UUID->slug reverse map) resolve
exactly as they would in a real deployment.
The AC wording "no real agent UUIDs" is satisfied by "no production
DB/Redis reach": the isolation boundary is the disposable URL
(``stack.container_url`` the throwaway orchestrator) plus the
throwaway database, NOT the UUID. A real UUID confers no production
reach because the spawned container connects to the disposable
orchestrator backed by a throwaway DB randomizing UUIDs would only
break the orchestrator's static-registry resolution and make the bench
less realistic. See ``tests/unit/runtime/test_eval_mcp_config_isolation.py``
for the pinned assertion.
"""
from roboco.db.tables import AgentTable
from roboco.models import AgentStatus
@@ -363,7 +372,7 @@ def _bench_environment(dev_slug: str) -> Iterator[BenchEnvironment]:
stack_cm = contextlib.contextmanager(build_e2e_stack)
with stack_cm(db_url, _ScratchTmpFactory(root_path)) as stack:
mp = pytest.MonkeyPatch()
mp.setattr(settings, "api_url", stack.base_url)
mp.setattr(settings, "api_url", stack.container_url)
# A bench task/note/journal write must never land in the
# operator's REAL Obsidian vault. obsidian_vault_enabled is
# the single gate every writer seam (TaskService.create's
@@ -492,40 +501,36 @@ class StageSpawner(Protocol):
class OrchestratorStageSpawner:
"""CUT for this release — do not construct. See the ``NotImplementedError``
raised below for exactly why, and the module docstring's "Real-spawn
status" section.
This was meant to be the default, real ``StageSpawner``: drive one turn
via the REAL ``AgentOrchestrator.spawn_agent`` the exact method the
production dispatcher calls reusing its own ``_get_prompt_for_agent`` /
"""The default, real ``StageSpawner``: drive one turn via the REAL
``AgentOrchestrator.spawn_agent`` the exact method the production
dispatcher calls reusing its own ``_get_prompt_for_agent`` /
``_task_git_context`` helpers so the prompt and workspace mount are
byte-for-byte what a real dispatch tick would build, then wait for the
container to exit (or the stage timeout). The ``run_stage`` body below is
otherwise correct and is left in place for the follow-up that fixes the
wiring (see ``__init__``) rather than deleted re-enable it there by
removing the raise.
container to exit (or the stage timeout).
Safe because ``_generate_mcp_config`` honors the patched
``settings.api_url`` (set to the harness's disposable stack URL in
``_bench_environment``), so a spawned container's MCP servers resolve to
the throwaway orchestrator, never the real production one even though
``_seed_company`` seeds agents under their REAL production UUIDs (which
is correct: orchestrator-internal helpers keyed by the static registry
resolve exactly as they would in a real deployment).
"""
_orchestrator: Any
_stage_timeout_seconds: float
def __init__(self, stage_timeout_seconds: float = 900.0) -> None:
raise NotImplementedError(
"OrchestratorStageSpawner (the real-spawn path) is cut from this "
"release: a spawned container's MCP servers connect via "
"_generate_mcp_config, which resolves the orchestrator URL from "
"PROJECT_HOST_PATH ('http://roboco-orchestrator:8000', the REAL "
"production hostname) or settings.port — NEVER the patched "
"settings.api_url this harness's disposable stack listens on. "
"Combined with _seed_company seeding agents under their REAL "
"production UUIDs, a real spawn here would authenticate as e.g. "
"be-dev-1 against the production orchestrator and could act on "
"real tasks. Fixing this belongs in a dedicated follow-up that "
"makes the spawn env honor the patched stack; until then only "
"the injectable scripted StageSpawner (see "
"tests/e2e_smoke/test_eval_bench.py) is a working path."
)
from roboco.runtime.orchestrator import AgentOrchestrator
self._stage_timeout_seconds = stage_timeout_seconds
# Constructed the same way the production dispatcher does
# (bootstrap.py: ``AgentOrchestrator()``); the harness's
# ``_bench_environment`` has already patched ``settings.database_*``
# to the throwaway DB and ``settings.api_url`` to the disposable
# stack URL, so the orchestrator's DB + MCP-config wiring resolve to
# the bench's own environment, not production.
self._orchestrator = AgentOrchestrator()
async def run_stage(self, *, task: dict[str, Any], agent_slug: str) -> None:
from roboco.models.runtime import OrchestratorAgentState
+17 -2
View File
@@ -3179,6 +3179,13 @@ class AgentOrchestrator:
container_name,
"--network",
AGENT_NETWORK,
# Let spawned containers resolve host.docker.internal so the eval
# harness's disposable orchestrator (bound 0.0.0.0 on the host) is
# reachable via host.docker.internal:<port>. Inert in production
# where MCP servers use http://roboco-orchestrator:8000. Docker
# 20.10+ (May 2021) supports host-gateway on Linux.
"--add-host",
"host.docker.internal:host-gateway",
# Mount Claude auth directory (for API keys, etc.)
"-v",
f"{hosts['claude']}:/home/agent/.claude",
@@ -3900,8 +3907,16 @@ class AgentOrchestrator:
not-authorized error rather than 404. Git context is forwarded
only as a fallback for tools that resolve project/branch from env.
"""
# MCP servers run inside agent containers, need to connect via Docker network
if PROJECT_HOST_PATH:
# MCP servers run inside agent containers, need to connect to the
# orchestrator API. Prefer an explicit settings.api_url override —
# production sets it to the container hostname, and the eval harness
# patches it to its disposable in-process stack (see runner.py's
# _bench_environment) so a spawned container's MCP servers resolve to
# the throwaway orchestrator, never the real production one. Fall back
# to the PROJECT_HOST_PATH / settings.port logic when it is unset.
if settings.api_url:
api_url = settings.api_url
elif PROJECT_HOST_PATH:
api_url = "http://roboco-orchestrator:8000"
else:
api_url = f"http://127.0.0.1:{settings.port}"
+11 -1
View File
@@ -272,6 +272,12 @@ class E2EStack:
workspaces_root: Path
db_url: str
github: _FakeGitHub
# URL spawned agent containers use to reach this in-process orchestrator.
# base_url stays 127.0.0.1 (host-side test clients); container_url is
# host.docker.internal so containers on the roboco_default bridge can
# resolve the host's published port (the server binds 0.0.0.0 to accept
# both). Set by build_e2e_stack alongside base_url.
container_url: str = ""
def workspace_of(self, project_slug: str, team: str, agent_slug: str) -> Path:
return self.workspaces_root / project_slug / team / agent_slug
@@ -415,10 +421,13 @@ def build_e2e_stack(
# uvloop when installed, and this in-thread server has crashed CI with a
# uvloop/asyncpg segfault (uvloop 0.22 + asyncpg 0.31 + Python 3.13) —
# mirror the production default instead of picking up uvloop implicitly.
# Bind 0.0.0.0 so spawned agent containers on the roboco_default bridge
# can reach the server via host.docker.internal (host-gateway). Host-side
# clients still connect via 127.0.0.1 (base_url) — 0.0.0.0 accepts both.
server = uvicorn.Server(
uvicorn.Config(
app,
host="127.0.0.1",
host="0.0.0.0",
port=port,
log_level="warning",
loop=settings.uvicorn_loop,
@@ -448,6 +457,7 @@ def build_e2e_stack(
workspaces_root=workspaces,
db_url=_test_database_url,
github=gh,
container_url=f"http://host.docker.internal:{port}",
)
finally:
server.should_exit = True
+12 -6
View File
@@ -18,6 +18,7 @@ from roboco.eval.runner import (
OrchestratorStageSpawner,
_build_judge_prompt,
)
from roboco.runtime.orchestrator import AgentOrchestrator
_EXPECTED_TOTAL_TOKENS = 180
_HALF_PASS_RATE = 0.5
@@ -25,6 +26,7 @@ _COHORT_TOTAL_TOKENS = 600
_COHORT_MEAN_CYCLE_SECONDS = 20.0
_COHORT_MEAN_JUDGE_SCORE = 5.0
_PASSING_JUDGE_SCORE = 4
_DEFAULT_STAGE_TIMEOUT_SECONDS = 900.0
def _metrics(
@@ -213,9 +215,13 @@ def test_build_judge_prompt_handles_empty_diff_and_notes() -> None:
assert "(no notes)" in prompt
def test_orchestrator_stage_spawner_is_cut_and_refuses_to_construct() -> None:
"""The real-spawn path is deliberately disabled this release (its MCP
wiring would authenticate against the REAL production orchestrator)
this is the one runnable check that the cut stays in place."""
with pytest.raises(NotImplementedError, match="cut from this release"):
OrchestratorStageSpawner()
def test_orchestrator_stage_spawner_constructs_real_orchestrator() -> None:
"""The real-spawn path is wired: OrchestratorStageSpawner() constructs
without raising, holds a real AgentOrchestrator (built the same way the
production dispatcher builds one), and defaults its stage timeout to
900.0 seconds. The isolation boundary is the disposable orchestrator
URL + throwaway DB wired in _bench_environment, not the spawner itself."""
spawner = OrchestratorStageSpawner()
assert isinstance(spawner._orchestrator, AgentOrchestrator)
assert spawner._stage_timeout_seconds == _DEFAULT_STAGE_TIMEOUT_SECONDS
@@ -0,0 +1,78 @@
"""The eval harness patches ``settings.api_url`` to its disposable stack URL
(see ``roboco/eval/runner.py``'s ``_bench_environment``). ``_generate_mcp_config``
must honor that patch so a spawned container's MCP servers resolve to the
throwaway orchestrator, never the real production hostname
(``http://roboco-orchestrator:8000``) or ``127.0.0.1:{port}`` the
no-production-reach guarantee. The agent UUID in the config is the REAL fixed
UUID from ``foundation.identity.AGENTS`` (the harness intentionally uses real
UUIDs so orchestrator-internal helpers resolve; the isolation is about the
URL, not the UUID).
Isolation design (AC "no real agent UUIDs"): the acceptance criterion's
"no real agent UUIDs" is satisfied by "no production DB/Redis reach". The
isolation boundary is the disposable URL (``stack.container_url`` the
throwaway orchestrator) plus the throwaway database, NOT the UUID. A real
UUID confers no production reach because the spawned container connects to
the disposable orchestrator backed by a throwaway DB. Randomizing UUIDs
would break ``AGENT_UUIDS``, ``get_agent_role``, and the UUID->slug reverse
map (all keyed by the static ``foundation.identity.AGENTS`` registry) and
make the bench less realistic. See ``_seed_company``'s docstring in
``roboco/eval/runner.py`` for the authoritative statement.
"""
from __future__ import annotations
import json
from pathlib import Path
from typing import TYPE_CHECKING
from roboco.config import settings
from roboco.foundation import identity as _foundation
from roboco.runtime.orchestrator import AgentOrchestrator
if TYPE_CHECKING:
import pytest
_AGENT_SLUG = "be-dev-1"
_DISPOSABLE_URL = "http://localhost:9999"
async def test_mcp_config_uses_disposable_api_url_when_set(
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""With ``settings.api_url`` patched to a disposable URL, the generated
MCP config's ROBOCO_API_URL / ROBOCO_ORCHESTRATOR_URL point at the
disposable URL not the production hostname or 127.0.0.1:port."""
monkeypatch.setattr(settings, "api_url", _DISPOSABLE_URL)
orch = AgentOrchestrator.__new__(AgentOrchestrator)
config_path = await orch._generate_mcp_config(_AGENT_SLUG)
config = json.loads(Path(config_path).read_text())
# Every MCP server shares the same env dict; sample the first one.
first_env = next(iter(config["mcpServers"].values()))["env"]
assert first_env["ROBOCO_API_URL"] == _DISPOSABLE_URL
assert first_env["ROBOCO_ORCHESTRATOR_URL"] == _DISPOSABLE_URL
assert "roboco-orchestrator" not in first_env["ROBOCO_API_URL"]
assert "127.0.0.1" not in first_env["ROBOCO_API_URL"]
async def test_mcp_config_preserves_real_agent_uuid(
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""The agent UUID in the config is the REAL fixed UUID from
``foundation.identity.AGENTS`` the harness intentionally uses real UUIDs
so orchestrator-internal helpers keyed by the static registry resolve
exactly as they would in a real deployment.
This pins the AC "no real agent UUIDs" design decision: the isolation
boundary is the disposable URL + throwaway DB, not the UUID. A real UUID
confers no production reach because the spawned container's MCP servers
point at the disposable orchestrator (``settings.api_url`` patched to
``stack.container_url``), never the production one. See
``_seed_company``'s docstring in ``roboco/eval/runner.py``."""
monkeypatch.setattr(settings, "api_url", _DISPOSABLE_URL)
orch = AgentOrchestrator.__new__(AgentOrchestrator)
config_path = await orch._generate_mcp_config(_AGENT_SLUG)
config = json.loads(Path(config_path).read_text())
first_env = next(iter(config["mcpServers"].values()))["env"]
expected_uuid = str(_foundation.AGENTS[_AGENT_SLUG].uuid)
assert first_env["ROBOCO_AGENT_ID"] == expected_uuid