feat(eval): golden-task eval harness + doctrine cohort stamp (#655)

* fix(notifications): exponential backoff + CAS claim for expired-unacked re-escalation

The sweep re-escalated every expired unacked ack-required notification
on every ~60s tick, forever — the live incident: 3 fresh blocker
escalations + Telegram DMs per minute from a static stale pile. Now
each notification carries reescalation_count / last_reescalated_at /
reescalation_delivered_count (migration 079): first fire at expiry,
then doubling intervals from 1h capped at 24h, hard stop after
ROBOCO_NOTIFICATION_MAX_REESCALATIONS (default 5) with one permanent
log carrying attempts-vs-delivered so 'seen and ignored' is
distinguishable from 'route never worked'. The due/wait/capped decision
is a pure function in foundation/policy/communications.py.

Per adversarial review, the attempt slot is claimed by compare-and-set
(UPDATE ... WHERE reescalation_count = :n) BEFORE delivery — the
previous draft leaned on the 60s dedup window, which never engages for
BLOCKER_ESCALATION (_LOOP_PRONE_TYPES excludes it), so concurrent
sweeps would have double-delivered. A lost claim skips delivery
outright. Legacy rows read as count=0 and keep today's first-fire
semantics. 61 tests incl. a two-session CAS race and a real alembic
upgrade/downgrade round trip.

* feat(budgets): per-task and per-project cost budgets (flag-gated)

tasks.budget_usd + projects.monthly_budget_usd (migration 080, chained
on 079; adds ix_agent_spawn_sessions_task_id since both enforcement
seams filter on bare task_id). Behind ROBOCO_TASK_BUDGETS_ENABLED
(default off, feature-flags card) — verifiably inert when off.

Claim-time: a project-month-spend guard applies to WORK-STARTING claims
only (i_will_work_on / i_will_plan) — per adversarial review, review/
doc/gate/inbound-PR claims are exempt so in-flight work can always
finish reviewing and merging at cap. Spend counts closed sessions'
estimated_cost_usd PLUS open sessions priced live from token snapshots
(the original closed-only sum read parallel long sessions as $0).

Sweep-side: the existing budget sweep also prices the active task's
spend vs budget_usd (TaskType defaults when null); on breach the task
is BLOCKED (HUMAN resolver, budget marker) BEFORE the graceful stop so
the unclaim no-ops and the dispatcher never respawns onto it, and the
CEO notification names both recovery steps. unblock on a budget-blocked
task re-checks live spend and refuses while still over — no silent
re-breach loop. Panel: budget inputs in both dialogs (0 rejected — a
zero budget silently blocks everything), spend logic consolidated in
TaskService.task_spend_usd. 42 new tests incl. a real-DB spend-query
suite and a two-tick non-refire sweep test.

* feat(eval): golden-task eval harness + doctrine cohort stamp

roboco/eval: 6 BenchTaskSpec fixtures run through the real lifecycle in
a disposable environment (the e2e_smoke harness's fake GitHub + local
git origin + throwaway DB catalog — real isolation, not convention),
scored deterministically (terminal status, revision_count, cycle time,
tokens/cost via the agent_spawn_sessions task_id join) plus a local-
model judge whose output is nested under a non_deterministic-marked
object so cohort diffs don't read judge noise as regression. CLI:
python -m roboco.eval run --role <slug> --cohort <name>. Source-
checkout-only by declared posture (deptry-scoped ignore + a hard
ImportError guard naming why; tests/ never ships in images or wheels).

agent_spawn_sessions.doctrine_version (migration 081, chained on 080)
is stamped at spawn-session finalize from the composed prompt layers —
with the session's model column it identifies a cohort durably.

Per adversarial review: bench runs patch the vault flags off (they were
writing real markdown into the operator's vault), and the real-spawn
OrchestratorStageSpawner is deliberately cut to NotImplementedError —
spawned containers' MCP wiring resolves to the production orchestrator
under real agent UUIDs, so real spawns wait for a dedicated follow-up;
the injectable scripted spawner is the working path. Full suite 13852
passed / 94% coverage in the source worktree; deptry/mypy/xenon clean.

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
This commit is contained in:
Renzo F
2026-07-23 00:06:50 +02:00
committed by GitHub
co-authored by Renn F
parent 7c8453e210
commit 10f039c36f
13 changed files with 2031 additions and 4 deletions
+57
View File
@@ -0,0 +1,57 @@
"""Doctrine-version stamp on agent_spawn_sessions, for the eval harness.
The eval harness (roboco/eval/) scores a (role, model/provider config) cohort
by replaying golden tasks through a real agent spawn. To attribute a quality
delta to a prompt/doctrine change (fable-mode, ponytail, a team-prompt edit,
...) the resulting spawn session needs to carry a fingerprint of exactly what
system prompt it ran with — otherwise two cohort runs are only comparable if
the operator remembers to keep everything else byte-for-byte identical.
``doctrine_version`` is a short hash of the composed system prompt (base +
role + team + identity + doctrine layers) for that spawn, stamped at
``_finalize_spawn_session`` in roboco/runtime/orchestrator.py — NOT at
``_record_spawn_session`` (spawn creation). The composed prompt string itself
is not passed through the AgentConfig the finalize call site holds, but the
file it was written to (``config.blueprint_path``, from
``_generate_composed_prompt``) is still on disk and unchanged at finalize
time (nothing in the spawn/stop path deletes it), so the finalize call reads
it back and hashes it there. Every provider gets one — ``_prepare_agent_spawn``
composes and writes the blueprint unconditionally, before provider/route
resolution, so GROK agents carry a real blueprint file too, same as Claude.
Nullable + additive: every existing row, and any row where the read
genuinely fails (a provider-parked stub instance that never actually
spawned — ``blueprint_path=Path()`` — an evicted temp dir, ...), simply gets
NULL — a pure quality-of-life addition to the sessions the eval harness
scores, never a hard requirement of the spawn/stop path.
Revision ID: 081_doctrine_version
Revises: 080_task_project_budgets
Create Date: 2026-07-22
Note: re-chained onto 080_task_project_budgets (sibling PRs #652/#654 own
079/080 at this branch's base commit, da4d9b33, where 078 was the head);
080 does not exist in this worktree, so the local migration-graph/enum-parity
tests are expected to fail here until this branch integrates alongside its
siblings — the same expected-failure posture the budgets sibling reported.
"""
from __future__ import annotations
import sqlalchemy as sa
from alembic import op
revision = "081_doctrine_version"
down_revision = "080_task_project_budgets"
branch_labels: dict[str, str] | None = None
depends_on: dict[str, str] | None = None
def upgrade() -> None:
op.add_column(
"agent_spawn_sessions",
sa.Column("doctrine_version", sa.String(length=32), nullable=True),
)
def downgrade() -> None:
op.drop_column("agent_spawn_sessions", "doctrine_version")
+24 -2
View File
@@ -198,6 +198,13 @@ select = [
# cost; ARG001 covers FastAPI path params the fake-GitHub handlers must
# name but not read.
"tests/e2e_smoke/*.py" = ["PLC0415", "ARG001"]
# The eval bench (offline CLI, not part of the served app) defers heavy/
# optional imports — tests.e2e_smoke.harness, roboco.runtime.orchestrator,
# roboco.services.task, asyncpg — to call time for the same reason
# tests/e2e_smoke and roboco/services do; PLR0913 covers the stage-driving
# and scoring call surfaces (task/spawner/role/timeout tuples), same
# rationale as roboco/services/gateway/**.
"roboco/eval/*.py" = ["PLC0415", "PLR0913"]
# PTH119: _grok_usage_json sanitizes the agent id with os.path.basename — the
# path-injection sanitizer CodeQL's query models; the pathlib equivalent
# (Path(...).name) is not recognized by that query, so we keep os.path here.
@@ -233,10 +240,18 @@ select = [
# signature for keyword-argument compatibility (mypy override check), but the
# stub bodies are empty — ARG002 would require renaming them, which breaks mypy.
"tests/unit/services/test_optimal_grounding.py" = ["ARG002"]
# _FakeJudge.score overrides BenchJudge.score — same override-signature
# rationale as test_optimal_grounding.py above (a fixed fixture/diff/notes
# stand-in body has nothing to do with those args).
"tests/e2e_smoke/test_eval_bench.py" = ["ARG002"]
# Collision-builder test helpers mirror the builder's many keyword inputs
# (parent/project/intends/migration/shared/sequence) — bundling them would
# hurt readability more than the arg count hurts.
"tests/unit/gateway/test_collision_context.py" = ["PLR0913"]
# _metrics()'s many optional kwargs mirror DeterministicMetrics' own field
# count (a plain, no-defaults dataclass) — same rationale as the collision
# builder above.
"tests/unit/eval/test_scoring.py" = ["PLR0913"]
# =============================================================================
# MyPy Configuration
@@ -452,8 +467,15 @@ DEP002 = [
"types-passlib",
"types-PyYAML",
]
# DEP003: Starlette is a transitive dep of FastAPI, but BaseHTTPMiddleware is needed
DEP003 = ["starlette"]
# DEP003: Starlette is a transitive dep of FastAPI, but BaseHTTPMiddleware is needed.
# "tests": roboco/eval/runner.py deliberately imports tests.e2e_smoke.harness/arcs
# (the offline eval bench's disposable-project machinery — see that module's
# docstring) — deptry sees the local `tests` package as an unresolvable
# transitive import since it isn't a PyPI dependency at all. The runtime side
# of this is guarded separately (an ImportError there raises a clear "this
# needs a source checkout" error), so this is a lint-posture ignore, not a
# correctness gap.
DEP003 = ["starlette", "tests"]
[dependency-groups]
dev = [
+6
View File
@@ -1904,6 +1904,12 @@ class AgentSpawnSessionTable(Base):
tool_calls: Mapped[int] = mapped_column(BigInteger, nullable=False, default=0)
exit_reason: Mapped[str | None] = mapped_column(String(100), nullable=True)
estimated_cost_usd: Mapped[float | None] = mapped_column(Float, nullable=True)
# Short hash of the composed system prompt this spawn ran with (migration
# 081_doctrine_version) — lets the eval harness (roboco/eval/) group spawn
# sessions by the exact prompt/doctrine version they ran, not just by
# model. Stamped at finalize (see _finalize_spawn_session); nullable —
# older rows and any read failure carry NULL.
doctrine_version: Mapped[str | None] = mapped_column(String(32), nullable=True)
# Relationship to snapshots (backref for convenience)
snapshots: Mapped[list["TokenUsageSnapshotTable"]] = relationship(
+1
View File
@@ -0,0 +1 @@
"""Golden-task quality bench — see ``roboco/eval/runner.py``."""
+95
View File
@@ -0,0 +1,95 @@
"""CLI entrypoint for the eval bench.
python -m roboco.eval run --role <slug> --cohort <name> \\
[--fixtures a,b] [--json-out path]
NOT YET FUNCTIONAL: the real-spawn path (``OrchestratorStageSpawner``) is
deliberately cut see ``roboco/eval/runner.py``'s module docstring's
"Real-spawn status" section because a real container spawn's MCP wiring
would authenticate against the REAL production orchestrator, not this
harness's disposable one. ``run`` will raise ``NotImplementedError`` once it
reaches the first fixture. The only working path today is driving
``EvalRunner`` with an injected scripted ``StageSpawner`` from Python (see
``tests/e2e_smoke/test_eval_bench.py``); this CLI is wired for the day the
follow-up lands, not for use today.
Offline dev/ops tool: no panel surface, no feature flag. Also runs from a
source checkout only (needs ``tests/e2e_smoke``, not shipped in containers
or wheels) plus the test Postgres (``ROBOCO_TEST_DB_*`` env vars, mirroring
the rest of the gate toolchain).
"""
from __future__ import annotations
import argparse
import sys
from pathlib import Path
from roboco.eval.fixtures import FIXTURES, BenchTaskSpec
from roboco.eval.runner import EvalRunner
def _parse_args(argv: list[str]) -> argparse.Namespace:
parser = argparse.ArgumentParser(prog="python -m roboco.eval")
subparsers = parser.add_subparsers(dest="command", required=True)
run = subparsers.add_parser(
"run",
help=(
"Replay the golden-task fixtures against one agent "
"[NOT YET FUNCTIONAL — real-spawn path is cut, see module docstring]"
),
)
run.add_argument(
"--role", required=True, help="Agent slug under test (e.g. be-dev-1)"
)
run.add_argument(
"--cohort",
required=True,
help="Label for this run, for before/after comparison (e.g. baseline)",
)
run.add_argument(
"--fixtures",
default=None,
help="Comma-separated fixture keys (default: every developer-role fixture)",
)
run.add_argument(
"--json-out",
default=None,
type=Path,
help="Write the scored cohort result as JSON to this path",
)
return parser.parse_args(argv)
def _select_fixtures(spec: str | None) -> tuple[list[BenchTaskSpec] | None, str | None]:
"""Resolve `--fixtures a,b` to a fixture list, or an error message for an
unknown key. `(None, None)` means "no filter — run every fixture"."""
if not spec:
return None, None
keys = {key.strip() for key in spec.split(",") if key.strip()}
fixtures = [f for f in FIXTURES if f.key in keys]
missing = keys - {f.key for f in fixtures}
if missing:
return None, f"Unknown fixture key(s): {sorted(missing)}"
return fixtures, None
def main(argv: list[str] | None = None) -> int:
args = _parse_args(sys.argv[1:] if argv is None else argv)
if args.command != "run":
return 1
fixtures, error = _select_fixtures(args.fixtures)
if error:
print(error, file=sys.stderr)
return 2
runner = EvalRunner()
runner.run_cohort(args.role, args.cohort, fixtures=fixtures, json_out=args.json_out)
return 0
if __name__ == "__main__":
raise SystemExit(main())
+262
View File
@@ -0,0 +1,262 @@
"""Golden-task fixtures for the eval bench (see ``roboco/eval/runner.py``).
Each ``BenchTaskSpec`` is a tiny, self-contained "golden task": a few
pre-seeded repo files, a task brief (title/description/acceptance criteria),
and a checked-in ``expectations`` note the local-model judge grades the final
PR diff + dev notes against. Fixture repo files are namespaced under
``bench/<key>/`` so every fixture can share one disposable project's git
history without colliding with the others (the runner seeds fixtures onto
the same project's default branch, one at a time).
"""
from __future__ import annotations
from dataclasses import dataclass
from roboco.models.base import TaskNature, TaskType
@dataclass(frozen=True)
class BenchTaskSpec:
"""One golden task: seeded repo state + brief + graded expectation.
``target_role`` is always ``"developer"`` the only role a task can be
freshly assigned to from PENDING with no prior work already done (QA /
documenter / cell-PM only ever pick up a task a developer has already
advanced through the lifecycle). Kept as an explicit field rather than
hardcoded at the call site so a future QA/PM-focused bench fixture has
somewhere to say otherwise.
"""
key: str
title: str
description: str
acceptance_criteria: tuple[str, ...]
task_type: TaskType
nature: TaskNature
repo_files: tuple[tuple[str, str], ...]
expectations: str
target_role: str = "developer"
FIXTURES: tuple[BenchTaskSpec, ...] = (
BenchTaskSpec(
key="bugfix-off-by-one",
title="Fix off-by-one in paginate()",
description=(
"`bench/bugfix-off-by-one/paginate.py`'s `paginate(items, page, "
"size)` drops the last item of every page because its slice end "
"is `page * size - 1` instead of `page * size`. Fix the slice "
"bound so every item appears exactly once across all pages."
),
acceptance_criteria=(
"paginate(list(range(10)), page=1, size=3) returns [0, 1, 2]",
"paginate(list(range(10)), page=4, size=3) returns [9] (the "
"last, previously-dropped item)",
"No item is duplicated or skipped across pages 1..4 for size=3",
),
task_type=TaskType.CODE,
nature=TaskNature.TECHNICAL,
repo_files=(
(
"bench/bugfix-off-by-one/paginate.py",
"def paginate(items, page, size):\n"
" start = (page - 1) * size\n"
" end = page * size - 1\n"
" return items[start:end]\n",
),
),
expectations=(
"The fix changes the slice end to `page * size` (or an "
"equivalent that includes the final item). A correct diff "
"touches only paginate.py's slice bound; no new dependency, no "
"unrelated rewrite. Commit/PR notes should describe the "
"off-by-one root cause, not just 'fixed a bug'."
),
),
BenchTaskSpec(
key="bugfix-null-check",
title="Fix crash on empty input in summarize()",
description=(
"`bench/bugfix-null-check/stats.py`'s `summarize(values)` "
"divides by `len(values)` unconditionally, so it raises "
"ZeroDivisionError on an empty list instead of returning a "
"sane empty-input result. Add a guard."
),
acceptance_criteria=(
"summarize([]) returns {'count': 0, 'total': 0, 'average': 0} "
"without raising",
"summarize([2, 4, 6]) still returns "
"{'count': 3, 'total': 12, 'average': 4}",
),
task_type=TaskType.CODE,
nature=TaskNature.TECHNICAL,
repo_files=(
(
"bench/bugfix-null-check/stats.py",
"def summarize(values):\n"
" total = sum(values)\n"
" count = len(values)\n"
" return {\n"
" 'count': count,\n"
" 'total': total,\n"
" 'average': total / count,\n"
" }\n",
),
),
expectations=(
"The fix adds an explicit empty-input guard (e.g. `if not "
"values: return {...}`) ahead of the division, without "
"changing the non-empty behavior. No new dependency; the "
"guard is the whole diff."
),
),
BenchTaskSpec(
key="small-feature-greet",
title="Add a greet() helper",
description=(
"`bench/small-feature-greet/greetings.py` has no greeting "
"helper yet. Add a `greet(name, formal=False)` function: "
"informal returns `f'Hi, {name}!'`, formal returns "
"`f'Good day, {name}.'`. Empty/whitespace-only `name` should "
"raise `ValueError`."
),
acceptance_criteria=(
"greet('Ada') == 'Hi, Ada!'",
"greet('Ada', formal=True) == 'Good day, Ada.'",
"greet('') and greet(' ') both raise ValueError",
),
task_type=TaskType.CODE,
nature=TaskNature.TECHNICAL,
repo_files=(
(
"bench/small-feature-greet/greetings.py",
"# Greeting helpers for bench/small-feature-greet.\n",
),
),
expectations=(
"greet() is added to greetings.py matching both the informal "
"and formal wording exactly, plus the empty-name ValueError "
"guard. A minimal, additive diff — no unrelated changes to "
"the file's header comment."
),
),
BenchTaskSpec(
key="refactor-duplicate-normalize",
title="De-duplicate normalize_a / normalize_b",
description=(
"`bench/refactor-duplicate-normalize/normalize.py` has two "
"near-identical functions, `normalize_a` and `normalize_b` — "
"both strip whitespace and lowercase a string, differing only "
"in which module used to call them. Refactor into one shared "
"helper both call, preserving both public names as thin "
"wrappers so existing callers are unaffected."
),
acceptance_criteria=(
"normalize_a(' Hello ') == 'hello'",
"normalize_b(' Hello ') == 'hello'",
"The duplicated strip/lower logic exists in exactly one place",
),
task_type=TaskType.CODE,
nature=TaskNature.TECHNICAL,
repo_files=(
(
"bench/refactor-duplicate-normalize/normalize.py",
"def normalize_a(text):\n"
" return text.strip().lower()\n"
"\n"
"\n"
"def normalize_b(text):\n"
" return text.strip().lower()\n",
),
),
expectations=(
"A single private helper (e.g. `_normalize`) holds the "
"strip/lower logic; normalize_a/normalize_b both delegate to "
"it and keep their existing signatures and return values "
"identical to before. No behavior change, pure de-duplication."
),
),
BenchTaskSpec(
key="docs-readme-flag",
title="Document the --dry-run flag",
description=(
"`bench/docs-readme-flag/cli.py` accepts a `--dry-run` flag "
"(prints what it would do instead of doing it) that isn't "
"mentioned anywhere in `bench/docs-readme-flag/README.md`. Add "
"a short section documenting it: what it does and an example "
"invocation."
),
acceptance_criteria=(
"README.md documents --dry-run's behavior in prose",
"README.md shows an example command line using --dry-run",
),
task_type=TaskType.DOCUMENTATION,
nature=TaskNature.TECHNICAL,
repo_files=(
(
"bench/docs-readme-flag/cli.py",
"import argparse\n"
"\n"
"\n"
"def build_parser():\n"
" parser = argparse.ArgumentParser()\n"
" parser.add_argument('--dry-run', action='store_true')\n"
" return parser\n",
),
(
"bench/docs-readme-flag/README.md",
"# bench/docs-readme-flag\n\nA tiny CLI fixture.\n",
),
),
expectations=(
"README.md gains a section documenting --dry-run's actual "
"behavior (prints instead of acting) with a runnable example "
"invocation. cli.py itself is unchanged — this is a docs-only "
"task."
),
),
BenchTaskSpec(
key="research-magic-constant",
title="Research the magic constant in legacy_calc.py",
description=(
"`bench/research-magic-constant/legacy_calc.py`'s `compute()` "
"multiplies by `1.10000001` instead of the obvious `1.1`. "
"Investigate the surrounding code/comments for why, and commit "
"your findings as `bench/research-magic-constant/NOTES.md` — "
"this is a research task, not a code fix: legacy_calc.py stays "
"unchanged."
),
acceptance_criteria=(
"NOTES.md exists and explains what the magic constant is "
"compensating for, based on the evidence in the file",
"NOTES.md recommends whether it's safe to simplify to 1.1, with reasoning",
"legacy_calc.py is not modified",
),
task_type=TaskType.RESEARCH,
nature=TaskNature.TECHNICAL,
repo_files=(
(
"bench/research-magic-constant/legacy_calc.py",
"# The 1.10000001 factor below is NOT a typo for 1.1 — it\n"
"# nudges the float rounding in compute()'s downstream\n"
"# int(...) truncation so historical invoice totals ending\n"
"# in .10 don't get truncated to one cent short. Changing\n"
"# this constant reopens ROBO-lore ticket #4471 (pre-git\n"
"# history) where 1.1 exactly caused a cent-level\n"
"# reconciliation mismatch on ~0.3% of invoices.\n"
"def compute(amount):\n"
" return int(amount * 1.10000001)\n",
),
),
expectations=(
"NOTES.md correctly identifies (from the file's own comment) "
"that the constant compensates for float-truncation rounding "
"in compute()'s int(...) cast, cites the historical "
"reconciliation-mismatch reasoning, and recommends AGAINST "
"simplifying to 1.1 without a broader fix — it should not "
"invent an unrelated explanation, and legacy_calc.py itself "
"must be untouched."
),
),
)
File diff suppressed because it is too large Load Diff
+31
View File
@@ -14,6 +14,7 @@ The orchestrator is the BRAIN of the system:
import asyncio
import contextlib
import hashlib
import json
import os
import shutil
@@ -6311,6 +6312,7 @@ class AgentOrchestrator:
if instance and instance.config:
model = instance.config.model or "unknown"
usage_session_id = instance.usage_session_id if instance else None
doctrine_version = self._doctrine_version_for_instance(instance)
cost = calculate_cost(
model=model,
@@ -6358,6 +6360,7 @@ class AgentOrchestrator:
tool_calls=tool_calls,
exit_reason=exit_reason,
estimated_cost_usd=cost,
doctrine_version=doctrine_version,
)
)
await db.commit()
@@ -6368,6 +6371,7 @@ class AgentOrchestrator:
tokens_input=tokens_input,
tokens_output=tokens_output,
estimated_cost_usd=cost,
doctrine_version=doctrine_version,
)
except Exception as exc:
logger.warning(
@@ -6376,6 +6380,33 @@ class AgentOrchestrator:
error=str(exc),
)
@staticmethod
def _doctrine_version_for_instance(instance: AgentInstance | None) -> str | None:
"""Short hash of the composed system prompt this spawn ran with.
Reads the SAME file ``_generate_composed_prompt`` wrote at spawn
preparation (``config.blueprint_path``) nothing on the spawn/stop
path deletes it, so it is still the exact prompt text this agent ran
with. Every provider gets one (the blueprint is composed and written
unconditionally in ``_prepare_agent_spawn``, before provider/route
resolution) GROK agents carry a real blueprint file too, same as
Claude. Best-effort: a missing/unreadable file (a provider-parked stub
instance that never actually spawned ``blueprint_path=Path()`` an
evicted temp dir, ...) is NULL, never a finalize failure the eval
harness treats an unstamped session as "doctrine unknown", not an
error.
"""
if instance is None or instance.config is None:
return None
blueprint_path = instance.config.blueprint_path
if not blueprint_path:
return None
try:
content = blueprint_path.read_text()
except OSError:
return None
return hashlib.sha256(content.encode("utf-8")).hexdigest()[:16]
@staticmethod
async def _fetch_agent_tokens(
client: httpx.AsyncClient, agent_id: str
+12
View File
@@ -706,6 +706,18 @@ ROADMAP_ITEM_SOURCE = "roadmap"
# the Main PM. The board routing is the start gate; no held-source skip.
VAULT_NOTE_SOURCE = "vault_note"
# Source tag for a golden-task fixture the offline eval bench
# (``roboco/eval/``) replays through the real delivery lifecycle to score a
# (role, model/provider config) cohort. Deliberately absent from every
# held-source / non-dev-dispatch set above: an eval_bench task is a normal,
# pre-assigned dev leaf task and must dispatch exactly like one (real spawn,
# real QA, real docs, real cell-PM merge) — every OTHER engine (self-heal,
# ci-watch, dep-update, docs-sync, release-manager, X, video, roadmap, vault-
# intake) only ever queries/dedupes by ITS OWN source constant above, so an
# eval_bench task is invisible to all of them by construction, not by an
# explicit exemption.
EVAL_BENCH_SOURCE = "eval_bench"
def extract_self_heal_fingerprint(task: Any) -> str | None:
"""The self-heal dedupe fingerprint from a task's markers, or None.
+20 -2
View File
@@ -45,8 +45,20 @@ if TYPE_CHECKING:
from collections.abc import Iterator
from pathlib import Path
from types import ModuleType
from typing import Protocol
from uuid import UUID
class TmpPathFactory(Protocol):
"""Structural stand-in for the one ``pytest.TempPathFactory`` method
``build_e2e_stack`` uses. pytest's real fixture value already
satisfies this shape, so it needs no adapter but it lets
``roboco/eval/runner.py`` (an offline CLI, not a pytest session) drive
this same stack-building machinery with a plain temp-dir factory
instead of constructing a real ``pytest.Config``."""
def mktemp(self, basename: str, numbered: bool = True) -> Path: ...
_OWNER = "e2e-smoke"
_REPO = "proj"
@@ -360,9 +372,15 @@ def _build_app(gh: _FakeGitHub) -> FastAPI:
def build_e2e_stack(
_test_database_url: str, tmp_path_factory: pytest.TempPathFactory
_test_database_url: str, tmp_path_factory: TmpPathFactory
) -> Iterator[E2EStack]:
"""Generator behind the ``e2e_stack`` fixture (defined in conftest)."""
"""Generator behind the ``e2e_stack`` fixture (defined in conftest).
``tmp_path_factory`` only needs ``.mktemp()`` (see ``TmpPathFactory``
above) pytest's real fixture satisfies it structurally, and
``roboco/eval/runner.py`` drives this same function with a plain
non-pytest factory to reuse this stack outside a test session.
"""
from roboco.config import settings
from roboco.db import base as db_base
+204
View File
@@ -0,0 +1,204 @@
"""Integration test for the eval bench's own orchestration/scoring plumbing.
The real ``StageSpawner`` (``OrchestratorStageSpawner``) drives a REAL agent
container via ``AgentOrchestrator.spawn_agent`` and needs a Docker daemon +
built agent images it cannot run here (see ``roboco/eval/runner.py``'s
module docstring). This test substitutes a scripted stand-in that drives the
SAME real MCP flow/do tool functions ``tests.e2e_smoke.harness.ScriptedAgent``
uses (via the existing ``dev_arc`` / ``qa_arc`` / ``doc_arc`` helpers, plus a
PM ``complete`` call) so it proves the runner's OWN code — its throwaway-DB +
disposable-project setup, its status-driven stage loop, its PM pre-claim,
its deterministic scoring, its JSON/table output without touching Docker.
Runs the smallest fixture (a single-file bug fix) end to end: PENDING ->
awaiting_qa -> awaiting_documentation -> awaiting_pm_review -> completed.
Gating: like every other module here, this is skipped unless
``ROBOCO_E2E_SMOKE=1`` (see ``tests/e2e_smoke/conftest.py``'s
``pytest_collection_modifyitems``) it needs the real test Postgres, which
``EvalRunner`` provisions its own throwaway copy of (see
``roboco/eval/runner.py``'s ``_scratch_database``), independent of this
package's shared session-scoped ``e2e_stack`` fixture.
"""
from __future__ import annotations
import asyncio
from typing import TYPE_CHECKING, Any
from uuid import UUID
from roboco.config import settings
from roboco.eval.fixtures import FIXTURES
from roboco.eval.runner import BenchJudge, EvalRunner, JudgeVerdict, _bench_environment
from tests.e2e_smoke.arcs import Company, dev_arc, doc_arc, qa_arc
from tests.e2e_smoke.harness import ScriptedAgent
if TYPE_CHECKING:
import pytest
from roboco.eval.fixtures import BenchTaskSpec
from tests.e2e_smoke.harness import E2EStack
_FIXTURE_KEY = "bugfix-off-by-one"
_FIXED_FIX = (
"def paginate(items, page, size):\n"
" start = (page - 1) * size\n"
" end = page * size\n"
" return items[start:end]\n"
)
def _fixture() -> BenchTaskSpec:
for f in FIXTURES:
if f.key == _FIXTURE_KEY:
return f
raise AssertionError(f"{_FIXTURE_KEY!r} fixture not found in FIXTURES")
def _stub_company() -> Company:
"""A ``Company`` carrying the FIXED uuids ``EvalRunner``'s own company
seeding uses (not fresh random ones, unlike ``arcs.seed_company``) the
"be-*" slugs are hardcoded inside ``dev_arc`` / ``qa_arc`` / ``doc_arc``
themselves, so this only needs to supply the matching ids."""
from roboco.foundation import identity as _foundation
company = Company()
company.dev_id = _foundation.AGENTS["be-dev-1"].uuid
company.qa_id = _foundation.AGENTS["be-qa"].uuid
company.doc_id = _foundation.AGENTS["be-doc"].uuid
company.cell_pm_id = _foundation.AGENTS["be-pm"].uuid
return company
class _ScriptedBenchSpawner:
"""Test-only ``StageSpawner``: applies the KNOWN correct fix via the real
MCP flow/do tool functions, standing in for a real container spawn.
``dev_arc`` / ``qa_arc`` / ``doc_arc`` (and ``ScriptedAgent`` itself) call
``E2EStack.run_db``, which runs its own ``asyncio.run()`` per call fine
from a plain sync pytest test, but ``run_stage`` is awaited from inside
``_drive_task_to_terminal``'s own event loop, where a nested
``asyncio.run()`` raises. Running the scripted turn on a worker thread
(``asyncio.to_thread``) gives it a thread with no running loop, exactly
like the sync test functions those helpers were written for.
"""
def __init__(self, stack: E2EStack) -> None:
self._stack = stack
self._company = _stub_company()
async def run_stage(self, *, task: dict[str, Any], agent_slug: str) -> None:
await asyncio.to_thread(self._run_stage_sync, task, agent_slug)
def _run_stage_sync(self, task: dict[str, Any], agent_slug: str) -> None:
from roboco.agents_config import get_agent_role
role = get_agent_role(agent_slug)
task_id = UUID(task["id"])
if role == "developer":
dev_arc(
self._stack,
self._company,
task["project_slug"],
task_id,
work=(f"bench/{_FIXTURE_KEY}/paginate.py", _FIXED_FIX),
)
elif role == "qa":
qa_arc(self._stack, self._company, task_id)
elif role == "documenter":
doc_arc(
self._stack,
self._company,
task_id,
filename=f"bench/{_FIXTURE_KEY}/paginate.py",
)
elif role == "cell_pm":
pm = ScriptedAgent(
self._stack, self._company.cell_pm_id, agent_slug, "cell_pm"
)
pm.flow(
"complete",
task_id=str(task_id),
notes="Scripted bench completion: QA passed, docs complete.",
)
else:
raise AssertionError(f"unexpected role for the scripted bench: {role!r}")
_EXPECTED_JUDGE_SCORE = 5
class _FakeJudge(BenchJudge):
"""Deterministic stand-in for the local-model judge — no network."""
async def score(
self, *, fixture: BenchTaskSpec, diff: str, notes: str
) -> JudgeVerdict:
return JudgeVerdict(
score=_EXPECTED_JUDGE_SCORE, rationale="scripted test: assumed correct"
)
def test_eval_runner_drives_a_fixture_to_completion_with_a_scripted_spawn() -> None:
runner = EvalRunner(
make_spawner=_ScriptedBenchSpawner,
judge=_FakeJudge(),
fixture_timeout_seconds=60.0,
)
cohort = runner.run_cohort(
"be-dev-1", "scripted-test", fixtures=[_fixture()], json_out=None
)
assert cohort.role_slug == "be-dev-1"
assert len(cohort.fixtures) == 1
result = cohort.fixtures[0]
assert result.fixture_key == _FIXTURE_KEY
assert result.metrics.final_status == "completed"
assert result.metrics.stalled is False
assert result.passed is True
assert result.metrics.revision_count == 0
assert result.judge.score == _EXPECTED_JUDGE_SCORE
assert cohort.pass_rate == 1.0
# No real container spawned, so no agent_spawn_sessions rows accrued for
# this task — the scripted stand-in proves the runner's DB/polling/
# scoring plumbing, not token/cost accounting (that needs a real spawn;
# see the module docstring).
assert result.metrics.total_tokens == 0
assert result.metrics.estimated_cost_usd == 0.0
def test_bench_environment_disables_vault_writes_even_when_ambient_flags_are_armed(
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""A bench run must never write into the operator's REAL Obsidian vault.
Simulates the compose-default posture (every vault flag armed True) and
asserts `_bench_environment` forces them all off for its duration, then
restores the prior values on exit the exact leak an adversarial review
flagged (TaskService.create / JournalService / A2AService all gate on
obsidian_vault_enabled first, so patching it is the load-bearing part;
the three sub-flags are patched too for defense-in-depth)."""
monkeypatch.setattr(settings, "obsidian_vault_enabled", True)
monkeypatch.setattr(settings, "vault_intake_enabled", True)
monkeypatch.setattr(settings, "vault_kb_enabled", True)
monkeypatch.setattr(settings, "vault_report_enabled", True)
armed = (
settings.obsidian_vault_enabled,
settings.vault_intake_enabled,
settings.vault_kb_enabled,
settings.vault_report_enabled,
)
with _bench_environment("be-dev-1"):
assert settings.obsidian_vault_enabled is False
assert settings.vault_intake_enabled is False
assert settings.vault_kb_enabled is False
assert settings.vault_report_enabled is False
# Restored to the (simulated ambient) armed state once the bench exits.
restored = (
settings.obsidian_vault_enabled,
settings.vault_intake_enabled,
settings.vault_kb_enabled,
settings.vault_report_enabled,
)
assert restored == armed == (True, True, True, True)
+74
View File
@@ -0,0 +1,74 @@
"""Schema checks for the golden-task fixtures (roboco/eval/fixtures.py).
Nothing here touches a DB or the network these are pure sanity checks on
the static FIXTURES tuple so a malformed fixture (a duplicate key, a fixture
file that escapes its own bench/<key>/ namespace and could collide with
another fixture's repo state, an empty brief) is caught before it ever
reaches the runner.
"""
from __future__ import annotations
import dataclasses
from typing import Any, cast
import pytest
from roboco.eval.fixtures import FIXTURES, BenchTaskSpec
_MIN_FIXTURES = 5
_MAX_FIXTURES = 8
def test_fixture_keys_are_unique() -> None:
keys = [f.key for f in FIXTURES]
assert len(keys) == len(set(keys)), f"duplicate fixture keys: {keys}"
def test_at_least_five_fixtures() -> None:
# The task calls for 5-8 canonical fixtures.
assert _MIN_FIXTURES <= len(FIXTURES) <= _MAX_FIXTURES, len(FIXTURES)
def test_every_fixture_has_a_non_empty_brief() -> None:
for f in FIXTURES:
assert f.title.strip(), f.key
assert f.description.strip(), f.key
assert f.acceptance_criteria, f"{f.key} has no acceptance criteria"
assert all(c.strip() for c in f.acceptance_criteria), f.key
assert f.expectations.strip(), f"{f.key} has no judge expectations note"
def test_repo_files_are_namespaced_under_bench_key() -> None:
"""Every fixture's seeded file lives under bench/<its own key>/ so
sequential fixtures sharing one project's git history never collide."""
for f in FIXTURES:
assert f.repo_files, f"{f.key} seeds no repo files"
prefix = f"bench/{f.key}/"
for rel_path, content in f.repo_files:
assert rel_path.startswith(prefix), (
f"{f.key}: {rel_path!r} escapes its own {prefix!r} namespace"
)
assert ".." not in rel_path, f"{f.key}: {rel_path!r} looks like a traversal"
assert content, f"{f.key}: {rel_path!r} has empty content"
def test_repo_file_paths_within_a_fixture_are_unique() -> None:
for f in FIXTURES:
paths = [rel_path for rel_path, _content in f.repo_files]
assert len(paths) == len(set(paths)), f"{f.key}: duplicate paths {paths}"
def test_target_role_is_developer_for_every_fixture() -> None:
"""Matches EvalRunner.run_cohort's current scope cut (see runner.py's
module docstring) every fixture must be runnable by the one role the
bench supports today."""
for f in FIXTURES:
assert f.target_role == "developer", f.key
def test_bench_task_spec_is_frozen() -> None:
spec = FIXTURES[0]
assert isinstance(spec, BenchTaskSpec)
mutable_view = cast("Any", spec)
with pytest.raises(dataclasses.FrozenInstanceError):
mutable_view.title = "mutated"
+221
View File
@@ -0,0 +1,221 @@
"""Unit tests for the eval bench's scorer math (roboco/eval/runner.py).
Pure dataclass/aggregate-property tests no DB, no network, no asyncio.
`_build_judge_prompt` and `BenchJudge`'s score-parsing regex are covered too
since both are pure string logic with no I/O.
"""
from __future__ import annotations
import pytest
from roboco.eval.fixtures import FIXTURES
from roboco.eval.runner import (
_JUDGE_SCORE_RE,
CohortResult,
DeterministicMetrics,
FixtureResult,
JudgeVerdict,
OrchestratorStageSpawner,
_build_judge_prompt,
)
_EXPECTED_TOTAL_TOKENS = 180
_HALF_PASS_RATE = 0.5
_COHORT_TOTAL_TOKENS = 600
_COHORT_MEAN_CYCLE_SECONDS = 20.0
_COHORT_MEAN_JUDGE_SCORE = 5.0
_PASSING_JUDGE_SCORE = 4
def _metrics(
*,
final_status: str = "completed",
stalled: bool = False,
cycle_time_seconds: float = 10.0,
tokens_input: int = 100,
tokens_output: int = 50,
tokens_cache_read: int = 0,
tokens_cache_write: int = 0,
estimated_cost_usd: float = 0.01,
) -> DeterministicMetrics:
return DeterministicMetrics(
final_status=final_status,
stalled=stalled,
revision_count=0,
cycle_time_seconds=cycle_time_seconds,
tokens_input=tokens_input,
tokens_output=tokens_output,
tokens_cache_read=tokens_cache_read,
tokens_cache_write=tokens_cache_write,
estimated_cost_usd=estimated_cost_usd,
)
def test_deterministic_metrics_total_tokens_sums_all_four_buckets() -> None:
m = _metrics(
tokens_input=100, tokens_output=50, tokens_cache_read=25, tokens_cache_write=5
)
assert m.total_tokens == _EXPECTED_TOTAL_TOKENS
def test_fixture_result_passed_requires_completed_and_not_stalled() -> None:
passed = FixtureResult(
fixture_key="a",
metrics=_metrics(final_status="completed", stalled=False),
judge=JudgeVerdict(score=None, rationale=None),
)
assert passed.passed is True
cancelled = FixtureResult(
fixture_key="b",
metrics=_metrics(final_status="cancelled", stalled=False),
judge=JudgeVerdict(score=None, rationale=None),
)
assert cancelled.passed is False
# A stall that happens to leave the row at "completed" is still not a
# pass — `stalled` overrides the status.
stalled_completed = FixtureResult(
fixture_key="c",
metrics=_metrics(final_status="completed", stalled=True),
judge=JudgeVerdict(score=None, rationale=None),
)
assert stalled_completed.passed is False
def _sample_cohort() -> CohortResult:
fixtures = [
FixtureResult(
fixture_key="a",
metrics=_metrics(
final_status="completed",
cycle_time_seconds=10.0,
estimated_cost_usd=0.10,
tokens_input=100,
tokens_output=100,
),
judge=JudgeVerdict(score=5, rationale="great"),
),
FixtureResult(
fixture_key="b",
metrics=_metrics(
final_status="needs_revision",
stalled=True,
cycle_time_seconds=30.0,
estimated_cost_usd=0.20,
tokens_input=200,
tokens_output=200,
),
judge=JudgeVerdict(score=None, rationale="judge unavailable"),
),
]
return CohortResult(role_slug="be-dev-1", cohort_name="baseline", fixtures=fixtures)
def test_cohort_pass_rate_and_totals() -> None:
cohort = _sample_cohort()
assert cohort.pass_rate == _HALF_PASS_RATE
assert cohort.total_cost_usd == pytest.approx(0.3)
assert cohort.total_tokens == _COHORT_TOTAL_TOKENS
assert cohort.mean_cycle_time_seconds == _COHORT_MEAN_CYCLE_SECONDS
# Only fixture "a" has a judge score; "b"'s None is excluded from the mean.
assert cohort.mean_judge_score == _COHORT_MEAN_JUDGE_SCORE
def test_cohort_mean_judge_score_is_none_when_no_fixture_was_scored() -> None:
fixtures = [
FixtureResult(
fixture_key="a",
metrics=_metrics(),
judge=JudgeVerdict(score=None, rationale="judge unavailable"),
)
]
cohort = CohortResult(role_slug="be-dev-1", cohort_name="x", fixtures=fixtures)
assert cohort.mean_judge_score is None
def test_cohort_with_no_fixtures_is_a_zero_result_not_a_crash() -> None:
cohort = CohortResult(role_slug="be-dev-1", cohort_name="x", fixtures=[])
assert cohort.pass_rate == 0.0
assert cohort.total_cost_usd == 0.0
assert cohort.total_tokens == 0
assert cohort.mean_cycle_time_seconds == 0.0
assert cohort.mean_judge_score is None
def test_cohort_as_dict_round_trips_every_fixture() -> None:
fixtures = [
FixtureResult(
fixture_key="a",
metrics=_metrics(),
judge=JudgeVerdict(_PASSING_JUDGE_SCORE, "solid"),
),
]
cohort = CohortResult(role_slug="be-dev-1", cohort_name="x", fixtures=fixtures)
payload = cohort.as_dict()
assert payload["role_slug"] == "be-dev-1"
assert payload["cohort_name"] == "x"
assert payload["aggregate"]["fixture_count"] == 1
assert payload["aggregate"]["pass_rate"] == 1.0
# Judge fields live under their own nested, explicitly-marked object —
# never flat beside deterministic metrics — so a naive diff can't read
# judge noise as a regression.
assert "mean_judge_score" not in payload["aggregate"]
assert payload["judge"] == {
"mean_score": _PASSING_JUDGE_SCORE,
"non_deterministic": True,
}
assert len(payload["fixtures"]) == 1
assert payload["fixtures"][0]["fixture_key"] == "a"
assert "judge_score" not in payload["fixtures"][0]
assert payload["fixtures"][0]["judge"] == {
"score": _PASSING_JUDGE_SCORE,
"rationale": "solid",
"non_deterministic": True,
}
def test_judge_score_regex_parses_the_required_reply_shape() -> None:
reply = "Score: 4\nRationale: matches the expectation closely.\n"
match = _JUDGE_SCORE_RE.search(reply)
assert match is not None
assert int(match.group(1)) == _PASSING_JUDGE_SCORE
def test_judge_score_regex_is_case_insensitive_and_tolerates_spacing() -> None:
assert _JUDGE_SCORE_RE.search("score:5") is not None
assert _JUDGE_SCORE_RE.search("SCORE : 3") is not None
def test_judge_score_regex_rejects_out_of_range_scores() -> None:
assert _JUDGE_SCORE_RE.search("Score: 0") is None
assert _JUDGE_SCORE_RE.search("Score: 6") is None
def test_build_judge_prompt_includes_the_expectation_and_acceptance_criteria() -> None:
fixture = FIXTURES[0]
prompt = _build_judge_prompt(fixture, diff="+ fixed line", notes="dev notes here")
assert fixture.title in prompt
assert fixture.expectations in prompt
for criterion in fixture.acceptance_criteria:
assert criterion in prompt
assert "+ fixed line" in prompt
assert "dev notes here" in prompt
def test_build_judge_prompt_handles_empty_diff_and_notes() -> None:
fixture = FIXTURES[0]
prompt = _build_judge_prompt(fixture, diff="", notes="")
assert "(empty diff)" in prompt
assert "(no notes)" in prompt
def test_orchestrator_stage_spawner_is_cut_and_refuses_to_construct() -> None:
"""The real-spawn path is deliberately disabled this release (its MCP
wiring would authenticate against the REAL production orchestrator)
this is the one runnable check that the cut stays in place."""
with pytest.raises(NotImplementedError, match="cut from this release"):
OrchestratorStageSpawner()