Files
roboco/tests/unit/services/test_coroner_engine.py
T
401f8a2cc9 feat(board): Board Programs — the complete twelve-program catalog (Phases 1-3) (#699)
* feat(board): Pest Control — the first project-scoped Board Program

The Product Owner hunts latent defects (what the org records but nobody
reads): a weekly cycle — accelerated off-schedule when the trailing-7-day
rework rate crosses pest_rework_threshold, with the cheap dedup/scope gates
evaluated before the metrics queries — opens one held exploration task
against the least-recently-explored opted-in project (deterministic
round-robin; opted_in_projects gains a stable ORDER BY), with server-
assembled evidence in the spawn prompt (rework hotspots, recurring-findings
and waived-minor ledger aggregates, all capped) plus prior-cycle LEARN
context. The PO calls the new PO-only propose_bug_hunt verb once: ≤5 items,
evidence required per item, targets validated against pest_control
participation. CEO decides per item — approve materializes a BACKLOG task
(source pest_control, never auto-starts), reject records the reason; both
feed the LEARN ledger by exploration task id; all-terminal completes the
cycle. Telegram queue pushes carry working Approve/Reject handlers
mirroring the roadmap kind. Doctrine: board.md Pest Control section +
product-owner verb entry + regenerated verb tables.

* feat(panel): Pest Control review queue

Command Center gains the pest review queue (per-item approve/reject with
reason, mirroring the roadmap queue); the Programs card and the project
settings participates-in checkboxes pick the new program up registry-driven
— the settings section renders for the first time now that a project-scoped
program exists.

* feat(board): Periscope — HoM market-research brief program

Weekly org-scoped cycle: a solo HoM spawn researches the market (web
research with mandatory source URLs — uncited findings are rejected) and
files one structured brief via the new HoM-only propose_market_brief verb:
headline, cited findings, threats/opportunities, positioning note, all
soup-checked and screened through the injection guard at persist time
(web-derived text later reaches prompts; flags recorded, content never
dropped). A brief is a report, not a proposal: the verb completes the
exploration in the same call (the x_feature asymmetry), the cycle ledger
auto-closes, and the CEO gets a best-effort notification with no
approve/reject surface (periscope deliberately never joins Telegram's
action kinds). The latest brief is injected into the roadmap exploration
prompt — Periscope feeds Printer, the first cross-role program input.

* feat(panel): Market Briefs tab (read-only)

Business page gains a Market Briefs tab listing Periscope briefs —
headline, cited findings, threats/opportunities — read-only by design; a
report has nothing to approve.

* feat(board): Coroner — event-triggered Auditor postmortems

The first EVENT program: no cron — three best-effort hooks open an autopsy
when a task bounces to its 3rd revision (the audit chokepoint), is
cancelled after work started, or is budget-blocked; all gated on arming +
one-open-autopsy dedup, none can fail the underlying transition. A solo
Auditor spawn reads the incident (server-assembled findings + transition
context) and files one propose_postmortem: incident summary, root cause,
failed stage (validated against the real status vocabulary), and ONE
process change — a playbook-kind change drafts via PlaybookService
directly into the normal pending-curation queue; the briefed draft_playbook
manifest grant was deliberately NOT added, preserving the existing
'auditor curates but never drafts' invariant test. Complete-at-propose
(report asymmetry), cycle ledger auto-closes, CEO notified link-only.
Integrated as a union with Periscope across the shared program surfaces.

* feat(panel): Coroner postmortems card

Read-only postmortems list under Business → Programs — incident, root
cause, failed stage, process change; nothing to approve, the process-change
artifact (a draft playbook) rides the existing curation queue.

* feat(board): Sentinel — Auditor drift-watch quality reports

Weekly org-scoped cycle: a solo Auditor spawn receives a server-assembled
drift context (waived-findings trend, open findings by severity,
conventions-violation hotspots, top spend — all capped, pure ORM) and files
one propose_quality_report: headline, 1-7 area-validated items with
evidence and suggested actions, overall assessment. Report semantics —
complete-at-propose, cycle auto-closes, CEO notified display-only (never on
Telegram's approve/reject surface); items are structured so a later
convert-to-task control is cheap. Integration adopts Sentinel's module-
level dict-dispatch for board-program routing (xenon-driven), folding all
prior programs in; app router mounting extracted to a helper for the same
budget.

* feat(panel): Quality Reports tab (read-only)

Business page gains the Sentinel quality-reports tab — headline, per-area
observations with evidence and suggested actions; read-only, a report has
nothing to approve.

* feat(board): Spackle — gap-fill audit program

Biweekly project-scoped PO cycle over the half-shipped surface area: API
routes without panel surfaces (and vice versa), armed flags without docs,
docs promises the code doesn't keep, dead-end tabs — the inventory diffing
is the PO's own read-tool work, ordered by the spawn prompt with file:line
citations required; the server injects only prior-cycle LEARN and the
rotation target. Rotation is now a shared module-level helper
(pick_rotation_target, parameterized by source) both project-scoped
engines use — pest_control delegates to it, behavior-identical, with a
cross-pollution test proving the two programs' rotations stay independent.
propose_gap_fill mirrors the bug-hunt verb (≤5 items, two-sided evidence
required, participation gate); per-item CEO decide materializes BACKLOG
source=spackle tasks; full Telegram kind incl. approve/reject handlers.
All seven program routers now mount from one helper.

* feat(panel): Spackle gap-fill review queue

Command Center gains the gap-fill queue mirroring the pest-control one —
per-item approve/reject with the two-sided gap evidence rendered.

* feat(board): Scales — monthly portfolio rebalance

Org-scoped PO cycle over the stale backlog: the spawn receives a capped
stale-task snapshot (BACKLOG/PENDING unclaimed >30 days) plus the charter
and prior-cycle LEARN, and files one propose_rebalance — 1-7 items, each a
resolvable task_ref with action reprioritize (validated new priority) or
cancel, rationale required. Per-item CEO decide: approve EXECUTES the
action (audited priority update, or the normal cancel path) — the first
program whose materializer mutates existing tasks instead of creating
them; reject records the reason; LEARN by exploration task id;
all-terminal completes the cycle. Full Telegram decide-kind wiring.
Integrated as the eight-program union (registry, dict dispatch, routers
helper, teardown enumerations).

* feat(panel): Scales rebalance review queue

Command Center gains the rebalance queue — per-item approve/reject with
the action, target task, and rationale rendered.

* feat(board): Mirror — quarterly positioning audit

Project-scoped HoM cycle over messaging surfaces: README claims vs shipped
reality, docs-site promises vs code, charter alignment — the audit is the
HoM's own read-tool work with citations required; the server injects the
charter, prior-cycle LEARN, and the shared rotation target. propose_
messaging_fixes mirrors the gap-fill verb (≤5 items, drift evidence naming
claim + contradicting reality, participation gate); per-item CEO decide
materializes BACKLOG source=mirror documentation tasks; full Telegram
decide-kind wiring. Nine-program union across the shared surfaces.

* feat(panel): Mirror messaging-fixes review queue

* feat(board): Megaphone — HoM standing editorial calendar

Cron cycle (3 days, org-scoped, gated on X credentials — drafting content
nobody can post is pointless): the HoM receives a shipped-this-week digest
plus Unreleased changelog bullets and files one propose_editorial_post
(angle-validated, ≤280, brand voice) that materializes a held x_editorial
draft through the SAME X-queue origination chokepoint release posts use —
zero new approval surface, notifications and CEO decide for free.
Complete-at-propose; cycle auto-closes. Ten-program union.

* feat(panel): x_editorial source labels in the X queue surfaces

* feat(board): Librarian — proactive playbook mining

Biweekly org-scoped Auditor cycle: mines recurring non-private learning
journals (≥2-count grouping with a recency fallback) against the existing
playbook-title inventory and files one propose_playbook_drafts — 1-3
drafts, each with the repeated-pattern evidence that justifies it,
duplicate titles rejected in-batch and against the live store. Drafts are
created via PlaybookService directly (the Coroner precedent — the
'auditor curates but never drafts' do-verb invariant stays intact and
tested) and land in the normal pending-curation queue the Auditor's own
triage already surfaces; no new panel surface. Complete-at-propose;
display-only CEO notification. Eleven-program union.

* feat(board): War Room — release campaign planning

EVENT program with a REAL originator (unlike coroner's stub): a release
publish hooks a campaign brief beside the release-post seam, and the CEO's
run-now originates on demand — the cron loop never fires it. The HoM
designs a 2-6 post arc (teaser → launch → follow-up → spotlight; 280-cap,
future strictly-ascending publish_after, stage vocabulary) and one
propose_campaign call materializes each post as a held x_campaign draft
through the X-queue chokepoint. V1 is manual-cadence by design: publish_
after renders as queue guidance and the CEO approves each post at its
moment — nothing auto-posts, ever; the auto-schedule upgrade is a
documented ceiling. Twelve-program union: full registry complete.

* feat(panel): x_campaign labels + publish-after guidance in the X queue

* feat(board): Barfly — adjacent-conversation replies

Cron cycle (2 days, org-scoped, X-credentials gated): the engine searches
X for conversations where RoboCo is relevant but unmentioned (new OAuth-
signed search_recent on the client; queries + candidate cap configurable),
screens every fetched tweet through the injection guard (stored unclamped
— a clamp was truncating the candidate under the envelope, caught by the
dev's own tests), dedupes via the existing x_seen_mentions ledger (no
migration; also prevents double-drafting against the mentions poll), and
opens one held HoM exploration carrying the screened candidates. propose_
conversation_replies enforces candidate-id-only replies (≤5, 280-cap);
each materializes a held x_barfly draft through the X-queue chokepoint,
threaded via a new in_reply_to seam on post_tweet that only x_barfly
drafts use. The X redraft machinery is now dict-dispatch over per-source
extractors with reply-ref carry for x_barfly. Thirteen-program registry.
War Room's test fakes gained the new abstract search_recent stub.

* feat(board): Dogfood — the PO walks the product

The fourteenth and final registry entry, completing the catalog. EVENT
program (release-publish hook beside the war-room hook + CEO run-now, both
through the same real originator; the cron loop never fires it), project-
scoped with shared rotation. The permission surface is the careful part:
the PO's dogfood spawn — and ONLY that spawn — gets the Playwright MCP
mounted, via a task-scoped fail-closed probe mirroring the video-authoring
precedent (a PO spawned for roadmap/pest/scales never sees browser tools;
tested both ways); the PM agent image bakes chromium unconditionally like
the ux image, the mount stays task-gated in code. The walk targets the
rotation target's live surfaces (panel_base_url only when the target is
the org's own project, honest degradation otherwise); propose_friction_
fixes files ≤5 walked-path-evidenced items; per-item CEO decide
materializes BACKLOG source=dogfood tasks; full Telegram decide kind.
Also: megaphone/librarian/war_room arming keys restored to the settings
validator — their panel toggles would have been rejected (dropped in
earlier unions; the same silent-arming class the drill killed once
already).

* feat(panel): Dogfood friction review queue

* chore(board): final whole-branch sweep fixes

The night's closing adversarial pass over the integrated fourteen-program
registry found ONE functional defect — the war-room test fakes' post_tweet
predated Barfly's in_reply_to_tweet_id kwarg (LSP violation, the only red
in an otherwise fully green gate) — plus doc/test drift, all fixed: the
source-parity test completes to fourteen (spackle/mirror were silently
absent while its neighboring comment claimed full coverage), the PO
identity doc gains its missing Dogfood verb, the auditor quick-list gains
propose_postmortem, three stale comments corrected (rotation docstring,
panel registry header, X source enumerations), the dogfood release-hook
gains the exception-swallow test its four sibling hooks already had, and
the CHANGELOG's Unreleased section documents the whole Board Programs
train. Full make quality: exit 0, all gates green.

* docs: full documentation sweep for the Board Programs train

CLAUDE.md's roadmap-engine entry superseded by the Board Program registry
entry (all fourteen programs, arming, scoping, LEARN, guardrails) with the
role verb tables and playwright row refreshed; docs/rag gains the agent-
facing architecture doc plus full propose_* call-shape sections in the
three board role docs, and corrects the strategy-engine section to shipped
reality (only idle→roadmap is wired); docs/map covers the registry + all
twelve engines with flags, gotchas, and drift notes. The 0.27.0 reference
inventory confirmed only the release-executor's canonical set carries the
version — left for the 0.28.0 cut.

* feat(board): human titles + descriptions on every program surface

Raw registry keys rendered as bare panel labels — an operator reading
x_feature had no idea what enabling or running it does. The registry
dataclass gains title/description (test-enforced non-empty for every
entry, unique titles), the API passes them through, and every surface
renders title-with-description-tooltip instead of the key: the Programs
card (label, toggle hint, run-now toast), and the project settings
participates-in/excluded-from checkboxes.

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-07-25 17:13:32 +02:00

283 lines
9.2 KiB
Python

"""Coroner engine: EVENT-triggered autopsy origination, deduped, never
authors content itself, and (unlike Pest Control) org-scoped — no
per-project opt-in gate. Mirrors test_pest_control_engine.py's shape.
"""
from __future__ import annotations
from typing import TYPE_CHECKING, cast
from uuid import uuid4
import pytest
import pytest_asyncio
from roboco.db.tables import (
AgentTable,
AuditLogTable,
BoardProgramCycleTable,
ProjectTable,
SystemSettingTable,
TaskReviewFindingTable,
TaskTable,
)
from roboco.foundation import identity as _foundation
from roboco.foundation.policy.content import markers
from roboco.models.base import AgentRole, AgentStatus, Complexity, Team
from roboco.models.base import TaskNature as TN
from roboco.models.base import TaskStatus as TS
from roboco.models.base import TaskType as TT
from roboco.services.coroner_engine import CoronerEngine
from roboco.services.task import (
CORONER_SOURCE,
PEST_CONTROL_SOURCE,
ROADMAP_SOURCE,
X_FEATURE_EXPLORATION_SOURCE,
TaskCreateRequest,
get_task_service,
)
from sqlalchemy import delete, select, update
if TYPE_CHECKING:
from uuid import UUID
from sqlalchemy.ext.asyncio import AsyncSession
SYSTEM_UUID = _foundation.AGENTS["system"].uuid
AUDITOR_UUID = _foundation.AGENTS["auditor"].uuid
SLUG = "backend-svc"
ONE = 1
@pytest_asyncio.fixture(autouse=True)
async def _purge_board_program_pollution(db_session: AsyncSession) -> None:
"""See test_board_program_engine.py's identical fixture: Board Program
settings-store rows / ledger rows / open exploration tasks are shared,
cross-test-persistent DB state. Purge before every test in this file."""
await db_session.execute(
delete(SystemSettingTable).where(SystemSettingTable.key.like("board_program.%"))
)
await db_session.execute(delete(BoardProgramCycleTable))
await db_session.execute(
update(TaskTable)
.where(
TaskTable.source.in_(
[
ROADMAP_SOURCE,
X_FEATURE_EXPLORATION_SOURCE,
PEST_CONTROL_SOURCE,
CORONER_SOURCE,
]
),
TaskTable.status.notin_([TS.COMPLETED, TS.CANCELLED]),
)
.values(status=TS.CANCELLED)
)
await db_session.commit()
async def _seed(session: AsyncSession) -> ProjectTable:
for uuid, slug, role, team in (
(SYSTEM_UUID, "system", AgentRole.SYSTEM, None),
(AUDITOR_UUID, "auditor", AgentRole.AUDITOR, Team.BOARD),
):
if await session.get(AgentTable, uuid) is None:
session.add(
AgentTable(
id=uuid,
name=slug,
slug=slug,
role=role,
team=team,
status=AgentStatus.ACTIVE,
model_config={},
system_prompt="x",
capabilities=[],
permissions={},
metrics={},
)
)
await session.flush()
project = ProjectTable(
name="Backend Service",
slug=SLUG,
git_url="https://github.com/x/backend-svc.git",
default_branch="master",
protected_branches=["master"],
assigned_cell=Team.BACKEND,
created_by=SYSTEM_UUID,
is_active=True,
)
session.add(project)
await session.flush()
return project
async def _incident(session: AsyncSession, project: ProjectTable) -> TaskTable:
"""A raw delivery task standing in for a bounced/cancelled/budget-blocked
incident — created PENDING, then raw-set to a terminal-ish status (never
routed through the transition validator, mirroring test_board_program_
engine.py's ``_make_exploration`` helper), since CoronerEngine only reads
the task, never transitions it."""
task = await get_task_service(session).create(
TaskCreateRequest(
title="Chronic task",
description="Bounced repeatedly",
acceptance_criteria=["it works"],
team=Team.BACKEND,
created_by=SYSTEM_UUID,
task_type=TT.CODE,
nature=TN.TECHNICAL,
estimated_complexity=Complexity.LOW,
project_id=cast("UUID", project.id),
status=TS.PENDING,
)
)
task.status = TS.NEEDS_REVISION
task.revision_count = 3
await session.flush()
return task
def _arm(session: AsyncSession) -> None:
session.add(SystemSettingTable(key="board_program.coroner.enabled", value="true"))
@pytest.mark.asyncio
async def test_disabled_creates_no_cycle(db_session: AsyncSession) -> None:
project = await _seed(db_session)
incident = await _incident(db_session, project)
await db_session.flush()
engine = CoronerEngine(db_session)
assert (
await engine.open_for_incident(cast("UUID", incident.id), kind="bounced")
is None
)
assert await get_task_service(db_session).list_open_coroner_cycles() == []
@pytest.mark.asyncio
async def test_enabled_originates_held_autopsy_task(db_session: AsyncSession) -> None:
project = await _seed(db_session)
incident = await _incident(db_session, project)
_arm(db_session)
await db_session.flush()
engine = CoronerEngine(db_session)
task = await engine.open_for_incident(cast("UUID", incident.id), kind="bounced")
assert task is not None
assert task.source == CORONER_SOURCE
assert task.status == TS.PENDING
assert task.assigned_to == AUDITOR_UUID
assert task.confirmed_by_human is False
ref = markers.get_coroner_incident(task)
assert ref is not None
assert ref["incident_task_id"] == str(incident.id)
assert ref["kind"] == "bounced"
rows = (await db_session.execute(select(BoardProgramCycleTable))).scalars().all()
assert len(rows) == ONE
assert rows[0].program_key == "coroner"
assert rows[0].exploration_task_id == task.id
@pytest.mark.asyncio
async def test_open_autopsy_blocks_a_second_incident(db_session: AsyncSession) -> None:
project = await _seed(db_session)
first_incident = await _incident(db_session, project)
second_incident = await _incident(db_session, project)
_arm(db_session)
await db_session.flush()
engine = CoronerEngine(db_session)
first_task = await engine.open_for_incident(
cast("UUID", first_incident.id), kind="bounced"
)
assert first_task is not None
second_task = await engine.open_for_incident(
cast("UUID", second_incident.id), kind="cancelled"
)
assert second_task is None
cycles = await get_task_service(db_session).list_open_coroner_cycles()
assert len(cycles) == ONE
@pytest.mark.asyncio
async def test_unresolvable_incident_returns_none(db_session: AsyncSession) -> None:
await _seed(db_session)
_arm(db_session)
await db_session.flush()
engine = CoronerEngine(db_session)
assert await engine.open_for_incident(uuid4(), kind="bounced") is None
@pytest.mark.asyncio
async def test_incident_context_renders_findings_and_transitions(
db_session: AsyncSession,
) -> None:
project = await _seed(db_session)
incident = await _incident(db_session, project)
await db_session.flush()
db_session.add(
TaskReviewFindingTable(
task_id=incident.id,
origin="qa",
round=1,
author_slug="be-qa",
file="roboco/app.py",
line=42,
severity="blocker",
criterion="AC 1",
expected="handles the edge case",
actual="raises",
fix="add a guard",
evidence="repro steps",
)
)
db_session.add(
AuditLogTable(
event_type="task.needs_revision",
target_type="task",
target_id=incident.id,
severity="info",
details={},
)
)
await db_session.flush()
engine = CoronerEngine(db_session)
context = await engine.incident_context(cast("UUID", incident.id))
assert "roboco/app.py:42" in context
assert "AC 1" in context
assert "task.needs_revision" in context
@pytest.mark.asyncio
async def test_incident_context_empty_with_no_history(db_session: AsyncSession) -> None:
project = await _seed(db_session)
incident = await _incident(db_session, project)
await db_session.flush()
engine = CoronerEngine(db_session)
assert await engine.incident_context(cast("UUID", incident.id)) == ""
@pytest.mark.asyncio
async def test_complete_with_postmortem_completes_the_task(
db_session: AsyncSession,
) -> None:
project = await _seed(db_session)
incident = await _incident(db_session, project)
_arm(db_session)
await db_session.flush()
engine = CoronerEngine(db_session)
task = await engine.open_for_incident(cast("UUID", incident.id), kind="bounced")
assert task is not None
payload = {
"incident_summary": "it bounced repeatedly",
"root_cause": "no venv-freshness check",
"failed_stage": "awaiting_qa",
"process_change": {"kind": "conventions_rule", "description": "add a check"},
"playbook_id": None,
}
await engine.complete_with_postmortem(task, payload)
assert task.status == TS.COMPLETED
assert markers.get_coroner_postmortem(task) == payload