Files
roboco/tests/e2e_smoke/test_barfly_loop.py
401f8a2cc9 feat(board): Board Programs — the complete twelve-program catalog (Phases 1-3) (#699)
* feat(board): Pest Control — the first project-scoped Board Program

The Product Owner hunts latent defects (what the org records but nobody
reads): a weekly cycle — accelerated off-schedule when the trailing-7-day
rework rate crosses pest_rework_threshold, with the cheap dedup/scope gates
evaluated before the metrics queries — opens one held exploration task
against the least-recently-explored opted-in project (deterministic
round-robin; opted_in_projects gains a stable ORDER BY), with server-
assembled evidence in the spawn prompt (rework hotspots, recurring-findings
and waived-minor ledger aggregates, all capped) plus prior-cycle LEARN
context. The PO calls the new PO-only propose_bug_hunt verb once: ≤5 items,
evidence required per item, targets validated against pest_control
participation. CEO decides per item — approve materializes a BACKLOG task
(source pest_control, never auto-starts), reject records the reason; both
feed the LEARN ledger by exploration task id; all-terminal completes the
cycle. Telegram queue pushes carry working Approve/Reject handlers
mirroring the roadmap kind. Doctrine: board.md Pest Control section +
product-owner verb entry + regenerated verb tables.

* feat(panel): Pest Control review queue

Command Center gains the pest review queue (per-item approve/reject with
reason, mirroring the roadmap queue); the Programs card and the project
settings participates-in checkboxes pick the new program up registry-driven
— the settings section renders for the first time now that a project-scoped
program exists.

* feat(board): Periscope — HoM market-research brief program

Weekly org-scoped cycle: a solo HoM spawn researches the market (web
research with mandatory source URLs — uncited findings are rejected) and
files one structured brief via the new HoM-only propose_market_brief verb:
headline, cited findings, threats/opportunities, positioning note, all
soup-checked and screened through the injection guard at persist time
(web-derived text later reaches prompts; flags recorded, content never
dropped). A brief is a report, not a proposal: the verb completes the
exploration in the same call (the x_feature asymmetry), the cycle ledger
auto-closes, and the CEO gets a best-effort notification with no
approve/reject surface (periscope deliberately never joins Telegram's
action kinds). The latest brief is injected into the roadmap exploration
prompt — Periscope feeds Printer, the first cross-role program input.

* feat(panel): Market Briefs tab (read-only)

Business page gains a Market Briefs tab listing Periscope briefs —
headline, cited findings, threats/opportunities — read-only by design; a
report has nothing to approve.

* feat(board): Coroner — event-triggered Auditor postmortems

The first EVENT program: no cron — three best-effort hooks open an autopsy
when a task bounces to its 3rd revision (the audit chokepoint), is
cancelled after work started, or is budget-blocked; all gated on arming +
one-open-autopsy dedup, none can fail the underlying transition. A solo
Auditor spawn reads the incident (server-assembled findings + transition
context) and files one propose_postmortem: incident summary, root cause,
failed stage (validated against the real status vocabulary), and ONE
process change — a playbook-kind change drafts via PlaybookService
directly into the normal pending-curation queue; the briefed draft_playbook
manifest grant was deliberately NOT added, preserving the existing
'auditor curates but never drafts' invariant test. Complete-at-propose
(report asymmetry), cycle ledger auto-closes, CEO notified link-only.
Integrated as a union with Periscope across the shared program surfaces.

* feat(panel): Coroner postmortems card

Read-only postmortems list under Business → Programs — incident, root
cause, failed stage, process change; nothing to approve, the process-change
artifact (a draft playbook) rides the existing curation queue.

* feat(board): Sentinel — Auditor drift-watch quality reports

Weekly org-scoped cycle: a solo Auditor spawn receives a server-assembled
drift context (waived-findings trend, open findings by severity,
conventions-violation hotspots, top spend — all capped, pure ORM) and files
one propose_quality_report: headline, 1-7 area-validated items with
evidence and suggested actions, overall assessment. Report semantics —
complete-at-propose, cycle auto-closes, CEO notified display-only (never on
Telegram's approve/reject surface); items are structured so a later
convert-to-task control is cheap. Integration adopts Sentinel's module-
level dict-dispatch for board-program routing (xenon-driven), folding all
prior programs in; app router mounting extracted to a helper for the same
budget.

* feat(panel): Quality Reports tab (read-only)

Business page gains the Sentinel quality-reports tab — headline, per-area
observations with evidence and suggested actions; read-only, a report has
nothing to approve.

* feat(board): Spackle — gap-fill audit program

Biweekly project-scoped PO cycle over the half-shipped surface area: API
routes without panel surfaces (and vice versa), armed flags without docs,
docs promises the code doesn't keep, dead-end tabs — the inventory diffing
is the PO's own read-tool work, ordered by the spawn prompt with file:line
citations required; the server injects only prior-cycle LEARN and the
rotation target. Rotation is now a shared module-level helper
(pick_rotation_target, parameterized by source) both project-scoped
engines use — pest_control delegates to it, behavior-identical, with a
cross-pollution test proving the two programs' rotations stay independent.
propose_gap_fill mirrors the bug-hunt verb (≤5 items, two-sided evidence
required, participation gate); per-item CEO decide materializes BACKLOG
source=spackle tasks; full Telegram kind incl. approve/reject handlers.
All seven program routers now mount from one helper.

* feat(panel): Spackle gap-fill review queue

Command Center gains the gap-fill queue mirroring the pest-control one —
per-item approve/reject with the two-sided gap evidence rendered.

* feat(board): Scales — monthly portfolio rebalance

Org-scoped PO cycle over the stale backlog: the spawn receives a capped
stale-task snapshot (BACKLOG/PENDING unclaimed >30 days) plus the charter
and prior-cycle LEARN, and files one propose_rebalance — 1-7 items, each a
resolvable task_ref with action reprioritize (validated new priority) or
cancel, rationale required. Per-item CEO decide: approve EXECUTES the
action (audited priority update, or the normal cancel path) — the first
program whose materializer mutates existing tasks instead of creating
them; reject records the reason; LEARN by exploration task id;
all-terminal completes the cycle. Full Telegram decide-kind wiring.
Integrated as the eight-program union (registry, dict dispatch, routers
helper, teardown enumerations).

* feat(panel): Scales rebalance review queue

Command Center gains the rebalance queue — per-item approve/reject with
the action, target task, and rationale rendered.

* feat(board): Mirror — quarterly positioning audit

Project-scoped HoM cycle over messaging surfaces: README claims vs shipped
reality, docs-site promises vs code, charter alignment — the audit is the
HoM's own read-tool work with citations required; the server injects the
charter, prior-cycle LEARN, and the shared rotation target. propose_
messaging_fixes mirrors the gap-fill verb (≤5 items, drift evidence naming
claim + contradicting reality, participation gate); per-item CEO decide
materializes BACKLOG source=mirror documentation tasks; full Telegram
decide-kind wiring. Nine-program union across the shared surfaces.

* feat(panel): Mirror messaging-fixes review queue

* feat(board): Megaphone — HoM standing editorial calendar

Cron cycle (3 days, org-scoped, gated on X credentials — drafting content
nobody can post is pointless): the HoM receives a shipped-this-week digest
plus Unreleased changelog bullets and files one propose_editorial_post
(angle-validated, ≤280, brand voice) that materializes a held x_editorial
draft through the SAME X-queue origination chokepoint release posts use —
zero new approval surface, notifications and CEO decide for free.
Complete-at-propose; cycle auto-closes. Ten-program union.

* feat(panel): x_editorial source labels in the X queue surfaces

* feat(board): Librarian — proactive playbook mining

Biweekly org-scoped Auditor cycle: mines recurring non-private learning
journals (≥2-count grouping with a recency fallback) against the existing
playbook-title inventory and files one propose_playbook_drafts — 1-3
drafts, each with the repeated-pattern evidence that justifies it,
duplicate titles rejected in-batch and against the live store. Drafts are
created via PlaybookService directly (the Coroner precedent — the
'auditor curates but never drafts' do-verb invariant stays intact and
tested) and land in the normal pending-curation queue the Auditor's own
triage already surfaces; no new panel surface. Complete-at-propose;
display-only CEO notification. Eleven-program union.

* feat(board): War Room — release campaign planning

EVENT program with a REAL originator (unlike coroner's stub): a release
publish hooks a campaign brief beside the release-post seam, and the CEO's
run-now originates on demand — the cron loop never fires it. The HoM
designs a 2-6 post arc (teaser → launch → follow-up → spotlight; 280-cap,
future strictly-ascending publish_after, stage vocabulary) and one
propose_campaign call materializes each post as a held x_campaign draft
through the X-queue chokepoint. V1 is manual-cadence by design: publish_
after renders as queue guidance and the CEO approves each post at its
moment — nothing auto-posts, ever; the auto-schedule upgrade is a
documented ceiling. Twelve-program union: full registry complete.

* feat(panel): x_campaign labels + publish-after guidance in the X queue

* feat(board): Barfly — adjacent-conversation replies

Cron cycle (2 days, org-scoped, X-credentials gated): the engine searches
X for conversations where RoboCo is relevant but unmentioned (new OAuth-
signed search_recent on the client; queries + candidate cap configurable),
screens every fetched tweet through the injection guard (stored unclamped
— a clamp was truncating the candidate under the envelope, caught by the
dev's own tests), dedupes via the existing x_seen_mentions ledger (no
migration; also prevents double-drafting against the mentions poll), and
opens one held HoM exploration carrying the screened candidates. propose_
conversation_replies enforces candidate-id-only replies (≤5, 280-cap);
each materializes a held x_barfly draft through the X-queue chokepoint,
threaded via a new in_reply_to seam on post_tweet that only x_barfly
drafts use. The X redraft machinery is now dict-dispatch over per-source
extractors with reply-ref carry for x_barfly. Thirteen-program registry.
War Room's test fakes gained the new abstract search_recent stub.

* feat(board): Dogfood — the PO walks the product

The fourteenth and final registry entry, completing the catalog. EVENT
program (release-publish hook beside the war-room hook + CEO run-now, both
through the same real originator; the cron loop never fires it), project-
scoped with shared rotation. The permission surface is the careful part:
the PO's dogfood spawn — and ONLY that spawn — gets the Playwright MCP
mounted, via a task-scoped fail-closed probe mirroring the video-authoring
precedent (a PO spawned for roadmap/pest/scales never sees browser tools;
tested both ways); the PM agent image bakes chromium unconditionally like
the ux image, the mount stays task-gated in code. The walk targets the
rotation target's live surfaces (panel_base_url only when the target is
the org's own project, honest degradation otherwise); propose_friction_
fixes files ≤5 walked-path-evidenced items; per-item CEO decide
materializes BACKLOG source=dogfood tasks; full Telegram decide kind.
Also: megaphone/librarian/war_room arming keys restored to the settings
validator — their panel toggles would have been rejected (dropped in
earlier unions; the same silent-arming class the drill killed once
already).

* feat(panel): Dogfood friction review queue

* chore(board): final whole-branch sweep fixes

The night's closing adversarial pass over the integrated fourteen-program
registry found ONE functional defect — the war-room test fakes' post_tweet
predated Barfly's in_reply_to_tweet_id kwarg (LSP violation, the only red
in an otherwise fully green gate) — plus doc/test drift, all fixed: the
source-parity test completes to fourteen (spackle/mirror were silently
absent while its neighboring comment claimed full coverage), the PO
identity doc gains its missing Dogfood verb, the auditor quick-list gains
propose_postmortem, three stale comments corrected (rotation docstring,
panel registry header, X source enumerations), the dogfood release-hook
gains the exception-swallow test its four sibling hooks already had, and
the CHANGELOG's Unreleased section documents the whole Board Programs
train. Full make quality: exit 0, all gates green.

* docs: full documentation sweep for the Board Programs train

CLAUDE.md's roadmap-engine entry superseded by the Board Program registry
entry (all fourteen programs, arming, scoping, LEARN, guardrails) with the
role verb tables and playwright row refreshed; docs/rag gains the agent-
facing architecture doc plus full propose_* call-shape sections in the
three board role docs, and corrects the strategy-engine section to shipped
reality (only idle→roadmap is wired); docs/map covers the registry + all
twelve engines with flags, gotchas, and drift notes. The 0.27.0 reference
inventory confirmed only the release-executor's canonical set carries the
version — left for the 0.28.0 cut.

* feat(board): human titles + descriptions on every program surface

Raw registry keys rendered as bare panel labels — an operator reading
x_feature had no idea what enabling or running it does. The registry
dataclass gains title/description (test-enforced non-empty for every
entry, unique titles), the API passes them through, and every surface
renders title-with-description-tooltip instead of the key: the Programs
card (label, toggle hint, run-now toast), and the project settings
participates-in/excluded-from checkboxes.

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-07-25 17:13:32 +02:00

368 lines
13 KiB
Python

"""Scenario: the Barfly (Board Program) loop end to end.
Mirrors test_periscope_loop.py's arm -> originate -> dedup shape AND
test_feature_spotlight.py's do_server wiring-regression check + real-route
propose call. The genuinely new pieces: (1) origination needs a configured X
client (search_recent), unlike periscope which needs none — faked here via
``roboco.services.barfly_engine.build_x_client`` (uvicorn runs in-thread, same
process, so the patch reaches the live route handler too); (2) unlike a
market brief (one report, completes its own task) a Barfly proposal
materializes N SEPARATE held x_barfly drafts through the shared
_originate_post chokepoint while completing the exploration task itself —
the x_feature-style asymmetry, multiplied.
"""
from __future__ import annotations
from typing import TYPE_CHECKING, Any
from unittest.mock import patch
from uuid import UUID, uuid4
from roboco.foundation import identity as _foundation
from tests.e2e_smoke.harness import ScriptedAgent, expect_ok
if TYPE_CHECKING:
from sqlalchemy.ext.asyncio import AsyncSession
from tests.e2e_smoke.harness import E2EStack
ONE = 1
TWO = 2
ZERO = 0
_CANDIDATES: list[dict[str, Any]] = [
{"id": "e2e-1", "text": "we should try a multi-agent coding org", "likes": 3},
{"id": "e2e-2", "text": "autonomous software teams are the future", "likes": 1},
]
# A second, disjoint id set for the reopen check below — the first batch's
# ids are already in the seen ledger by then, so reusing them would starve
# the reopened cycle of any surviving candidate.
_CANDIDATES_ROUND_2: list[dict[str, Any]] = [
{
"id": "e2e-3",
"text": "curious how agent teams handle merge conflicts",
"likes": 2,
},
]
class _FakeSearchClient:
"""Configured stub returned by the patched ``build_x_client`` — never
touches the network."""
def __init__(self, candidates: list[dict[str, Any]] | None = None) -> None:
self._candidates = candidates if candidates is not None else _CANDIDATES
@property
def configured(self) -> bool:
return True
async def search_recent(self, query: str, max_results: int) -> list[Any]:
from roboco.services.x_client import XMention
_ = (query, max_results)
return [
XMention(
id=c["id"],
author_id=f"author-{c['id']}",
text=c["text"],
like_count=c["likes"],
reply_count=0,
retweet_count=0,
)
for c in self._candidates
]
async def post_tweet(self, text: str, **_kwargs: Any) -> Any:
raise AssertionError(f"this e2e scenario never approves/posts: {text!r}")
async def fetch_mentions(self, since_id: str | None, max_results: int) -> list[Any]:
_ = (since_id, max_results)
return []
def _seed_system_and_hom(stack: E2EStack) -> str:
"""Seed ``system`` + ``head-marketing`` + ``secretary-1`` at their FIXED
foundation UUIDs (BarflyEngine/_originate_post write created_by/
assigned_to straight from the static identity registry), plus a project
those own — the exploration task's FK anchor. Returns the project's slug.
"""
from roboco.db.tables import AgentTable, ProjectTable
from roboco.models import AgentRole, AgentStatus, Team
slug = f"e2e-barfly-{uuid4().hex[:8]}"
async def _run(session: AsyncSession) -> None:
for agent_uuid, agent_slug, role, team in (
(_foundation.AGENTS["system"].uuid, "system", AgentRole.SYSTEM, None),
(
_foundation.AGENTS["head-marketing"].uuid,
"head-marketing",
AgentRole.HEAD_MARKETING,
Team.BOARD,
),
(
_foundation.AGENTS["secretary-1"].uuid,
"secretary-1",
AgentRole.SECRETARY,
None,
),
):
if await session.get(AgentTable, agent_uuid) is not None:
continue
session.add(
AgentTable(
id=agent_uuid,
name=agent_slug,
slug=agent_slug,
role=role,
team=team,
status=AgentStatus.ACTIVE,
model_config={},
system_prompt=agent_slug,
capabilities=[],
permissions={},
metrics={},
)
)
session.add(
ProjectTable(
id=uuid4(),
name="RoboCo",
slug=slug,
git_url="https://example.com/roboco.git",
default_branch="master",
protected_branches=["master"],
assigned_cell=Team.BACKEND,
created_by=_foundation.AGENTS["system"].uuid,
is_active=True,
)
)
stack.run_db(_run)
return slug
def _arm(stack: E2EStack, project_slug: str) -> None:
"""Arm via the settings-store key (the ONLY arming path — no legacy env
flag exists for barfly) + stored X credentials (BarflyEngine's creds
gate) + point ``self_heal_project_slug`` at the seeded project."""
from roboco.config import settings as cfg
from roboco.db.tables import SystemSettingTable
from roboco.services.x_credentials import get_x_credentials_service
cfg.self_heal_project_slug = project_slug
async def _run(session: AsyncSession) -> None:
session.add(
SystemSettingTable(key="board_program.barfly.enabled", value="true")
)
await get_x_credentials_service(session).set_credentials(
api_key="k",
api_secret="s",
access_token="t",
access_token_secret="ts",
)
stack.run_db(_run)
def _run_due_programs(stack: E2EStack) -> list[str]:
from roboco.services.board_programs import get_board_program_engine
async def _run(session: AsyncSession) -> list[str]:
return await get_board_program_engine(session).run_due_programs()
result: list[str] = stack.run_db(_run)
return result
def _find_barfly_task(stack: E2EStack) -> dict[str, Any]:
from roboco.db.tables import TaskTable
from roboco.foundation.policy.content import markers
from roboco.services.task import BARFLY_SOURCE
from sqlalchemy import select
async def _run(session: AsyncSession) -> dict[str, Any]:
rows = (
(
await session.execute(
select(TaskTable).where(TaskTable.source == BARFLY_SOURCE)
)
)
.scalars()
.all()
)
return {
"count": len(rows),
"rows": [
{
"id": r.id,
"status": str(r.status),
"assigned_to": r.assigned_to,
"source": r.source,
"project_id": r.project_id,
"confirmed_by_human": r.confirmed_by_human,
"candidates": markers.get_barfly_candidates(r),
}
for r in rows
],
}
state: dict[str, Any] = stack.run_db(_run)
return state
def _x_barfly_drafts(stack: E2EStack) -> list[dict[str, Any]]:
from roboco.db.tables import TaskTable
from roboco.foundation.policy.content import markers
from sqlalchemy import select
async def _run(session: AsyncSession) -> list[dict[str, Any]]:
rows = (
(
await session.execute(
select(TaskTable).where(TaskTable.source == "x_barfly")
)
)
.scalars()
.all()
)
return [
{
"id": r.id,
"status": str(r.status),
"confirmed_by_human": r.confirmed_by_human,
"body": markers.get_x_draft_body(r),
"reply_ref": markers.get_barfly_reply_ref(r),
}
for r in rows
]
result: list[dict[str, Any]] = stack.run_db(_run)
return result
def test_barfly_loop_originates_dedups_proposes_and_materializes(
e2e_stack: E2EStack,
) -> None:
stack = e2e_stack
project_slug = _seed_system_and_hom(stack)
_arm(stack, project_slug)
with patch(
"roboco.services.barfly_engine.build_x_client",
return_value=_FakeSearchClient(),
):
opened = _run_due_programs(stack)
assert opened == ["barfly"], opened
state = _find_barfly_task(stack)
assert state["count"] == ONE, state
row = state["rows"][0]
assert row["status"] == "pending"
assert row["assigned_to"] == _foundation.AGENTS["head-marketing"].uuid
assert row["confirmed_by_human"] is False
assert row["project_id"] is not None
candidate_ids = {c["id"] for c in row["candidates"]}
assert candidate_ids == {"e2e-1", "e2e-2"}
# board_barfly is board-dispatched (one-shot HoM spawn), never handed to
# the generic dev dispatch loop's give_me_work/claim path.
from roboco.runtime.orchestrator import _is_non_dev_dispatch_source
assert _is_non_dev_dispatch_source({"source": row["source"]}) is True
# Second tick — a client whose configured check passes but whose
# search_recent errors if ever called: the open-cycle dedup must block
# BEFORE a real search runs, even though the creds/configured check
# itself still runs ahead of dedup (mirrors XEngine's own ordering).
class _BoomIfSearched(_FakeSearchClient):
async def search_recent(self, query: str, max_results: int) -> list[Any]:
raise AssertionError(
f"must not search (query={query!r}, max_results={max_results}) "
"while dedup-blocked"
)
with patch(
"roboco.services.barfly_engine.build_x_client",
return_value=_BoomIfSearched(),
):
opened_again = _run_due_programs(stack)
assert opened_again == [], opened_again
state_after = _find_barfly_task(stack)
assert state_after["count"] == ONE, state_after
# The Head of Marketing authors replies through the REAL do_server module
# + /api/v1/do/propose_conversation_replies route + ContentActions +
# XEngine wiring — the wiring-regression class check test_feature_
# spotlight.py guards (a verb granted in role_config but missing from
# do_server._TOOLS is unreachable over MCP).
hom = ScriptedAgent(
stack,
_foundation.AGENTS["head-marketing"].uuid,
"head-marketing",
"head_marketing",
)
do_module = hom._module("roboco.mcp.do_server")
assert "propose_conversation_replies" in do_module._TOOLS, (
"propose_conversation_replies missing from do_server._TOOLS — the "
"MCP server has no way to expose it to any role"
)
assert "propose_conversation_replies" in do_module._REGISTERED_TOOLS, (
"propose_conversation_replies is granted to head_marketing in "
"role_config but absent from this agent's _register_tools() output "
"— the manifest -> _register_tools -> callable chain dropped it"
)
env = expect_ok(
hom.do(
"propose_conversation_replies",
items=[
{
"tweet_id": "e2e-1",
"reply_body": "That's exactly what request_sandbox() gives you.",
"rationale": "Directly answers what they're describing.",
},
{
"tweet_id": "e2e-2",
"reply_body": "We built exactly that — 25 agents, one CEO.",
"rationale": "Speaks directly to the premise of their post.",
},
],
),
"hom propose_conversation_replies",
)
assert env.get("status") == "conversation_replies_proposed", env
assert env.get("task_id") == str(row["id"])
materialized_ids = env.get("context_briefing", {}).get("materialized_task_ids", [])
assert len(materialized_ids) == TWO, env
from tests.e2e_smoke.arcs import task_state
exploration_id = UUID(env["task_id"])
assert task_state(stack, exploration_id)["status"] == "completed"
drafts = _x_barfly_drafts(stack)
assert len(drafts) == TWO, drafts
tweet_ids = {d["reply_ref"]["tweet_id"] for d in drafts}
assert tweet_ids == {"e2e-1", "e2e-2"}
for draft in drafts:
assert draft["status"] == "pending"
assert draft["confirmed_by_human"] is False # HELD; the CEO decides
assert draft["body"]
# The exploration going terminal auto-closes the LEARN ledger row — a
# fresh cycle can open off-schedule (enabled + dedup only) proving it.
from roboco.services.board_programs import get_board_program_engine
async def _reopen(session: AsyncSession) -> Any:
return await get_board_program_engine(session).open_program_cycle("barfly")
with patch(
"roboco.services.barfly_engine.build_x_client",
return_value=_FakeSearchClient(_CANDIDATES_ROUND_2),
):
reopened_id = stack.run_db(_reopen)
assert reopened_id is not None
assert reopened_id != row["id"]