Files
roboco/tests/unit/services/test_scales_service.py
T
401f8a2cc9 feat(board): Board Programs — the complete twelve-program catalog (Phases 1-3) (#699)
* feat(board): Pest Control — the first project-scoped Board Program

The Product Owner hunts latent defects (what the org records but nobody
reads): a weekly cycle — accelerated off-schedule when the trailing-7-day
rework rate crosses pest_rework_threshold, with the cheap dedup/scope gates
evaluated before the metrics queries — opens one held exploration task
against the least-recently-explored opted-in project (deterministic
round-robin; opted_in_projects gains a stable ORDER BY), with server-
assembled evidence in the spawn prompt (rework hotspots, recurring-findings
and waived-minor ledger aggregates, all capped) plus prior-cycle LEARN
context. The PO calls the new PO-only propose_bug_hunt verb once: ≤5 items,
evidence required per item, targets validated against pest_control
participation. CEO decides per item — approve materializes a BACKLOG task
(source pest_control, never auto-starts), reject records the reason; both
feed the LEARN ledger by exploration task id; all-terminal completes the
cycle. Telegram queue pushes carry working Approve/Reject handlers
mirroring the roadmap kind. Doctrine: board.md Pest Control section +
product-owner verb entry + regenerated verb tables.

* feat(panel): Pest Control review queue

Command Center gains the pest review queue (per-item approve/reject with
reason, mirroring the roadmap queue); the Programs card and the project
settings participates-in checkboxes pick the new program up registry-driven
— the settings section renders for the first time now that a project-scoped
program exists.

* feat(board): Periscope — HoM market-research brief program

Weekly org-scoped cycle: a solo HoM spawn researches the market (web
research with mandatory source URLs — uncited findings are rejected) and
files one structured brief via the new HoM-only propose_market_brief verb:
headline, cited findings, threats/opportunities, positioning note, all
soup-checked and screened through the injection guard at persist time
(web-derived text later reaches prompts; flags recorded, content never
dropped). A brief is a report, not a proposal: the verb completes the
exploration in the same call (the x_feature asymmetry), the cycle ledger
auto-closes, and the CEO gets a best-effort notification with no
approve/reject surface (periscope deliberately never joins Telegram's
action kinds). The latest brief is injected into the roadmap exploration
prompt — Periscope feeds Printer, the first cross-role program input.

* feat(panel): Market Briefs tab (read-only)

Business page gains a Market Briefs tab listing Periscope briefs —
headline, cited findings, threats/opportunities — read-only by design; a
report has nothing to approve.

* feat(board): Coroner — event-triggered Auditor postmortems

The first EVENT program: no cron — three best-effort hooks open an autopsy
when a task bounces to its 3rd revision (the audit chokepoint), is
cancelled after work started, or is budget-blocked; all gated on arming +
one-open-autopsy dedup, none can fail the underlying transition. A solo
Auditor spawn reads the incident (server-assembled findings + transition
context) and files one propose_postmortem: incident summary, root cause,
failed stage (validated against the real status vocabulary), and ONE
process change — a playbook-kind change drafts via PlaybookService
directly into the normal pending-curation queue; the briefed draft_playbook
manifest grant was deliberately NOT added, preserving the existing
'auditor curates but never drafts' invariant test. Complete-at-propose
(report asymmetry), cycle ledger auto-closes, CEO notified link-only.
Integrated as a union with Periscope across the shared program surfaces.

* feat(panel): Coroner postmortems card

Read-only postmortems list under Business → Programs — incident, root
cause, failed stage, process change; nothing to approve, the process-change
artifact (a draft playbook) rides the existing curation queue.

* feat(board): Sentinel — Auditor drift-watch quality reports

Weekly org-scoped cycle: a solo Auditor spawn receives a server-assembled
drift context (waived-findings trend, open findings by severity,
conventions-violation hotspots, top spend — all capped, pure ORM) and files
one propose_quality_report: headline, 1-7 area-validated items with
evidence and suggested actions, overall assessment. Report semantics —
complete-at-propose, cycle auto-closes, CEO notified display-only (never on
Telegram's approve/reject surface); items are structured so a later
convert-to-task control is cheap. Integration adopts Sentinel's module-
level dict-dispatch for board-program routing (xenon-driven), folding all
prior programs in; app router mounting extracted to a helper for the same
budget.

* feat(panel): Quality Reports tab (read-only)

Business page gains the Sentinel quality-reports tab — headline, per-area
observations with evidence and suggested actions; read-only, a report has
nothing to approve.

* feat(board): Spackle — gap-fill audit program

Biweekly project-scoped PO cycle over the half-shipped surface area: API
routes without panel surfaces (and vice versa), armed flags without docs,
docs promises the code doesn't keep, dead-end tabs — the inventory diffing
is the PO's own read-tool work, ordered by the spawn prompt with file:line
citations required; the server injects only prior-cycle LEARN and the
rotation target. Rotation is now a shared module-level helper
(pick_rotation_target, parameterized by source) both project-scoped
engines use — pest_control delegates to it, behavior-identical, with a
cross-pollution test proving the two programs' rotations stay independent.
propose_gap_fill mirrors the bug-hunt verb (≤5 items, two-sided evidence
required, participation gate); per-item CEO decide materializes BACKLOG
source=spackle tasks; full Telegram kind incl. approve/reject handlers.
All seven program routers now mount from one helper.

* feat(panel): Spackle gap-fill review queue

Command Center gains the gap-fill queue mirroring the pest-control one —
per-item approve/reject with the two-sided gap evidence rendered.

* feat(board): Scales — monthly portfolio rebalance

Org-scoped PO cycle over the stale backlog: the spawn receives a capped
stale-task snapshot (BACKLOG/PENDING unclaimed >30 days) plus the charter
and prior-cycle LEARN, and files one propose_rebalance — 1-7 items, each a
resolvable task_ref with action reprioritize (validated new priority) or
cancel, rationale required. Per-item CEO decide: approve EXECUTES the
action (audited priority update, or the normal cancel path) — the first
program whose materializer mutates existing tasks instead of creating
them; reject records the reason; LEARN by exploration task id;
all-terminal completes the cycle. Full Telegram decide-kind wiring.
Integrated as the eight-program union (registry, dict dispatch, routers
helper, teardown enumerations).

* feat(panel): Scales rebalance review queue

Command Center gains the rebalance queue — per-item approve/reject with
the action, target task, and rationale rendered.

* feat(board): Mirror — quarterly positioning audit

Project-scoped HoM cycle over messaging surfaces: README claims vs shipped
reality, docs-site promises vs code, charter alignment — the audit is the
HoM's own read-tool work with citations required; the server injects the
charter, prior-cycle LEARN, and the shared rotation target. propose_
messaging_fixes mirrors the gap-fill verb (≤5 items, drift evidence naming
claim + contradicting reality, participation gate); per-item CEO decide
materializes BACKLOG source=mirror documentation tasks; full Telegram
decide-kind wiring. Nine-program union across the shared surfaces.

* feat(panel): Mirror messaging-fixes review queue

* feat(board): Megaphone — HoM standing editorial calendar

Cron cycle (3 days, org-scoped, gated on X credentials — drafting content
nobody can post is pointless): the HoM receives a shipped-this-week digest
plus Unreleased changelog bullets and files one propose_editorial_post
(angle-validated, ≤280, brand voice) that materializes a held x_editorial
draft through the SAME X-queue origination chokepoint release posts use —
zero new approval surface, notifications and CEO decide for free.
Complete-at-propose; cycle auto-closes. Ten-program union.

* feat(panel): x_editorial source labels in the X queue surfaces

* feat(board): Librarian — proactive playbook mining

Biweekly org-scoped Auditor cycle: mines recurring non-private learning
journals (≥2-count grouping with a recency fallback) against the existing
playbook-title inventory and files one propose_playbook_drafts — 1-3
drafts, each with the repeated-pattern evidence that justifies it,
duplicate titles rejected in-batch and against the live store. Drafts are
created via PlaybookService directly (the Coroner precedent — the
'auditor curates but never drafts' do-verb invariant stays intact and
tested) and land in the normal pending-curation queue the Auditor's own
triage already surfaces; no new panel surface. Complete-at-propose;
display-only CEO notification. Eleven-program union.

* feat(board): War Room — release campaign planning

EVENT program with a REAL originator (unlike coroner's stub): a release
publish hooks a campaign brief beside the release-post seam, and the CEO's
run-now originates on demand — the cron loop never fires it. The HoM
designs a 2-6 post arc (teaser → launch → follow-up → spotlight; 280-cap,
future strictly-ascending publish_after, stage vocabulary) and one
propose_campaign call materializes each post as a held x_campaign draft
through the X-queue chokepoint. V1 is manual-cadence by design: publish_
after renders as queue guidance and the CEO approves each post at its
moment — nothing auto-posts, ever; the auto-schedule upgrade is a
documented ceiling. Twelve-program union: full registry complete.

* feat(panel): x_campaign labels + publish-after guidance in the X queue

* feat(board): Barfly — adjacent-conversation replies

Cron cycle (2 days, org-scoped, X-credentials gated): the engine searches
X for conversations where RoboCo is relevant but unmentioned (new OAuth-
signed search_recent on the client; queries + candidate cap configurable),
screens every fetched tweet through the injection guard (stored unclamped
— a clamp was truncating the candidate under the envelope, caught by the
dev's own tests), dedupes via the existing x_seen_mentions ledger (no
migration; also prevents double-drafting against the mentions poll), and
opens one held HoM exploration carrying the screened candidates. propose_
conversation_replies enforces candidate-id-only replies (≤5, 280-cap);
each materializes a held x_barfly draft through the X-queue chokepoint,
threaded via a new in_reply_to seam on post_tweet that only x_barfly
drafts use. The X redraft machinery is now dict-dispatch over per-source
extractors with reply-ref carry for x_barfly. Thirteen-program registry.
War Room's test fakes gained the new abstract search_recent stub.

* feat(board): Dogfood — the PO walks the product

The fourteenth and final registry entry, completing the catalog. EVENT
program (release-publish hook beside the war-room hook + CEO run-now, both
through the same real originator; the cron loop never fires it), project-
scoped with shared rotation. The permission surface is the careful part:
the PO's dogfood spawn — and ONLY that spawn — gets the Playwright MCP
mounted, via a task-scoped fail-closed probe mirroring the video-authoring
precedent (a PO spawned for roadmap/pest/scales never sees browser tools;
tested both ways); the PM agent image bakes chromium unconditionally like
the ux image, the mount stays task-gated in code. The walk targets the
rotation target's live surfaces (panel_base_url only when the target is
the org's own project, honest degradation otherwise); propose_friction_
fixes files ≤5 walked-path-evidenced items; per-item CEO decide
materializes BACKLOG source=dogfood tasks; full Telegram decide kind.
Also: megaphone/librarian/war_room arming keys restored to the settings
validator — their panel toggles would have been rejected (dropped in
earlier unions; the same silent-arming class the drill killed once
already).

* feat(panel): Dogfood friction review queue

* chore(board): final whole-branch sweep fixes

The night's closing adversarial pass over the integrated fourteen-program
registry found ONE functional defect — the war-room test fakes' post_tweet
predated Barfly's in_reply_to_tweet_id kwarg (LSP violation, the only red
in an otherwise fully green gate) — plus doc/test drift, all fixed: the
source-parity test completes to fourteen (spackle/mirror were silently
absent while its neighboring comment claimed full coverage), the PO
identity doc gains its missing Dogfood verb, the auditor quick-list gains
propose_postmortem, three stale comments corrected (rotation docstring,
panel registry header, X source enumerations), the dogfood release-hook
gains the exception-swallow test its four sibling hooks already had, and
the CHANGELOG's Unreleased section documents the whole Board Programs
train. Full make quality: exit 0, all gates green.

* docs: full documentation sweep for the Board Programs train

CLAUDE.md's roadmap-engine entry superseded by the Board Program registry
entry (all fourteen programs, arming, scoping, LEARN, guardrails) with the
role verb tables and playwright row refreshed; docs/rag gains the agent-
facing architecture doc plus full propose_* call-shape sections in the
three board role docs, and corrects the strategy-engine section to shipped
reality (only idle→roadmap is wired); docs/map covers the registry + all
twelve engines with flags, gotchas, and drift notes. The 0.27.0 reference
inventory confirmed only the release-executor's canonical set carries the
version — left for the 0.28.0 cut.

* feat(board): human titles + descriptions on every program surface

Raw registry keys rendered as bare panel labels — an operator reading
x_feature had no idea what enabling or running it does. The registry
dataclass gains title/description (test-enforced non-empty for every
entry, unique titles), the API passes them through, and every surface
renders title-with-description-tooltip instead of the key: the Programs
card (label, toggle hint, run-now toast), and the project settings
participates-in/excluded-from checkboxes.

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-07-25 17:13:32 +02:00

582 lines
20 KiB
Python

"""ScalesService coverage: per-item approve EXECUTES the item's action
against a live target task (reprioritize its priority, or cancel it) —
idempotent; reject records a reason (idempotent, never touches the target);
the exploration task completes once every item is terminal.
Mirrors test_pest_control_service.py — the structural difference is that
approval mutates an existing task in place instead of materializing a new
one, so there is no ``materialized_task_id`` and no project-participation
gate (a rebalance item targets any live task, not a new draft against an
opted-in project).
"""
from __future__ import annotations
from datetime import UTC, datetime
from typing import TYPE_CHECKING, cast
from uuid import uuid4
import pytest
import pytest_asyncio
from roboco.db.tables import (
AgentTable,
AuditLogTable,
BoardProgramCycleTable,
ProjectTable,
SystemSettingTable,
TaskTable,
)
from roboco.foundation import identity as _foundation
from roboco.foundation.policy.content import markers
from roboco.models.base import (
AgentRole,
AgentStatus,
Complexity,
Team,
)
from roboco.models.base import TaskNature as TN
from roboco.models.base import TaskStatus as TS
from roboco.models.base import TaskType as TT
from roboco.services import board_programs as bp_module
from roboco.services.scales_service import ScalesService, get_scales_service
from roboco.services.task import (
CORONER_SOURCE,
PEST_CONTROL_SOURCE,
ROADMAP_SOURCE,
SCALES_SOURCE,
X_FEATURE_EXPLORATION_SOURCE,
)
from sqlalchemy import delete, select, update
if TYPE_CHECKING:
from uuid import UUID
from sqlalchemy.ext.asyncio import AsyncSession
SYSTEM_UUID = _foundation.AGENTS["system"].uuid
PO_UUID = _foundation.AGENTS["product-owner"].uuid
CEO_UUID = _foundation.AGENTS["ceo"].uuid
ONE = 1
TWO = 2
@pytest_asyncio.fixture(autouse=True)
async def _purge_board_program_pollution(db_session: AsyncSession) -> None:
"""See test_board_program_engine.py's identical fixture."""
await db_session.execute(
delete(SystemSettingTable).where(SystemSettingTable.key.like("board_program.%"))
)
await db_session.execute(delete(BoardProgramCycleTable))
await db_session.execute(
update(TaskTable)
.where(
TaskTable.source.in_(
[
ROADMAP_SOURCE,
X_FEATURE_EXPLORATION_SOURCE,
PEST_CONTROL_SOURCE,
CORONER_SOURCE,
SCALES_SOURCE,
]
),
TaskTable.status.notin_([TS.COMPLETED, TS.CANCELLED]),
)
.values(status=TS.CANCELLED)
)
await db_session.commit()
async def _seed_agents(session: AsyncSession) -> None:
for uuid, slug, role, team in (
(SYSTEM_UUID, "system", AgentRole.SYSTEM, None),
(PO_UUID, "product-owner", AgentRole.PRODUCT_OWNER, Team.BOARD),
(CEO_UUID, "ceo", AgentRole.CEO, None),
):
if await session.get(AgentTable, uuid) is None:
session.add(
AgentTable(
id=uuid,
name=slug,
slug=slug,
role=role,
team=team,
status=AgentStatus.ACTIVE,
model_config={},
system_prompt="x",
capabilities=[],
permissions={},
metrics={},
)
)
await session.flush()
async def _seed_project(session: AsyncSession, slug: str) -> ProjectTable:
await _seed_agents(session)
project = ProjectTable(
id=uuid4(),
name=slug,
slug=slug,
git_url=f"https://example.com/{slug}.git",
assigned_cell=Team.BACKEND,
created_by=SYSTEM_UUID,
)
session.add(project)
await session.flush()
return project
async def _seed_target_task(
session: AsyncSession,
project: ProjectTable,
*,
title: str,
status: TS = TS.BACKLOG,
priority: int = 2,
) -> TaskTable:
task = TaskTable(
id=uuid4(),
title=title,
description="A live backlog task the rebalance plan targets",
acceptance_criteria=["done"],
status=status,
priority=priority,
task_type=TT.CODE,
nature=TN.TECHNICAL,
estimated_complexity=Complexity.LOW,
created_by=SYSTEM_UUID,
project_id=project.id,
team=Team.BACKEND,
)
session.add(task)
await session.flush()
return task
def _item(
idx: int,
target: TaskTable,
*,
action: str = "reprioritize",
new_priority: int | None = 0,
status: str = "proposed",
) -> dict:
return {
"id": f"item-{idx}",
"task_ref": str(target.id)[:8],
"target_task_id": str(target.id),
"target_task_title": target.title,
"action": action,
"new_priority": new_priority if action == "reprioritize" else None,
"rationale": f"Rationale for item {idx} — a substantive reason",
"status": status,
"reject_reason": None,
"executed_detail": None,
}
async def _seed_cycle(session: AsyncSession, *, items: list[dict]) -> TaskTable:
await _seed_agents(session)
task = TaskTable(
id=uuid4(),
title="Scales portfolio-rebalance cycle",
description="Review the live backlog and propose a rebalance.",
acceptance_criteria=["propose_rebalance() called once"],
status=TS.PENDING,
priority=2,
task_type=TT.ADMINISTRATIVE,
nature=TN.NON_TECHNICAL,
estimated_complexity=Complexity.LOW,
created_by=SYSTEM_UUID,
assigned_to=PO_UUID,
team=Team.BOARD,
source=SCALES_SOURCE,
confirmed_by_human=False,
)
session.add(task)
await session.flush()
markers.set_rebalance_plan(task, {"items": items})
await session.flush()
return task
def _svc(session: AsyncSession) -> ScalesService:
return get_scales_service(session)
def _id(task: TaskTable) -> UUID:
return cast("UUID", task.id)
@pytest.mark.asyncio
async def test_approve_reprioritize_changes_target_priority(
db_session: AsyncSession,
) -> None:
project = await _seed_project(db_session, "backend-svc")
target = await _seed_target_task(
db_session, project, title="Stale task", priority=2
)
cycle = await _seed_cycle(
db_session, items=[_item(0, target, action="reprioritize", new_priority=0)]
)
result = await _svc(db_session).approve_item(
_id(cycle), "item-0", created_by=CEO_UUID
)
assert result is not None
assert result.status == "approved"
assert result.executed_detail == "priority changed to P0"
await db_session.refresh(target)
assert target.priority == 0
assert target.status == TS.BACKLOG # untouched otherwise
await db_session.refresh(cycle)
payload = markers.get_rebalance_plan(cycle)
assert payload is not None
item0 = next(i for i in payload["items"] if i["id"] == "item-0")
assert item0["status"] == "approved"
assert item0["executed_detail"] == "priority changed to P0"
@pytest.mark.asyncio
async def test_approve_cancel_cancels_target_task(db_session: AsyncSession) -> None:
project = await _seed_project(db_session, "backend-svc")
target = await _seed_target_task(db_session, project, title="Dead weight task")
cycle = await _seed_cycle(db_session, items=[_item(0, target, action="cancel")])
result = await _svc(db_session).approve_item(
_id(cycle), "item-0", created_by=CEO_UUID
)
assert result is not None
assert result.status == "approved"
assert result.executed_detail == "cancelled"
await db_session.refresh(target)
assert target.status == TS.CANCELLED
@pytest.mark.asyncio
async def test_approve_is_idempotent(db_session: AsyncSession) -> None:
project = await _seed_project(db_session, "backend-svc")
target = await _seed_target_task(db_session, project, title="Stale task")
cycle = await _seed_cycle(
db_session, items=[_item(0, target, action="reprioritize", new_priority=0)]
)
svc = _svc(db_session)
first = await svc.approve_item(_id(cycle), "item-0", created_by=CEO_UUID)
second = await svc.approve_item(_id(cycle), "item-0", created_by=CEO_UUID)
assert first is not None
assert second is not None
assert second.status == "already_approved"
assert second.executed_detail == first.executed_detail
await db_session.refresh(target)
assert target.priority == 0 # only ever applied once
@pytest.mark.asyncio
async def test_reject_records_reason_and_leaves_target_untouched(
db_session: AsyncSession,
) -> None:
project = await _seed_project(db_session, "backend-svc")
target = await _seed_target_task(
db_session, project, title="Stale task", priority=2
)
cycle = await _seed_cycle(
db_session, items=[_item(0, target, action="reprioritize", new_priority=0)]
)
result = await _svc(db_session).reject_item(
_id(cycle), "item-0", "still worth doing at current priority"
)
assert result is not None
assert result.status == "rejected"
await db_session.refresh(target)
assert target.priority == TWO # untouched
await db_session.refresh(cycle)
payload = markers.get_rebalance_plan(cycle)
assert payload is not None
item0 = next(i for i in payload["items"] if i["id"] == "item-0")
assert item0["status"] == "rejected"
assert item0["reject_reason"] == "still worth doing at current priority"
@pytest.mark.asyncio
async def test_reject_is_idempotent(db_session: AsyncSession) -> None:
project = await _seed_project(db_session, "backend-svc")
target = await _seed_target_task(db_session, project, title="Stale task")
cycle = await _seed_cycle(db_session, items=[_item(0, target, action="cancel")])
svc = _svc(db_session)
await svc.reject_item(_id(cycle), "item-0", "reason one")
second = await svc.reject_item(_id(cycle), "item-0", "reason two")
assert second is not None
assert second.status == "already_rejected"
@pytest.mark.asyncio
async def test_cannot_reject_an_approved_item(db_session: AsyncSession) -> None:
project = await _seed_project(db_session, "backend-svc")
target = await _seed_target_task(db_session, project, title="Stale task")
cycle = await _seed_cycle(db_session, items=[_item(0, target, action="cancel")])
svc = _svc(db_session)
await svc.approve_item(_id(cycle), "item-0", created_by=CEO_UUID)
result = await svc.reject_item(_id(cycle), "item-0", "changed my mind")
assert result is not None
assert result.status == "invalid_state"
@pytest.mark.asyncio
async def test_cannot_approve_a_rejected_item(db_session: AsyncSession) -> None:
project = await _seed_project(db_session, "backend-svc")
target = await _seed_target_task(db_session, project, title="Stale task")
cycle = await _seed_cycle(db_session, items=[_item(0, target, action="cancel")])
svc = _svc(db_session)
await svc.reject_item(_id(cycle), "item-0", "not now")
result = await svc.approve_item(_id(cycle), "item-0", created_by=CEO_UUID)
assert result is not None
assert result.status == "invalid_state"
@pytest.mark.asyncio
async def test_all_items_terminal_completes_exploration_task(
db_session: AsyncSession,
) -> None:
project = await _seed_project(db_session, "backend-svc")
target0 = await _seed_target_task(db_session, project, title="Task A")
target1 = await _seed_target_task(db_session, project, title="Task B")
cycle = await _seed_cycle(
db_session,
items=[
_item(0, target0, action="cancel"),
_item(1, target1, action="reprioritize", new_priority=1),
],
)
svc = _svc(db_session)
await svc.approve_item(_id(cycle), "item-0", created_by=CEO_UUID)
assert cycle.status == TS.PENDING # one item still proposed
await svc.reject_item(_id(cycle), "item-1", "not now")
assert cycle.status == TS.COMPLETED # both items terminal
@pytest.mark.asyncio
async def test_approve_target_already_claimed_is_invalid_state(
db_session: AsyncSession,
) -> None:
"""The target left BACKLOG/PENDING since propose time (e.g. a PM claimed
it) — approve refuses instead of silently mutating in-flight work."""
project = await _seed_project(db_session, "backend-svc")
target = await _seed_target_task(
db_session, project, title="Now claimed", status=TS.IN_PROGRESS
)
cycle = await _seed_cycle(
db_session, items=[_item(0, target, action="reprioritize", new_priority=0)]
)
result = await _svc(db_session).approve_item(
_id(cycle), "item-0", created_by=CEO_UUID
)
assert result is not None
assert result.status == "invalid_state"
@pytest.mark.asyncio
async def test_approve_missing_target_is_invalid_state(
db_session: AsyncSession,
) -> None:
await _seed_agents(db_session)
ghost = {
"id": "item-0",
"task_ref": "deadbeef",
"target_task_id": str(uuid4()),
"target_task_title": "Gone",
"action": "cancel",
"new_priority": None,
"rationale": "Doesn't matter, target is gone",
"status": "proposed",
"reject_reason": None,
"executed_detail": None,
}
cycle = await _seed_cycle(db_session, items=[ghost])
result = await _svc(db_session).approve_item(
_id(cycle), "item-0", created_by=CEO_UUID
)
assert result is not None
assert result.status == "invalid_state"
@pytest.mark.asyncio
async def test_unknown_task_returns_none(db_session: AsyncSession) -> None:
result = await _svc(db_session).approve_item(uuid4(), "item-0", created_by=CEO_UUID)
assert result is None
@pytest.mark.asyncio
async def test_unknown_item_id_returns_none(db_session: AsyncSession) -> None:
project = await _seed_project(db_session, "backend-svc")
target = await _seed_target_task(db_session, project, title="Stale task")
cycle = await _seed_cycle(db_session, items=[_item(0, target, action="cancel")])
result = await _svc(db_session).approve_item(
_id(cycle), "item-999", created_by=CEO_UUID
)
assert result is None
@pytest.mark.asyncio
async def test_list_open_cycles_excludes_completed(db_session: AsyncSession) -> None:
project = await _seed_project(db_session, "backend-svc")
target = await _seed_target_task(db_session, project, title="Stale task")
cycle = await _seed_cycle(db_session, items=[_item(0, target, action="cancel")])
svc = _svc(db_session)
open_before = await svc.list_open_cycles()
assert cycle.id in {t.id for t in open_before}
await svc.approve_item(_id(cycle), "item-0", created_by=CEO_UUID)
open_after = await svc.list_open_cycles()
assert cycle.id not in {t.id for t in open_after}
@pytest.mark.asyncio
async def test_maybe_complete_cycle_emits_audit(db_session: AsyncSession) -> None:
project = await _seed_project(db_session, "backend-svc")
target0 = await _seed_target_task(db_session, project, title="Task A")
target1 = await _seed_target_task(db_session, project, title="Task B")
cycle = await _seed_cycle(
db_session,
items=[
_item(0, target0, action="cancel"),
_item(1, target1, action="reprioritize", new_priority=1),
],
)
svc = _svc(db_session)
await svc.approve_item(_id(cycle), "item-0", created_by=CEO_UUID)
await svc.reject_item(_id(cycle), "item-1", "not now")
assert cycle.status == TS.COMPLETED
rows = (
(
await db_session.execute(
select(AuditLogTable).where(AuditLogTable.target_id == cycle.id)
)
)
.scalars()
.all()
)
audit = [
r
for r in rows
if r.event_type == "task.completed"
or str(r.details.get("to_status", "")).lower() == "completed"
]
assert audit, (
"expected a task.completed audit row for the PENDING -> COMPLETED transition"
)
@pytest.mark.asyncio
async def test_approve_emits_rebalance_execution_audit(
db_session: AsyncSession,
) -> None:
project = await _seed_project(db_session, "backend-svc")
target = await _seed_target_task(db_session, project, title="Stale task")
cycle = await _seed_cycle(
db_session, items=[_item(0, target, action="reprioritize", new_priority=0)]
)
await _svc(db_session).approve_item(_id(cycle), "item-0", created_by=CEO_UUID)
rows = (
(
await db_session.execute(
select(AuditLogTable).where(AuditLogTable.target_id == target.id)
)
)
.scalars()
.all()
)
assert any(r.event_type == "task.scales_rebalance" for r in rows)
# --------------------------------------------------------------------------- #
# LEARN wiring: approve/reject best-effort record onto the open
# board_program_cycles row for "scales".
# --------------------------------------------------------------------------- #
async def _seed_cycle_ledger_row(session: AsyncSession, task: TaskTable) -> None:
session.add(
BoardProgramCycleTable(
program_key="scales",
exploration_task_id=task.id,
opened_at=datetime.now(UTC),
)
)
await session.flush()
@pytest.mark.asyncio
async def test_approve_records_learn_decision(db_session: AsyncSession) -> None:
project = await _seed_project(db_session, "backend-svc")
target = await _seed_target_task(db_session, project, title="Stale task")
cycle = await _seed_cycle(
db_session, items=[_item(0, target, action="reprioritize", new_priority=0)]
)
await _seed_cycle_ledger_row(db_session, cycle)
await _svc(db_session).approve_item(_id(cycle), "item-0", created_by=CEO_UUID)
row = (
await db_session.execute(
select(BoardProgramCycleTable).where(
BoardProgramCycleTable.program_key == "scales"
)
)
).scalar_one()
assert row.items_approved == ONE
assert {
"item_ref": "item-0",
"verdict": "approved",
"reason": None,
} in row.decisions
@pytest.mark.asyncio
async def test_reject_records_learn_decision_with_reason(
db_session: AsyncSession,
) -> None:
project = await _seed_project(db_session, "backend-svc")
target = await _seed_target_task(db_session, project, title="Stale task")
cycle = await _seed_cycle(db_session, items=[_item(0, target, action="cancel")])
await _seed_cycle_ledger_row(db_session, cycle)
await _svc(db_session).reject_item(_id(cycle), "item-0", "not a priority")
row = (
await db_session.execute(
select(BoardProgramCycleTable).where(
BoardProgramCycleTable.program_key == "scales"
)
)
).scalar_one()
assert row.items_rejected == ONE
assert {
"item_ref": "item-0",
"verdict": "rejected",
"reason": "not a priority",
} in row.decisions
@pytest.mark.asyncio
async def test_approve_survives_learn_recording_failure(
db_session: AsyncSession, monkeypatch: pytest.MonkeyPatch
) -> None:
"""A record_decision blow-up must never break the CEO's approve."""
project = await _seed_project(db_session, "backend-svc")
target = await _seed_target_task(db_session, project, title="Stale task")
cycle = await _seed_cycle(db_session, items=[_item(0, target, action="cancel")])
await _seed_cycle_ledger_row(db_session, cycle)
async def _boom(_self: object, *_args: object, **_kwargs: object) -> None:
raise RuntimeError("learn boom")
monkeypatch.setattr(bp_module.BoardProgramEngine, "record_decision", _boom)
result = await _svc(db_session).approve_item(
_id(cycle), "item-0", created_by=CEO_UUID
)
assert result is not None
assert result.status == "approved"