Files
roboco/roboco/services/gateway/evidence_builder.py
T
cfde4369b1 Token optimization levers — claim-scoped briefing, payload caps, role-scoped optimal, notification-spawn cooldown (#292)
* feat(gateway): claim-scoped context briefing — heavy sections only on context-acquisition verbs

* feat(gateway): cap unbounded LLM-facing payloads — embedded diffs, notification bodies, handoff journal content, north star

* feat(mcp): role-scope the optimal server's tool groups; index management becomes dev/test-only

* feat(mcp): cap per-result content on kb/error/learning search, mentor sources, rag citations

* refactor(gateway): extract heavy-briefing sections + clip helper to keep xenon ranks

* feat(orchestrator): cross-tick cooldown for notification-triggered spawns

* feat(usage,orchestrator): scope spawn-waste to anthropic sessions; cap agent Bash output via settings env

* docs: claim-scoped briefing, payload caps, optimal role-scoping, notification-spawn cooldown

* test(mcp): type the mixed-item cap fixture explicitly

* fix(orchestrator): lazy-init the notification-spawn cooldown store

* fix(lifecycle): admin-override claim reconciliation + PM request_changes verb (S6 postmortem B3+B4)

B3 — admin_set_status now reconciles claim ownership when leaving BLOCKED:
review/queue targets clear claimed_by/claimed_at/active_claimant_id and
consume the pre-block snapshot (a stale escalation claim was stranding the
next claimant: give_me_work handed the task out while note() bounced
not_authorized — the live b8fe0494 wedge). The pending/in_progress restore
path also syncs active_claimant_id, and a REST PATCH unassign releases the
claim with it.

B4 — new PM verb request_changes: awaiting_pm_review -> needs_revision with
concrete issues. The PM previously had no reject at merge review (only
complete/escalate), so an AC/scope violation looped i_am_blocked->escalate
4x live. Full vertical: lifecycle transition + ActionSpec + IntentSpec,
TaskService.request_changes (routes like a QA fail — original dev for a
leaf, revision PM for assembled; issues appended to dev_notes), verb-runner
compose, choreographer verb (spec gate + non-empty issues + soup check +
a2a delivery of the reject reason), HTTP routes on both PM flows, MCP tool,
journal:decision tracing, PM prompts, regenerated lifecycle artifacts.

* fix(panel): stop scorecard fetches for fallback-roster placeholder ids

useAgents() serves the static AGENT_ROSTER (ids "1".."22") while agent
definitions load; the Scorecards tab fetched a member scorecard per row
immediately, firing 22 guaranteed-422 requests per refetch cycle. Through
the browser's per-origin connection limit those queued every metrics-page
query behind them (~10s of skeletons on every tab). Gate the fetch on a
real member id (agent UUID or the "ceo" alias).

* Upgraded uv.lock

* fix(sequencing): declared deps become real edges + full loop-breaker coverage + assembled-branch freshness (S6 postmortem B1/B2/B6 + breaker)

B1a — code delegations REQUIRE a collision surface: new TASK_AT_DELEGATE
completeness spec (conditional FieldRequirement, when=('task_type','code'))
enforced at the gateway delegate gate. A no-surface code sibling is
'parallel to everything' by analyzer design, which is how two devs ran the
CEO's explicitly-ordered work out of order (f3e1afc5: seq#1 started before
seq#0, zero dependency edges). PM prompts updated; REST/manual creation
(TASK_AT_CREATE) unchanged.

B1b — the CEO's declared 'Depends on' lists become real edges: DraftSurface
gains declared_depends_on; SequencingService.analyze unions declared edges
(validated: self/out-of-range rejected) with the derived collision rules,
cycle-checked by the existing toposort. confirm_live_batch/preview_batch
map each draft's depends_on through (string indices coerced); intake tool
doc + prompter role prompt instruct verbatim copying. The live S6 root got
1 of its 3 declared in-batch edges and started alongside still-running R3.

Breaker coverage — the progress-aware respawn circuit breaker
(_pm_respawn_should_gate: strike counting, status-advance reset,
tracing-gap budget, DB durability, one-shot CEO notification) was consulted
by only 3 spawn paths; the doc/QA/dev/PR-review/PR-gate/revision/board
paths spawned unguarded at fixed cadence (the 26-respawn fe-doc loop,
~$7.20). Now consulted at every task-keyed spawn site (14 total).

B2 — assembled-branch freshness: submit_up/submit_root auto-sync the
assembled branch when it has fallen behind its base (children are terminal
at submit time, so the rebase is safe; master is never written). A rebase
conflict is a hard reject naming the files instead of a blind re-review —
kills the needs_revision↔awaiting_pr_review ping-pong of re-submitting a
stale head. Leaf i_am_done already had the behind-base gate; claim-time
fetch-fresh cut already existed.

B6 — documenter revision-pass loop: the awaiting_documentation bail
rejections (i_am_blocked/unclaim) now name the actual exit (i_documented
re-affirm) and the documenter prompt gets an explicit revision-pass rule.

* fix(orchestration): assembly-integrity gate + dispatcher heartbeat (incidents #11, #1)

Assembly integrity — submit_up/submit_root refuse when a completed child's
commits are not patch-present in the assembled branch (git cherry —
rebase-safe; branch pruned after merge or any git error fails open). Live
incident #11: a completed revert subtask's merge was lost from the cell
branch and the review gate re-flagged the exact violation the revert fixed,
spawning another revision cycle.

Dispatcher heartbeat — a dispatcher.alive audit row every 5 minutes from
the dispatch loop. The 2026-07-01 outage was 4h25m of fleet-wide silence
with no way to distinguish 'loop dead' from 'no work'; the loop's stdout
died with the container while audit_log survives. CHANGELOG for tonight's
full sweep included.

* style: ruff format for the orchestration sweep

* refactor(gateway): fold the assembled-submit guards + trim complexity under the xenon gate

_assembled_submit_guards combines the #11 integrity check and B2 freshen for
submit_up/submit_root; lifecycle's invalid-source remediate and git's
per-child cherry probe extracted into helpers. Test harnesses built via
__new__ stub the respawn tracker (the breaker now runs on their paths).

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-07-02 05:46:01 +02:00

245 lines
9.5 KiB
Python

"""Build the `evidence` and `context_briefing` blocks for verb responses.
`evidence` is task-scoped: PR + commits + files + journal highlights.
`context_briefing` is agent-scoped: unread A2As, mentions, notifications,
task metadata gaps, recent team activity, blockers in lane.
This module is pure: it takes already-fetched lists and assembles them.
The choreographer queries the data via existing services and passes it in.
"""
from __future__ import annotations
from dataclasses import asdict, dataclass, field
from typing import Any
BRIEFING_LIST_CAP = 10
@dataclass(frozen=True)
class EvidencePayload:
pr_number: int | None
pr_url: str | None
pr_diff_summary: str | None
commits: list[dict[str, Any]]
files_changed: list[str]
dev_summary: str | None
journal_highlights: list[dict[str, Any]]
acceptance_criteria_status: list[dict[str, Any]]
# Architectural-conventions validator findings on the changed files, so QA
# can flag a misplaced definition / suppression. Empty when the subsystem
# is off; a single ``could_not_run`` entry surfaces a fail-loud explicitly.
convention_findings: list[dict[str, Any]] = field(default_factory=list)
def as_dict(self) -> dict[str, Any]:
return asdict(self)
@dataclass(frozen=True)
class BriefingInputs:
"""All agent-scoped context collected before assembling `context_briefing`."""
unread_a2a: list[dict[str, Any]]
unread_mentions: list[dict[str, Any]]
pending_notifications: list[dict[str, Any]]
task_metadata_gaps: list[str]
recent_team_activity: list[dict[str, Any]]
blockers_in_my_lane: list[dict[str, Any]]
# Prior-work digest for the briefed task (None when there is no task in
# scope or no prior work to resume from). Pushed so a freshly spawned or
# respawned agent picks up where the previous worker left off instead of
# re-exploring the codebase from cold.
task_handoff: dict[str, Any] | None = None
# Compact company charter (north star + objectives + operating policy), or
# None when unset — injected so every agent's work is goal-aware.
company_goals: dict[str, Any] | None = None
# Char cap for a unified diff embedded in an LLM-facing payload (~5K tokens).
# The head is kept (file headers + earliest hunks carry the most signal); the
# marker points at the full diff so a reviewer is never silently blinded.
EVIDENCE_DIFF_CAP_CHARS = 20_000
def truncate_diff(diff: str | None, limit: int = EVIDENCE_DIFF_CAP_CHARS) -> str | None:
"""Cap a diff for envelope embedding; annotate what was omitted."""
if not diff or len(diff) <= limit:
return diff
omitted = len(diff) - limit
return (
diff[:limit]
+ f"\n… [diff truncated: {omitted} chars omitted — read the full diff on"
" the PR, or scope roboco_git_diff to a single file_path]"
)
def build_evidence_for_task(
task: Any,
*,
journal_highlights: list[dict[str, Any]],
files_changed: list[str],
pr_diff_summary: str | None = None,
convention_findings: list[dict[str, Any]] | None = None,
) -> EvidencePayload:
"""Compose an EvidencePayload from a Task model + supplemental data."""
return EvidencePayload(
pr_number=task.pr_number,
pr_url=task.pr_url,
pr_diff_summary=truncate_diff(pr_diff_summary),
commits=list(task.commits or []),
files_changed=list(files_changed),
dev_summary=task.dev_notes,
journal_highlights=list(journal_highlights),
acceptance_criteria_status=list(task.acceptance_criteria_status or []),
convention_findings=list(convention_findings or []),
)
def _typed(value: Any, expected: type | tuple[type, ...], default: Any) -> Any:
"""Return ``value`` only when it is the expected type, else ``default``.
Keeps every handoff field serialisable: a bare mock or unexpected attribute
type degrades to a safe default rather than leaking a non-JSON object.
"""
return value if isinstance(value, expected) else default
def _has_prior_work(
commits: list,
acceptance: list,
highlights: list,
pr_number: int | None,
dev_summary: str | None,
completed_deps: list,
pr_review: dict[str, Any] | None,
) -> bool:
"""True when any resumable prior-work signal is present on the task."""
return bool(
commits
or acceptance
or highlights
or pr_number is not None
or dev_summary
or completed_deps
or pr_review is not None
)
def build_task_handoff(
task: Any, journal_highlights: list[dict[str, Any]]
) -> dict[str, Any] | None:
"""Compose a compact prior-work digest for the briefed task.
Returns ``None`` when there is no task or no prior work worth resuming
from, so the briefing only carries a handoff when one genuinely exists.
DB-only by design — no git diff — so it is cheap enough to attach to
every task-scoped briefing.
"""
if task is None:
return None
commits = _typed(task.commits, list, [])
acceptance = _typed(task.acceptance_criteria_status, list, [])
highlights = _typed(journal_highlights, list, [])
pr_number = _typed(task.pr_number, int, None)
dev_summary = _typed(task.dev_notes, str, None)
# Upstream dependencies that completed and were cleared — present only on a
# just-unblocked task, so the revived dependent knows what it can build on.
completed_deps = _typed(getattr(task, "completed_dependency_ids", None), list, [])
# The persisted in-path PR-review gate verdict + concrete issues.
# ``pr_fail`` writes ``notes_structured.pr_review``; surfacing it here puts
# the concrete issues in every PM briefing so a respawned PM doesn't
# re-submit the same PR blind.
pr_review = _extract_pr_review(getattr(task, "notes_structured", None))
if not _has_prior_work(
commits,
acceptance,
highlights,
pr_number,
dev_summary,
completed_deps,
pr_review,
):
return None
handoff: dict[str, Any] = {
"pr_number": pr_number,
"pr_url": _typed(task.pr_url, str, None),
"branch_name": _typed(task.branch_name, str, None),
"commit_count": len(commits),
"recent_commits": commits[-BRIEFING_LIST_CAP:],
"dev_summary": dev_summary,
"acceptance_criteria_status": acceptance[:BRIEFING_LIST_CAP],
"journal_highlights": highlights[:BRIEFING_LIST_CAP],
"completed_dependency_ids": [
str(d) for d in completed_deps[:BRIEFING_LIST_CAP]
],
}
if pr_review is not None:
handoff["pr_review"] = pr_review
return handoff
def _extract_pr_review(notes_structured: Any) -> dict[str, Any] | None:
"""Pull the canonical ``pr_review`` slot out of ``notes_structured``.
Returns ``None`` when there is no structured note, no ``pr_review`` key, or
the slot isn't a dict — so the handoff omits the field entirely (no
misleading empty slot) for a task with no prior gate verdict. Only the
well-typed scalar/list fields the gate writes are forwarded; anything else
degrades to a safe default so a malformed slot never leaks a non-JSON
object into the briefing.
"""
if not isinstance(notes_structured, dict):
return None
raw = notes_structured.get("pr_review")
if not isinstance(raw, dict):
return None
verdict = _typed(raw.get("verdict"), str, None)
summary = _typed(raw.get("summary"), str, None)
issues = _typed(raw.get("issues"), list, [])
head_sha = _typed(raw.get("head_sha"), str, None)
if not (verdict or summary or issues or head_sha):
return None
surface: dict[str, Any] = {"issues": list(issues[:BRIEFING_LIST_CAP])}
if verdict:
surface["verdict"] = verdict
if summary:
surface["summary"] = summary
if head_sha:
surface["head_sha"] = head_sha
return surface
def build_context_briefing(inputs: BriefingInputs) -> dict[str, Any]:
"""Compose the context_briefing dict; caps each list at BRIEFING_LIST_CAP items.
Empty sections (empty lists / dicts) are omitted: the agent reads this payload
on every verb response, and an absent section reads as "nothing here" to the
model, identical to an empty one, without the per-call token cost.
"""
briefing: dict[str, Any] = {
"unread_a2a": inputs.unread_a2a[:BRIEFING_LIST_CAP],
"unread_mentions": inputs.unread_mentions[:BRIEFING_LIST_CAP],
"pending_notifications": inputs.pending_notifications[:BRIEFING_LIST_CAP],
"task_metadata_gaps": list(inputs.task_metadata_gaps),
"recent_team_activity": inputs.recent_team_activity[:BRIEFING_LIST_CAP],
"blockers_in_my_lane": inputs.blockers_in_my_lane[:BRIEFING_LIST_CAP],
"task_handoff": inputs.task_handoff,
"company_goals": inputs.company_goals,
}
return {key: value for key, value in briefing.items() if value}
def shape_memory_query(role: str, title: str, task_type: str) -> str:
"""Role-shape the institutional-memory query so each role retrieves what it
actually needs: implementation lessons for a dev, decomposition lessons for a
PM, defect patterns for QA, doc patterns for a documenter."""
if role == "developer":
return f"implementation lessons and playbooks for {title} ({task_type})"
if role in ("cell_pm", "main_pm"):
return f"decomposition and planning lessons for {title}"
if role == "qa":
return f"recurring defects and review feedback for {task_type}"
if role == "documenter":
return f"documentation patterns for {task_type}"
return f"lessons and playbooks for {title}"