mirror of
https://github.com/rennf93/roboco.git
synced 2026-08-03 07:23:24 +02:00
Delegation detail-fidelity + PM-loop hardening (#541)
* feat(gateway): delegation detail-fidelity — details survive hand-off, both directions
Details thinned out at every delegation hop: a PM child task mapped to no
parent criterion was legal (coverage only surfaced at submit_up, after the
whole wave ran — a 12-subtask docs tree grew through 8 review rounds that
way, one child titled 'docs page and route wrapper' shipping only the
page), and QA could pass work on a gestalt read (a 4-scene video brief
shipped 3 scenes past every gate because the features existed only in
prose). Three chokepoint gates:
- delegate (down): every child must declare covers_parent_criteria
resolving against the parent's real acceptance criteria — no mapping or
an unresolvable ref rejects naming every offending child and the valid
criteria; the success envelope carries parent_ac_coverage
{covered, uncovered} so a wave-planning PM sees remaining gaps in the
same turn. Full coverage stays enforced at submit_up (waves stay legal).
- pass_review (up): mandatory criteria_verified — one {criterion,
evidence} entry per task AC, matched by the findings ledger's
id-or-exact-text matcher, evidence soup-checked and capped; rejects
naming the unverified criteria; entries render deterministically into
qa_notes as '[AC] <criterion> — verified: <evidence>' lines. The old
count-only ac_verdicts gate is superseded (arg kept for back-compat).
- video briefs (structured detail at origination): an enumerable feature
list (release highlights, or input_props.highlights carried onto a
reject re-author) becomes its own scene acceptance criterion, bounded to
the AC caps; a re-author without highlights carries the
feedback-addressed criterion instead.
Extracted findings.py's criterion matcher into shared unmatched_criteria /
uncovered_acceptance_criteria instead of duplicating it; criteria_verified
joins the WAF free-text exclusion set like findings/issues.
* fix(gateway): break the block/unblock wedge — four hardening fixes from the live PM loop
A cell task looped fe-pm/main-pm block/unblock for hours (10 cycles, 43
spawns): a transient GitHub API error resolving CI became an unwaivable
blocker finding whose own fix text said no code change was required, the
submit freshness guard then demanded a commit no finding called for,
escalate_up auto-blocked, and main-pm's correct recovery plan 422'd on
the approach length cap, degrading it to a bare unblock. Four fixes:
- pr_pass CI-unresolvable refusal is now explicitly transient-worded:
retry pr_pass shortly, do NOT pr_fail over a CI-status lookup error —
a platform blip is not a code finding
- submit freshness guard grants ONE unchanged-head resubmission per
head sha when the findings ledger has zero open rows (all addressed
without code changes) — stamped via the resubmit_unchanged_head
marker so the same head can never loop a second time
- unblock carries a flip breaker: block_flip_count marker, and at the
third flip a one-shot CEO notification flags the task as structurally
wedged (unblock itself still succeeds — the breaker signals, it does
not wedge recovery)
- i_will_plan's approach cap truncates at 800 chars instead of
rejecting — an over-detailed plan must never cost the PM its turn
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
This commit is contained in:
@@ -8,6 +8,7 @@ from uuid import uuid4
|
||||
import pytest
|
||||
from pydantic import ValidationError
|
||||
from roboco.api.schemas.v1.flow import (
|
||||
_APPROACH_MAX_CHARS,
|
||||
DelegateRequest,
|
||||
IWillPlanRequest,
|
||||
IWillWorkOnRequest,
|
||||
@@ -165,13 +166,35 @@ def test_i_will_plan_request_rejects_overlong_plan() -> None:
|
||||
assert "plan" in str(exc.value)
|
||||
|
||||
|
||||
def test_i_will_plan_request_rejects_overlong_approach() -> None:
|
||||
"""approach >800 chars is rejected — no ceiling was the bloat bug."""
|
||||
def test_i_will_plan_request_truncates_overlong_approach() -> None:
|
||||
"""approach >800 chars is truncated to 800 (797 + "..."), never rejected —
|
||||
a hard 422 here used to throw away an otherwise-good plan and degrade the
|
||||
PM to a bare unblock (live incident)."""
|
||||
req = IWillPlanRequest(
|
||||
task_id=uuid4(),
|
||||
plan="plan",
|
||||
approach="a" * (_APPROACH_MAX_CHARS + 1),
|
||||
)
|
||||
assert len(req.approach) == _APPROACH_MAX_CHARS
|
||||
assert req.approach.endswith("...")
|
||||
assert req.approach == "a" * (_APPROACH_MAX_CHARS - 3) + "..."
|
||||
|
||||
|
||||
def test_i_will_plan_request_approach_exactly_800_untruncated() -> None:
|
||||
"""Exactly the ceiling is the boundary — passes through byte-for-byte."""
|
||||
approach = "a" * _APPROACH_MAX_CHARS
|
||||
req = IWillPlanRequest(task_id=uuid4(), plan="plan", approach=approach)
|
||||
assert req.approach == approach
|
||||
assert not req.approach.endswith("...")
|
||||
|
||||
|
||||
def test_i_will_plan_request_rejects_thin_approach() -> None:
|
||||
"""approach <150 chars is still a hard reject — a thin plan IS a defect."""
|
||||
with pytest.raises(ValidationError) as exc:
|
||||
IWillPlanRequest(
|
||||
task_id=uuid4(),
|
||||
plan="plan",
|
||||
approach="a" * 801,
|
||||
approach="a" * 149,
|
||||
)
|
||||
assert "approach" in str(exc.value)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user