Files
roboco/tests/e2e_smoke/test_pm_plan_guardrails.py
T
2a9d9e25d9 feat(tasks): task-content guardrails — structured plans + constraints split (#328)
* feat(tasks): task-content guardrails — structured plans + constraints split

Bound task PLANNING content the way journals/notes already are, fixing the
poor task quality flagged 2026-07-07 (degenerate roots, over-decomposed
leaves, descriptions bloated by an auto-attached conventions dump).

Phase A — plan/AC guardrails (no migration):
- _pm_sub_tasks_gate: cap sub_tasks at 7; per-subtask ceilings (title <=200,
  description <=600) enforced at both the Pydantic boundary and the gate.
  Dropped the min-2-roots and no-subtasks-on-code rules: both contradict the
  2026-05-08 rule (test_cell_pm_can_plan_code_typed_parent_via_i_will_plan)
  and break legitimate single-cell roots. Long comment in the gate explains.
- IWillPlanRequest: plan <=2000, approach <=800 (floor 150 kept), typed
  SubTaskCreate/RiskCreate/OpenQuestionCreate replacing loose list[dict].
- DelegateRequest + task_completeness: acceptance_criteria capped at 7 items,
  each <=200 chars. New FieldRule.MAX_LENGTH_LIST + _post_rule_reject helper
  (extracted to keep the gate under xenon B).
- Routes dump typed models to dicts for the existing rich_plan shaper.

Phase B — conventions split (migration 068):
- New nullable tasks.constraints Text column; _attach_baseline_constraints
  now writes the ## Constraints block there instead of appending to
  description, so description is the human-authored instruction only. The
  conventions still reach the agent independently at spawn via the ambient
  block, so agent correctness is unaffected.
- TaskResponse / Task model / panel Task type carry constraints; panel shows
  a read-only Constraints card. Field is optional on the TS type (backend
  returns null for flag-off / pre-migration rows).

Tests: 5 new gate unit tests, 7 schema tests, 3 AC policy tests, 3 e2e smoke
scenarios; 4 baseline-constraints integration tests updated. ruff/mypy/xenon
clean; 10026 unit+foundation+e2e green; panel typecheck clean.

Refs: plan breezy-imagining-kahn

* test(tasks): use typed SubTaskCreate instead of dict literals in plan tests

make quality runs mypy over tests/ (1079 files), not just roboco/ — the
four sites passing dict literals to the now-typed sub_tasks: list[SubTaskCreate]
field failed mypy. Construct SubTaskCreate directly; the typed model raising
ValidationError IS the boundary the rejection tests assert.

* fix(deps): drop unused python-jose — clears PYSEC-2026-1325 (ecdsa, no fix)

CI's pip-audit went red on a freshly-published advisory PYSEC-2026-1325
against ecdsa 0.19.2 (no fix published — 0.19.2 is the latest). ecdsa is a
transitive dep of python-jose, which is a DIRECT dep of roboco but is NOT
imported anywhere in roboco/ or tests/ (grep-verified). The actual JWT path
uses PyJWT (import jwt) + fastapi_users.jwt, not python-jose.

So python-jose is a dead dependency. Removing it (deletion over an
--ignore-vuln waiver) drops ecdsa + rsa + pyasn1 + their type stubs from the
lockfile, eliminating the CVE at the source. deptry roboco/ stays clean
(no missing-dep), mypy clean, auth + schema tests pass.

Master CI was green 9h before this PR's run, so the advisory published in
that window would red any run including master — this fix unblocks both.

* chore(prompts): regenerate verb tables for typed plan sub_tasks

Phase A's IWillPlanRequest schema change (sub_tasks/risks/open_questions from
loose list[dict] to typed SubTaskCreate/RiskCreate/OpenQuestionCreate) made
the auto-generated verb tables stale. Regenerated via
scripts/regenerate_verb_tables.py — the diff is purely the signature
reflection (list[str|str] -> list[SubTaskCreate], etc.). Required by the
foundation-check gate (Makefile:559).

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-07-08 02:01:23 +02:00

162 lines
5.9 KiB
Python

"""Scenario: the PM plan-content guardrails reject over-decomposed plans.
The 2026-07-07 task-quality defects (root 79d686f0 decomposed into 1 subtask,
code leaf 55376b8a carrying a 5-subtask plan, descriptions bloated to 3000+
chars) were structural — the planning verb barely guardrailed its content.
The fix adds ceilings (plan <= 2000, approach <= 800, sub-task title <= 200,
sub-task description <= 600) and an over-decomposition cap (>7 sub_tasks
rejected) at both the HTTP Pydantic boundary and the choreographer gate.
This scenario drives a MAIN_PM planning root through the real API and asserts
each guardrail surfaces a clean ``incomplete_input`` envelope with a
remediation hint the agent can act on — not a 500, not a silent accept. The
gate is the load-bearing layer (direct service callers bypass Pydantic), so
the assertions land on the envelope, not the HTTP status.
"""
from __future__ import annotations
from typing import TYPE_CHECKING, Any
from tests.e2e_smoke.arcs import (
origin_branch,
seed_company,
seed_project,
seed_task,
set_branch_name,
)
from tests.e2e_smoke.harness import ScriptedAgent, expect_error
if TYPE_CHECKING:
from uuid import UUID
from tests.e2e_smoke.arcs import Company
from tests.e2e_smoke.harness import E2EStack
# Pydantic's IWillPlanRequest.approach enforces 150..800 chars at the HTTP
# boundary, so every i_will_plan call needs a compliant approach.
_APPROACH = (
"Plan and delegate the page-scoped refresh button work to the frontend "
"cell: land the provider/hook, add the navbar button, remove the inline "
"buttons, and route one planning subtask to fe-pm for delivery. "
"Sequenced strictly; no cross-cell dependencies for this slice."
)
_GOOD_SUB = {
"title": "Frontend cell: refresh button",
"description": (
"Delegate the navbar refresh button to fe-pm: land the provider/hook "
"and wire the click handler into the page, then open the leaf PR."
),
}
_PLAN = "Land the refresh button via the frontend cell."
def _seed_planning_root(
stack: E2EStack, company: Company
) -> tuple[ScriptedAgent, UUID]:
"""Seed a PENDING MAIN_PM planning root + its origin branch."""
from roboco.models import Team
from roboco.models.base import TaskStatus, TaskType
project_id, _project_slug = seed_project(stack, company)
main_pm = ScriptedAgent(stack, company.main_pm_id, "main-pm", "main_pm")
task_id = seed_task(
stack,
title="Root: page-scoped refresh button",
description="Frontend-only root: provider/hook + navbar button.",
acceptance_criteria=["the refresh button lands on master"],
task_type=TaskType.PLANNING,
team=Team.MAIN_PM,
project_id=project_id,
created_by=company.main_pm_id,
assigned_to=company.main_pm_id,
status=TaskStatus.PENDING,
)
branch = f"feature/main_pm/{str(task_id)[:8]}"
origin_branch(stack, branch, start="master")
set_branch_name(stack, task_id, branch)
return main_pm, task_id
def _plan_with(sub_tasks: list[dict[str, Any]]) -> dict[str, Any]:
return {
"plan": _PLAN,
"approach": _APPROACH,
"sub_tasks": sub_tasks,
}
def test_pm_plan_over_decomposition_cap_rejected(e2e_stack: E2EStack) -> None:
"""A plan with >7 sub_tasks is over-decomposition — the gate rejects it
with incomplete_input + a 'split into sibling coordination tasks' hint."""
stack = e2e_stack
company = seed_company(stack)
main_pm, task_id = _seed_planning_root(stack, company)
env = main_pm.flow(
"i_will_plan",
task_id=str(task_id),
plan=_PLAN,
approach=_APPROACH,
sub_tasks=[dict(_GOOD_SUB, title=f"Slice {i}") for i in range(8)],
)
body = expect_error(env, "incomplete_input", "8 sub_tasks rejected")
assert "sub_tasks" in (body.get("missing") or []), body
assert "at most 7" in str(body.get("field_hints", {})), body
def test_pm_plan_overlong_subtask_description_rejected(
e2e_stack: E2EStack,
) -> None:
"""A sub-task description >600 chars is the bloat defect — rejected at
the Pydantic boundary (422), so the envelope carries the validation
``detail`` rather than the gate's ``field_hints``. Both layers are the
guardrail working; this test pins the boundary layer."""
stack = e2e_stack
company = seed_company(stack)
main_pm, task_id = _seed_planning_root(stack, company)
bloated = dict(_GOOD_SUB, description="x" * 700)
env = main_pm.flow(
"i_will_plan",
task_id=str(task_id),
plan=_PLAN,
approach=_APPROACH,
sub_tasks=[bloated],
)
body = expect_error(env, "incomplete_input", "over-long subtask desc")
# Boundary 422: missing is [] but detail carries the Pydantic error
# naming sub_tasks + the 600-char cap.
detail = str(body.get("detail"))
assert "sub_tasks" in detail, body
assert "600" in detail, body
def test_pm_plan_valid_plan_passes_gate(e2e_stack: E2EStack) -> None:
"""A well-formed plan (2 sub_tasks, bounded fields) passes the guardrails
and transitions the root to in_progress — the happy path stays green."""
stack = e2e_stack
company = seed_company(stack)
main_pm, task_id = _seed_planning_root(stack, company)
env = main_pm.flow(
"i_will_plan",
task_id=str(task_id),
plan=_PLAN,
approach=_APPROACH,
sub_tasks=[
_GOOD_SUB,
{
"title": "Frontend cell: remove inline buttons",
"description": (
"Delegate removal of the stale inline refresh buttons to "
"fe-pm so the navbar button is the single source of truth."
),
},
],
)
# The gate must not fire; the root moves to in_progress. Downstream may
# raise a different error (e.g. tracing_gap) but NOT incomplete_input.
body = env
assert body.get("error") != "incomplete_input", body