feat(tasks): task-content guardrails — structured plans + constraints split (#328)

* feat(tasks): task-content guardrails — structured plans + constraints split

Bound task PLANNING content the way journals/notes already are, fixing the
poor task quality flagged 2026-07-07 (degenerate roots, over-decomposed
leaves, descriptions bloated by an auto-attached conventions dump).

Phase A — plan/AC guardrails (no migration):
- _pm_sub_tasks_gate: cap sub_tasks at 7; per-subtask ceilings (title <=200,
  description <=600) enforced at both the Pydantic boundary and the gate.
  Dropped the min-2-roots and no-subtasks-on-code rules: both contradict the
  2026-05-08 rule (test_cell_pm_can_plan_code_typed_parent_via_i_will_plan)
  and break legitimate single-cell roots. Long comment in the gate explains.
- IWillPlanRequest: plan <=2000, approach <=800 (floor 150 kept), typed
  SubTaskCreate/RiskCreate/OpenQuestionCreate replacing loose list[dict].
- DelegateRequest + task_completeness: acceptance_criteria capped at 7 items,
  each <=200 chars. New FieldRule.MAX_LENGTH_LIST + _post_rule_reject helper
  (extracted to keep the gate under xenon B).
- Routes dump typed models to dicts for the existing rich_plan shaper.

Phase B — conventions split (migration 068):
- New nullable tasks.constraints Text column; _attach_baseline_constraints
  now writes the ## Constraints block there instead of appending to
  description, so description is the human-authored instruction only. The
  conventions still reach the agent independently at spawn via the ambient
  block, so agent correctness is unaffected.
- TaskResponse / Task model / panel Task type carry constraints; panel shows
  a read-only Constraints card. Field is optional on the TS type (backend
  returns null for flag-off / pre-migration rows).

Tests: 5 new gate unit tests, 7 schema tests, 3 AC policy tests, 3 e2e smoke
scenarios; 4 baseline-constraints integration tests updated. ruff/mypy/xenon
clean; 10026 unit+foundation+e2e green; panel typecheck clean.

Refs: plan breezy-imagining-kahn

* test(tasks): use typed SubTaskCreate instead of dict literals in plan tests

make quality runs mypy over tests/ (1079 files), not just roboco/ — the
four sites passing dict literals to the now-typed sub_tasks: list[SubTaskCreate]
field failed mypy. Construct SubTaskCreate directly; the typed model raising
ValidationError IS the boundary the rejection tests assert.

* fix(deps): drop unused python-jose — clears PYSEC-2026-1325 (ecdsa, no fix)

CI's pip-audit went red on a freshly-published advisory PYSEC-2026-1325
against ecdsa 0.19.2 (no fix published — 0.19.2 is the latest). ecdsa is a
transitive dep of python-jose, which is a DIRECT dep of roboco but is NOT
imported anywhere in roboco/ or tests/ (grep-verified). The actual JWT path
uses PyJWT (import jwt) + fastapi_users.jwt, not python-jose.

So python-jose is a dead dependency. Removing it (deletion over an
--ignore-vuln waiver) drops ecdsa + rsa + pyasn1 + their type stubs from the
lockfile, eliminating the CVE at the source. deptry roboco/ stays clean
(no missing-dep), mypy clean, auth + schema tests pass.

Master CI was green 9h before this PR's run, so the advisory published in
that window would red any run including master — this fix unblocks both.

* chore(prompts): regenerate verb tables for typed plan sub_tasks

Phase A's IWillPlanRequest schema change (sub_tasks/risks/open_questions from
loose list[dict] to typed SubTaskCreate/RiskCreate/OpenQuestionCreate) made
the auto-generated verb tables stale. Regenerated via
scripts/regenerate_verb_tables.py — the diff is purely the signature
reflection (list[str|str] -> list[SubTaskCreate], etc.). Required by the
foundation-check gate (Makefile:559).

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
This commit is contained in:
Renzo F
2026-07-08 02:01:23 +02:00
committed by GitHub
co-authored by Renn F
parent 92ab13bce0
commit 2a9d9e25d9
23 changed files with 789 additions and 139 deletions
@@ -13,7 +13,7 @@ from fastapi import FastAPI
from fastapi.testclient import TestClient
from roboco.api.deps import get_choreographer
from roboco.api.routes.v1.flow_main_pm import router
from roboco.api.schemas.v1.flow import IWillPlanRequest
from roboco.api.schemas.v1.flow import IWillPlanRequest, SubTaskCreate
_AGENT_ID = "00000000-0000-0000-0004-000000000001"
_HEADERS = {"X-Agent-ID": _AGENT_ID, "X-Agent-Role": "main_pm"}
@@ -111,7 +111,9 @@ def test_i_will_plan_schema_accepts_rich_plan() -> None:
task_id=uuid4(),
plan="Route to backend",
approach=_GOOD_APPROACH,
sub_tasks=[{"title": "Backend slice", "description": _GOOD_SUBTASK_DESC}],
sub_tasks=[
SubTaskCreate(title="Backend slice", description=_GOOD_SUBTASK_DESC)
],
risks=[],
open_questions=[],
)
+99
View File
@@ -11,6 +11,7 @@ from roboco.api.schemas.v1.flow import (
DelegateRequest,
IWillPlanRequest,
IWillWorkOnRequest,
SubTaskCreate,
)
from roboco.models.base import Complexity
@@ -146,3 +147,101 @@ def test_strlist_drops_non_string_junk_instead_of_crashing() -> None:
technical_considerations=technical_considerations,
)
assert req.technical_considerations == ["real note"]
# ---------------------------------------------------------------------------
# Task-content guardrails (2026-07-07): ceilings on plan content + AC caps.
# ---------------------------------------------------------------------------
def test_i_will_plan_request_rejects_overlong_plan() -> None:
"""plan >2000 chars is rejected at the boundary — the bloat defect."""
with pytest.raises(ValidationError) as exc:
IWillPlanRequest(
task_id=uuid4(),
plan="x" * 2001,
approach="a" * 150,
)
assert "plan" in str(exc.value)
def test_i_will_plan_request_rejects_overlong_approach() -> None:
"""approach >800 chars is rejected — no ceiling was the bloat bug."""
with pytest.raises(ValidationError) as exc:
IWillPlanRequest(
task_id=uuid4(),
plan="plan",
approach="a" * 801,
)
assert "approach" in str(exc.value)
def test_i_will_plan_request_rejects_overlong_subtask_title() -> None:
with pytest.raises(ValidationError):
IWillPlanRequest(
task_id=uuid4(),
plan="plan",
approach="a" * 150,
sub_tasks=[SubTaskCreate(title="t" * 201, description="d" * 30)],
)
def test_i_will_plan_request_rejects_overlong_subtask_description() -> None:
with pytest.raises(ValidationError):
IWillPlanRequest(
task_id=uuid4(),
plan="plan",
approach="a" * 150,
sub_tasks=[SubTaskCreate(title="ok title", description="d" * 601)],
)
def test_i_will_plan_request_rejects_thin_subtask_description() -> None:
"""A sub_task description <20 chars fails the typed SubTaskCreate model."""
with pytest.raises(ValidationError):
IWillPlanRequest(
task_id=uuid4(),
plan="plan",
approach="a" * 150,
sub_tasks=[SubTaskCreate(title="ok title", description="too short")],
)
def test_delegate_request_rejects_overlong_acceptance_criterion() -> None:
"""An AC item >200 chars is rejected — a criterion that long is a restated
description, not a verifiable outcome."""
long_ac = "x" * 201
with pytest.raises(ValidationError) as exc:
DelegateRequest.model_validate(
{
"parent_task_id": uuid4(),
"title": "t",
"description": "add the new endpoint plus tests",
"assigned_to": "be-dev-1",
"team": "backend",
"task_type": "code",
"nature": "technical",
"estimated_complexity": "medium",
"acceptance_criteria": [long_ac],
}
)
assert "acceptance_criteria" in str(exc.value)
def test_delegate_request_rejects_too_many_acceptance_criteria() -> None:
"""An AC list >7 items is rejected — over-decomposition of criteria."""
with pytest.raises(ValidationError) as exc:
DelegateRequest.model_validate(
{
"parent_task_id": uuid4(),
"title": "t",
"description": "add the new endpoint plus tests",
"assigned_to": "be-dev-1",
"team": "backend",
"task_type": "code",
"nature": "technical",
"estimated_complexity": "medium",
"acceptance_criteria": [f"criterion {i}" for i in range(8)],
}
)
assert "acceptance_criteria" in str(exc.value)
@@ -32,6 +32,8 @@ _GOOD_SUBTASK_DESC = (
"README H1 leaving the rest untouched, commits with the task-id prefix, "
"and opens the leaf PR for QA."
)
_LONG_SUBTASK_DESC = "x" * 700 # exceeds _PM_SUBTASK_DESC_MAX_LEN (600)
_OVERLONG_TITLE = "t" * 250 # exceeds _PM_SUBTASK_TITLE_MAX_LEN (200)
# ---------------------------------------------------------------------------
@@ -429,3 +431,199 @@ async def test_pm_reentry_in_progress_short_circuits_before_gate() -> None:
"The gate is firing before _handle_pm_reentry — ordering bug not fixed."
)
assert body.get("status") == "in_progress", body
# ---------------------------------------------------------------------------
# Test 7: over-decomposition cap — >7 sub_tasks rejected (2026-07-07 defect)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_pm_with_too_many_subtasks_gets_incomplete_input() -> None:
"""A plan with >7 sub_tasks is over-decomposition — rejected with a hint
to split into sibling coordination tasks instead of one giant plan."""
pm_id = uuid4()
task_id = uuid4()
task_svc = _pm_task_svc(task_id, role="cell_pm")
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_plan(
pm_id,
task_id,
plan="decompose work",
rich_plan={
"approach": _GOOD_APPROACH,
"sub_tasks": [
{"title": f"Slice {i}", "description": _GOOD_SUBTASK_DESC}
for i in range(8) # 8 > _PM_SUBTASKS_MAX (7)
],
},
)
body = env.as_dict()
assert body["error"] == "incomplete_input", body
assert "sub_tasks" in (body.get("missing") or []), body
assert "at most 7" in str(body.get("field_hints", {})), body
@pytest.mark.asyncio
async def test_pm_with_seven_subtasks_passes_cap() -> None:
"""Exactly 7 sub_tasks is the cap — must not be rejected by the cap."""
pm_id = uuid4()
task_id = uuid4()
task_svc = _pm_task_svc(task_id, role="cell_pm")
claimed = MagicMock(
id=task_id,
status="claimed",
plan=None,
assigned_to=pm_id,
task_type="planning",
)
started = MagicMock(
id=task_id,
status="in_progress",
plan={"text": "x"},
assigned_to=pm_id,
task_type="planning",
)
task_svc.claim.return_value = claimed
task_svc.set_plan.return_value = claimed
task_svc.start.return_value = started
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_plan(
pm_id,
task_id,
plan="decompose work",
rich_plan={
"approach": _GOOD_APPROACH,
"sub_tasks": [
{"title": f"Slice {i}", "description": _GOOD_SUBTASK_DESC}
for i in range(7) # exactly the cap
],
},
)
body = env.as_dict()
assert body.get("error") != "incomplete_input", body
# ---------------------------------------------------------------------------
# Test 8: per-sub_task ceilings — over-long title / description rejected
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_pm_with_overlong_subtask_title_gets_incomplete_input() -> None:
"""A sub_task title >200 chars is rejected — titles are single-line labels."""
pm_id = uuid4()
task_id = uuid4()
task_svc = _pm_task_svc(task_id, role="cell_pm")
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_plan(
pm_id,
task_id,
plan="decompose work",
rich_plan={
"approach": _GOOD_APPROACH,
"sub_tasks": [
{"title": _OVERLONG_TITLE, "description": _GOOD_SUBTASK_DESC}
],
},
)
body = env.as_dict()
assert body["error"] == "incomplete_input", body
assert "sub_tasks" in (body.get("missing") or []), body
assert "title is too long" in str(body.get("field_hints", {})), body
@pytest.mark.asyncio
async def test_pm_with_overlong_subtask_description_gets_incomplete_input() -> None:
"""A sub_task description >600 chars is rejected — over-long sub-task
prose is the bloat defect (descriptions dominated by one giant step)."""
pm_id = uuid4()
task_id = uuid4()
task_svc = _pm_task_svc(task_id, role="cell_pm")
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_plan(
pm_id,
task_id,
plan="decompose work",
rich_plan={
"approach": _GOOD_APPROACH,
"sub_tasks": [
{"title": "Backend slice", "description": _LONG_SUBTASK_DESC}
],
},
)
body = env.as_dict()
assert body["error"] == "incomplete_input", body
assert "sub_tasks" in (body.get("missing") or []), body
assert "description is too long" in str(body.get("field_hints", {})), body
# ---------------------------------------------------------------------------
# Test 9: a code-typed parent WITH sub_tasks is NOT rejected — PMs legitimately
# plan code-typed parents into dev subtasks (the 2026-05-08 rule). The gate
# must not ban sub_tasks on code tasks.
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_pm_can_plan_code_typed_parent_with_subtasks() -> None:
"""A code-typed task planned by a PM with sub_tasks is legitimate
decomposition (test_cell_pm_can_plan_code_typed_parent_via_i_will_plan
establishes the 2026-05-08 rule). The cap/ceilings gate must NOT reject
merely because task_type is 'code'."""
pm_id = uuid4()
task_id = uuid4()
task_svc = AsyncMock()
task_svc.get.return_value = MagicMock(
id=task_id,
status="pending",
plan=None,
assigned_to=None,
task_type="code",
parent_task_id=None,
sequence=0,
team="backend",
commits=[],
pr_number=None,
branch_name=None,
quick_context=None,
)
task_svc.agent_for.return_value = MagicMock(
id=pm_id, role="cell_pm", team="backend", slug=None
)
task_svc.list_in_progress_for_agent.return_value = []
task_svc.list_paused_for_agent.return_value = []
task_svc.get_subtasks.return_value = []
task_svc.session = MagicMock()
task_svc.session.begin_nested = MagicMock(
return_value=MagicMock(
__aenter__=AsyncMock(return_value=None),
__aexit__=AsyncMock(return_value=False),
)
)
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_plan(
pm_id,
task_id,
plan="decompose the code parent",
rich_plan={
"approach": _GOOD_APPROACH,
"sub_tasks": [
{"title": "API slice", "description": _GOOD_SUBTASK_DESC},
{"title": "Test slice", "description": _GOOD_SUBTASK_DESC},
],
},
)
body = env.as_dict()
# The gate must not reject merely for task_type='code' with sub_tasks.
assert body.get("error") != "incomplete_input", body