mirror of
https://github.com/rennf93/roboco.git
synced 2026-08-03 07:23:24 +02:00
feat(tasks): task-content guardrails — structured plans + constraints split (#328)
* feat(tasks): task-content guardrails — structured plans + constraints split Bound task PLANNING content the way journals/notes already are, fixing the poor task quality flagged 2026-07-07 (degenerate roots, over-decomposed leaves, descriptions bloated by an auto-attached conventions dump). Phase A — plan/AC guardrails (no migration): - _pm_sub_tasks_gate: cap sub_tasks at 7; per-subtask ceilings (title <=200, description <=600) enforced at both the Pydantic boundary and the gate. Dropped the min-2-roots and no-subtasks-on-code rules: both contradict the 2026-05-08 rule (test_cell_pm_can_plan_code_typed_parent_via_i_will_plan) and break legitimate single-cell roots. Long comment in the gate explains. - IWillPlanRequest: plan <=2000, approach <=800 (floor 150 kept), typed SubTaskCreate/RiskCreate/OpenQuestionCreate replacing loose list[dict]. - DelegateRequest + task_completeness: acceptance_criteria capped at 7 items, each <=200 chars. New FieldRule.MAX_LENGTH_LIST + _post_rule_reject helper (extracted to keep the gate under xenon B). - Routes dump typed models to dicts for the existing rich_plan shaper. Phase B — conventions split (migration 068): - New nullable tasks.constraints Text column; _attach_baseline_constraints now writes the ## Constraints block there instead of appending to description, so description is the human-authored instruction only. The conventions still reach the agent independently at spawn via the ambient block, so agent correctness is unaffected. - TaskResponse / Task model / panel Task type carry constraints; panel shows a read-only Constraints card. Field is optional on the TS type (backend returns null for flag-off / pre-migration rows). Tests: 5 new gate unit tests, 7 schema tests, 3 AC policy tests, 3 e2e smoke scenarios; 4 baseline-constraints integration tests updated. ruff/mypy/xenon clean; 10026 unit+foundation+e2e green; panel typecheck clean. Refs: plan breezy-imagining-kahn * test(tasks): use typed SubTaskCreate instead of dict literals in plan tests make quality runs mypy over tests/ (1079 files), not just roboco/ — the four sites passing dict literals to the now-typed sub_tasks: list[SubTaskCreate] field failed mypy. Construct SubTaskCreate directly; the typed model raising ValidationError IS the boundary the rejection tests assert. * fix(deps): drop unused python-jose — clears PYSEC-2026-1325 (ecdsa, no fix) CI's pip-audit went red on a freshly-published advisory PYSEC-2026-1325 against ecdsa 0.19.2 (no fix published — 0.19.2 is the latest). ecdsa is a transitive dep of python-jose, which is a DIRECT dep of roboco but is NOT imported anywhere in roboco/ or tests/ (grep-verified). The actual JWT path uses PyJWT (import jwt) + fastapi_users.jwt, not python-jose. So python-jose is a dead dependency. Removing it (deletion over an --ignore-vuln waiver) drops ecdsa + rsa + pyasn1 + their type stubs from the lockfile, eliminating the CVE at the source. deptry roboco/ stays clean (no missing-dep), mypy clean, auth + schema tests pass. Master CI was green 9h before this PR's run, so the advisory published in that window would red any run including master — this fix unblocks both. * chore(prompts): regenerate verb tables for typed plan sub_tasks Phase A's IWillPlanRequest schema change (sub_tasks/risks/open_questions from loose list[dict] to typed SubTaskCreate/RiskCreate/OpenQuestionCreate) made the auto-generated verb tables stale. Regenerated via scripts/regenerate_verb_tables.py — the diff is purely the signature reflection (list[str|str] -> list[SubTaskCreate], etc.). Required by the foundation-check gate (Makefile:559). --------- Co-authored-by: Renn F <rennf93@users.noreply.github.com>
This commit is contained in:
@@ -0,0 +1,161 @@
|
||||
"""Scenario: the PM plan-content guardrails reject over-decomposed plans.
|
||||
|
||||
The 2026-07-07 task-quality defects (root 79d686f0 decomposed into 1 subtask,
|
||||
code leaf 55376b8a carrying a 5-subtask plan, descriptions bloated to 3000+
|
||||
chars) were structural — the planning verb barely guardrailed its content.
|
||||
The fix adds ceilings (plan <= 2000, approach <= 800, sub-task title <= 200,
|
||||
sub-task description <= 600) and an over-decomposition cap (>7 sub_tasks
|
||||
rejected) at both the HTTP Pydantic boundary and the choreographer gate.
|
||||
|
||||
This scenario drives a MAIN_PM planning root through the real API and asserts
|
||||
each guardrail surfaces a clean ``incomplete_input`` envelope with a
|
||||
remediation hint the agent can act on — not a 500, not a silent accept. The
|
||||
gate is the load-bearing layer (direct service callers bypass Pydantic), so
|
||||
the assertions land on the envelope, not the HTTP status.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import TYPE_CHECKING, Any
|
||||
|
||||
from tests.e2e_smoke.arcs import (
|
||||
origin_branch,
|
||||
seed_company,
|
||||
seed_project,
|
||||
seed_task,
|
||||
set_branch_name,
|
||||
)
|
||||
from tests.e2e_smoke.harness import ScriptedAgent, expect_error
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from uuid import UUID
|
||||
|
||||
from tests.e2e_smoke.arcs import Company
|
||||
from tests.e2e_smoke.harness import E2EStack
|
||||
|
||||
# Pydantic's IWillPlanRequest.approach enforces 150..800 chars at the HTTP
|
||||
# boundary, so every i_will_plan call needs a compliant approach.
|
||||
_APPROACH = (
|
||||
"Plan and delegate the page-scoped refresh button work to the frontend "
|
||||
"cell: land the provider/hook, add the navbar button, remove the inline "
|
||||
"buttons, and route one planning subtask to fe-pm for delivery. "
|
||||
"Sequenced strictly; no cross-cell dependencies for this slice."
|
||||
)
|
||||
_GOOD_SUB = {
|
||||
"title": "Frontend cell: refresh button",
|
||||
"description": (
|
||||
"Delegate the navbar refresh button to fe-pm: land the provider/hook "
|
||||
"and wire the click handler into the page, then open the leaf PR."
|
||||
),
|
||||
}
|
||||
_PLAN = "Land the refresh button via the frontend cell."
|
||||
|
||||
|
||||
def _seed_planning_root(
|
||||
stack: E2EStack, company: Company
|
||||
) -> tuple[ScriptedAgent, UUID]:
|
||||
"""Seed a PENDING MAIN_PM planning root + its origin branch."""
|
||||
from roboco.models import Team
|
||||
from roboco.models.base import TaskStatus, TaskType
|
||||
|
||||
project_id, _project_slug = seed_project(stack, company)
|
||||
main_pm = ScriptedAgent(stack, company.main_pm_id, "main-pm", "main_pm")
|
||||
task_id = seed_task(
|
||||
stack,
|
||||
title="Root: page-scoped refresh button",
|
||||
description="Frontend-only root: provider/hook + navbar button.",
|
||||
acceptance_criteria=["the refresh button lands on master"],
|
||||
task_type=TaskType.PLANNING,
|
||||
team=Team.MAIN_PM,
|
||||
project_id=project_id,
|
||||
created_by=company.main_pm_id,
|
||||
assigned_to=company.main_pm_id,
|
||||
status=TaskStatus.PENDING,
|
||||
)
|
||||
branch = f"feature/main_pm/{str(task_id)[:8]}"
|
||||
origin_branch(stack, branch, start="master")
|
||||
set_branch_name(stack, task_id, branch)
|
||||
return main_pm, task_id
|
||||
|
||||
|
||||
def _plan_with(sub_tasks: list[dict[str, Any]]) -> dict[str, Any]:
|
||||
return {
|
||||
"plan": _PLAN,
|
||||
"approach": _APPROACH,
|
||||
"sub_tasks": sub_tasks,
|
||||
}
|
||||
|
||||
|
||||
def test_pm_plan_over_decomposition_cap_rejected(e2e_stack: E2EStack) -> None:
|
||||
"""A plan with >7 sub_tasks is over-decomposition — the gate rejects it
|
||||
with incomplete_input + a 'split into sibling coordination tasks' hint."""
|
||||
stack = e2e_stack
|
||||
company = seed_company(stack)
|
||||
main_pm, task_id = _seed_planning_root(stack, company)
|
||||
|
||||
env = main_pm.flow(
|
||||
"i_will_plan",
|
||||
task_id=str(task_id),
|
||||
plan=_PLAN,
|
||||
approach=_APPROACH,
|
||||
sub_tasks=[dict(_GOOD_SUB, title=f"Slice {i}") for i in range(8)],
|
||||
)
|
||||
body = expect_error(env, "incomplete_input", "8 sub_tasks rejected")
|
||||
assert "sub_tasks" in (body.get("missing") or []), body
|
||||
assert "at most 7" in str(body.get("field_hints", {})), body
|
||||
|
||||
|
||||
def test_pm_plan_overlong_subtask_description_rejected(
|
||||
e2e_stack: E2EStack,
|
||||
) -> None:
|
||||
"""A sub-task description >600 chars is the bloat defect — rejected at
|
||||
the Pydantic boundary (422), so the envelope carries the validation
|
||||
``detail`` rather than the gate's ``field_hints``. Both layers are the
|
||||
guardrail working; this test pins the boundary layer."""
|
||||
stack = e2e_stack
|
||||
company = seed_company(stack)
|
||||
main_pm, task_id = _seed_planning_root(stack, company)
|
||||
|
||||
bloated = dict(_GOOD_SUB, description="x" * 700)
|
||||
env = main_pm.flow(
|
||||
"i_will_plan",
|
||||
task_id=str(task_id),
|
||||
plan=_PLAN,
|
||||
approach=_APPROACH,
|
||||
sub_tasks=[bloated],
|
||||
)
|
||||
body = expect_error(env, "incomplete_input", "over-long subtask desc")
|
||||
# Boundary 422: missing is [] but detail carries the Pydantic error
|
||||
# naming sub_tasks + the 600-char cap.
|
||||
detail = str(body.get("detail"))
|
||||
assert "sub_tasks" in detail, body
|
||||
assert "600" in detail, body
|
||||
|
||||
|
||||
def test_pm_plan_valid_plan_passes_gate(e2e_stack: E2EStack) -> None:
|
||||
"""A well-formed plan (2 sub_tasks, bounded fields) passes the guardrails
|
||||
and transitions the root to in_progress — the happy path stays green."""
|
||||
stack = e2e_stack
|
||||
company = seed_company(stack)
|
||||
main_pm, task_id = _seed_planning_root(stack, company)
|
||||
|
||||
env = main_pm.flow(
|
||||
"i_will_plan",
|
||||
task_id=str(task_id),
|
||||
plan=_PLAN,
|
||||
approach=_APPROACH,
|
||||
sub_tasks=[
|
||||
_GOOD_SUB,
|
||||
{
|
||||
"title": "Frontend cell: remove inline buttons",
|
||||
"description": (
|
||||
"Delegate removal of the stale inline refresh buttons to "
|
||||
"fe-pm so the navbar button is the single source of truth."
|
||||
),
|
||||
},
|
||||
],
|
||||
)
|
||||
# The gate must not fire; the root moves to in_progress. Downstream may
|
||||
# raise a different error (e.g. tracing_gap) but NOT incomplete_input.
|
||||
body = env
|
||||
assert body.get("error") != "incomplete_input", body
|
||||
@@ -152,3 +152,34 @@ def test_fill_parent_from_active_task_does_not_overwrite_explicit() -> None:
|
||||
payload = {"parent_task_id": "explicit-id"}
|
||||
result = tc.fill_parent_from_active_task(payload, "active-id")
|
||||
assert result["parent_task_id"] == "explicit-id"
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# AC discipline (2026-07-07 task-quality defect): per-item cap + max count.
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_task_at_create_rejects_too_many_acceptance_criteria() -> None:
|
||||
"""An AC list >7 items is over-decomposition of criteria."""
|
||||
result = tc.check(
|
||||
tc.TASK_AT_CREATE, _task(acceptance_criteria=[f"c{i}" for i in range(8)])
|
||||
)
|
||||
assert result.passed is False
|
||||
assert "acceptance_criteria" in result.missing
|
||||
|
||||
|
||||
def test_task_at_create_accepts_seven_acceptance_criteria() -> None:
|
||||
"""Exactly 7 AC items is the cap — must pass."""
|
||||
result = tc.check(
|
||||
tc.TASK_AT_CREATE, _task(acceptance_criteria=[f"c{i}" for i in range(7)])
|
||||
)
|
||||
assert result.passed is True
|
||||
|
||||
|
||||
def test_task_at_create_rejects_overlong_acceptance_criterion() -> None:
|
||||
"""An AC item >200 chars is a restated description, not a criterion."""
|
||||
long_ac = "x" * 201
|
||||
result = tc.check(tc.TASK_AT_CREATE, _task(acceptance_criteria=[long_ac]))
|
||||
assert result.passed is False
|
||||
assert "acceptance_criteria" in result.missing
|
||||
assert any("200" in h for h in result.field_hints.values())
|
||||
|
||||
@@ -68,9 +68,12 @@ async def test_baseline_attached_when_flag_on(
|
||||
monkeypatch.setattr(settings, "conventions_enabled", True)
|
||||
agent, project = await _seed(db_session)
|
||||
task = await TaskService(db_session).create(_req(agent, project, "Do the work"))
|
||||
assert task.description is not None
|
||||
assert "## Constraints" in task.description
|
||||
assert "no lint suppressions" in task.description
|
||||
# The conventions block lives in `constraints`, not `description` (the
|
||||
# 2026-07-07 fix: description is the human-authored instruction only).
|
||||
assert task.description == "Do the work"
|
||||
assert task.constraints is not None
|
||||
assert "## Constraints" in task.constraints
|
||||
assert "no lint suppressions" in task.constraints
|
||||
|
||||
|
||||
async def test_flag_off_attaches_nothing(
|
||||
@@ -80,20 +83,24 @@ async def test_flag_off_attaches_nothing(
|
||||
agent, project = await _seed(db_session)
|
||||
task = await TaskService(db_session).create(_req(agent, project, "Do the work"))
|
||||
assert task.description == "Do the work"
|
||||
assert task.constraints is None
|
||||
|
||||
|
||||
async def test_baseline_not_suppressed_by_agent_constraints_section(
|
||||
db_session: AsyncSession, monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
# An agent-authored ## Constraints section must NOT suppress the mandatory
|
||||
# server baseline — both are present.
|
||||
# An agent-authored ## Constraints section in the description must NOT
|
||||
# suppress the mandatory server baseline — the description keeps the
|
||||
# agent's block and `constraints` carries the server baseline (separate
|
||||
# field, so neither suppresses the other).
|
||||
monkeypatch.setattr(settings, "conventions_enabled", True)
|
||||
agent, project = await _seed(db_session)
|
||||
seeded = "Do the work\n\n## Constraints\n- a task-specific note"
|
||||
task = await TaskService(db_session).create(_req(agent, project, seeded))
|
||||
assert task.description is not None
|
||||
assert "a task-specific note" in task.description
|
||||
assert "no lint suppressions" in task.description
|
||||
assert task.constraints is not None
|
||||
assert "no lint suppressions" in task.constraints
|
||||
|
||||
|
||||
async def test_baseline_attach_is_idempotent(
|
||||
@@ -103,8 +110,8 @@ async def test_baseline_attach_is_idempotent(
|
||||
agent, project = await _seed(db_session)
|
||||
svc = TaskService(db_session)
|
||||
task = await svc.create(_req(agent, project, "Do the work"))
|
||||
before = task.description
|
||||
before = task.constraints
|
||||
await svc._attach_baseline_constraints(task)
|
||||
assert task.description == before
|
||||
assert task.description is not None
|
||||
assert task.description.count("no lint suppressions") == 1
|
||||
assert task.constraints == before
|
||||
assert task.constraints is not None
|
||||
assert task.constraints.count("no lint suppressions") == 1
|
||||
|
||||
@@ -13,7 +13,7 @@ from fastapi import FastAPI
|
||||
from fastapi.testclient import TestClient
|
||||
from roboco.api.deps import get_choreographer
|
||||
from roboco.api.routes.v1.flow_main_pm import router
|
||||
from roboco.api.schemas.v1.flow import IWillPlanRequest
|
||||
from roboco.api.schemas.v1.flow import IWillPlanRequest, SubTaskCreate
|
||||
|
||||
_AGENT_ID = "00000000-0000-0000-0004-000000000001"
|
||||
_HEADERS = {"X-Agent-ID": _AGENT_ID, "X-Agent-Role": "main_pm"}
|
||||
@@ -111,7 +111,9 @@ def test_i_will_plan_schema_accepts_rich_plan() -> None:
|
||||
task_id=uuid4(),
|
||||
plan="Route to backend",
|
||||
approach=_GOOD_APPROACH,
|
||||
sub_tasks=[{"title": "Backend slice", "description": _GOOD_SUBTASK_DESC}],
|
||||
sub_tasks=[
|
||||
SubTaskCreate(title="Backend slice", description=_GOOD_SUBTASK_DESC)
|
||||
],
|
||||
risks=[],
|
||||
open_questions=[],
|
||||
)
|
||||
|
||||
@@ -11,6 +11,7 @@ from roboco.api.schemas.v1.flow import (
|
||||
DelegateRequest,
|
||||
IWillPlanRequest,
|
||||
IWillWorkOnRequest,
|
||||
SubTaskCreate,
|
||||
)
|
||||
from roboco.models.base import Complexity
|
||||
|
||||
@@ -146,3 +147,101 @@ def test_strlist_drops_non_string_junk_instead_of_crashing() -> None:
|
||||
technical_considerations=technical_considerations,
|
||||
)
|
||||
assert req.technical_considerations == ["real note"]
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Task-content guardrails (2026-07-07): ceilings on plan content + AC caps.
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_i_will_plan_request_rejects_overlong_plan() -> None:
|
||||
"""plan >2000 chars is rejected at the boundary — the bloat defect."""
|
||||
with pytest.raises(ValidationError) as exc:
|
||||
IWillPlanRequest(
|
||||
task_id=uuid4(),
|
||||
plan="x" * 2001,
|
||||
approach="a" * 150,
|
||||
)
|
||||
assert "plan" in str(exc.value)
|
||||
|
||||
|
||||
def test_i_will_plan_request_rejects_overlong_approach() -> None:
|
||||
"""approach >800 chars is rejected — no ceiling was the bloat bug."""
|
||||
with pytest.raises(ValidationError) as exc:
|
||||
IWillPlanRequest(
|
||||
task_id=uuid4(),
|
||||
plan="plan",
|
||||
approach="a" * 801,
|
||||
)
|
||||
assert "approach" in str(exc.value)
|
||||
|
||||
|
||||
def test_i_will_plan_request_rejects_overlong_subtask_title() -> None:
|
||||
with pytest.raises(ValidationError):
|
||||
IWillPlanRequest(
|
||||
task_id=uuid4(),
|
||||
plan="plan",
|
||||
approach="a" * 150,
|
||||
sub_tasks=[SubTaskCreate(title="t" * 201, description="d" * 30)],
|
||||
)
|
||||
|
||||
|
||||
def test_i_will_plan_request_rejects_overlong_subtask_description() -> None:
|
||||
with pytest.raises(ValidationError):
|
||||
IWillPlanRequest(
|
||||
task_id=uuid4(),
|
||||
plan="plan",
|
||||
approach="a" * 150,
|
||||
sub_tasks=[SubTaskCreate(title="ok title", description="d" * 601)],
|
||||
)
|
||||
|
||||
|
||||
def test_i_will_plan_request_rejects_thin_subtask_description() -> None:
|
||||
"""A sub_task description <20 chars fails the typed SubTaskCreate model."""
|
||||
with pytest.raises(ValidationError):
|
||||
IWillPlanRequest(
|
||||
task_id=uuid4(),
|
||||
plan="plan",
|
||||
approach="a" * 150,
|
||||
sub_tasks=[SubTaskCreate(title="ok title", description="too short")],
|
||||
)
|
||||
|
||||
|
||||
def test_delegate_request_rejects_overlong_acceptance_criterion() -> None:
|
||||
"""An AC item >200 chars is rejected — a criterion that long is a restated
|
||||
description, not a verifiable outcome."""
|
||||
long_ac = "x" * 201
|
||||
with pytest.raises(ValidationError) as exc:
|
||||
DelegateRequest.model_validate(
|
||||
{
|
||||
"parent_task_id": uuid4(),
|
||||
"title": "t",
|
||||
"description": "add the new endpoint plus tests",
|
||||
"assigned_to": "be-dev-1",
|
||||
"team": "backend",
|
||||
"task_type": "code",
|
||||
"nature": "technical",
|
||||
"estimated_complexity": "medium",
|
||||
"acceptance_criteria": [long_ac],
|
||||
}
|
||||
)
|
||||
assert "acceptance_criteria" in str(exc.value)
|
||||
|
||||
|
||||
def test_delegate_request_rejects_too_many_acceptance_criteria() -> None:
|
||||
"""An AC list >7 items is rejected — over-decomposition of criteria."""
|
||||
with pytest.raises(ValidationError) as exc:
|
||||
DelegateRequest.model_validate(
|
||||
{
|
||||
"parent_task_id": uuid4(),
|
||||
"title": "t",
|
||||
"description": "add the new endpoint plus tests",
|
||||
"assigned_to": "be-dev-1",
|
||||
"team": "backend",
|
||||
"task_type": "code",
|
||||
"nature": "technical",
|
||||
"estimated_complexity": "medium",
|
||||
"acceptance_criteria": [f"criterion {i}" for i in range(8)],
|
||||
}
|
||||
)
|
||||
assert "acceptance_criteria" in str(exc.value)
|
||||
|
||||
@@ -32,6 +32,8 @@ _GOOD_SUBTASK_DESC = (
|
||||
"README H1 leaving the rest untouched, commits with the task-id prefix, "
|
||||
"and opens the leaf PR for QA."
|
||||
)
|
||||
_LONG_SUBTASK_DESC = "x" * 700 # exceeds _PM_SUBTASK_DESC_MAX_LEN (600)
|
||||
_OVERLONG_TITLE = "t" * 250 # exceeds _PM_SUBTASK_TITLE_MAX_LEN (200)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
@@ -429,3 +431,199 @@ async def test_pm_reentry_in_progress_short_circuits_before_gate() -> None:
|
||||
"The gate is firing before _handle_pm_reentry — ordering bug not fixed."
|
||||
)
|
||||
assert body.get("status") == "in_progress", body
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Test 7: over-decomposition cap — >7 sub_tasks rejected (2026-07-07 defect)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_pm_with_too_many_subtasks_gets_incomplete_input() -> None:
|
||||
"""A plan with >7 sub_tasks is over-decomposition — rejected with a hint
|
||||
to split into sibling coordination tasks instead of one giant plan."""
|
||||
pm_id = uuid4()
|
||||
task_id = uuid4()
|
||||
task_svc = _pm_task_svc(task_id, role="cell_pm")
|
||||
deps = _make_deps(task=task_svc)
|
||||
c = Choreographer(deps)
|
||||
|
||||
env = await c.i_will_plan(
|
||||
pm_id,
|
||||
task_id,
|
||||
plan="decompose work",
|
||||
rich_plan={
|
||||
"approach": _GOOD_APPROACH,
|
||||
"sub_tasks": [
|
||||
{"title": f"Slice {i}", "description": _GOOD_SUBTASK_DESC}
|
||||
for i in range(8) # 8 > _PM_SUBTASKS_MAX (7)
|
||||
],
|
||||
},
|
||||
)
|
||||
body = env.as_dict()
|
||||
assert body["error"] == "incomplete_input", body
|
||||
assert "sub_tasks" in (body.get("missing") or []), body
|
||||
assert "at most 7" in str(body.get("field_hints", {})), body
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_pm_with_seven_subtasks_passes_cap() -> None:
|
||||
"""Exactly 7 sub_tasks is the cap — must not be rejected by the cap."""
|
||||
pm_id = uuid4()
|
||||
task_id = uuid4()
|
||||
task_svc = _pm_task_svc(task_id, role="cell_pm")
|
||||
claimed = MagicMock(
|
||||
id=task_id,
|
||||
status="claimed",
|
||||
plan=None,
|
||||
assigned_to=pm_id,
|
||||
task_type="planning",
|
||||
)
|
||||
started = MagicMock(
|
||||
id=task_id,
|
||||
status="in_progress",
|
||||
plan={"text": "x"},
|
||||
assigned_to=pm_id,
|
||||
task_type="planning",
|
||||
)
|
||||
task_svc.claim.return_value = claimed
|
||||
task_svc.set_plan.return_value = claimed
|
||||
task_svc.start.return_value = started
|
||||
deps = _make_deps(task=task_svc)
|
||||
c = Choreographer(deps)
|
||||
|
||||
env = await c.i_will_plan(
|
||||
pm_id,
|
||||
task_id,
|
||||
plan="decompose work",
|
||||
rich_plan={
|
||||
"approach": _GOOD_APPROACH,
|
||||
"sub_tasks": [
|
||||
{"title": f"Slice {i}", "description": _GOOD_SUBTASK_DESC}
|
||||
for i in range(7) # exactly the cap
|
||||
],
|
||||
},
|
||||
)
|
||||
body = env.as_dict()
|
||||
assert body.get("error") != "incomplete_input", body
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Test 8: per-sub_task ceilings — over-long title / description rejected
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_pm_with_overlong_subtask_title_gets_incomplete_input() -> None:
|
||||
"""A sub_task title >200 chars is rejected — titles are single-line labels."""
|
||||
pm_id = uuid4()
|
||||
task_id = uuid4()
|
||||
task_svc = _pm_task_svc(task_id, role="cell_pm")
|
||||
deps = _make_deps(task=task_svc)
|
||||
c = Choreographer(deps)
|
||||
|
||||
env = await c.i_will_plan(
|
||||
pm_id,
|
||||
task_id,
|
||||
plan="decompose work",
|
||||
rich_plan={
|
||||
"approach": _GOOD_APPROACH,
|
||||
"sub_tasks": [
|
||||
{"title": _OVERLONG_TITLE, "description": _GOOD_SUBTASK_DESC}
|
||||
],
|
||||
},
|
||||
)
|
||||
body = env.as_dict()
|
||||
assert body["error"] == "incomplete_input", body
|
||||
assert "sub_tasks" in (body.get("missing") or []), body
|
||||
assert "title is too long" in str(body.get("field_hints", {})), body
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_pm_with_overlong_subtask_description_gets_incomplete_input() -> None:
|
||||
"""A sub_task description >600 chars is rejected — over-long sub-task
|
||||
prose is the bloat defect (descriptions dominated by one giant step)."""
|
||||
pm_id = uuid4()
|
||||
task_id = uuid4()
|
||||
task_svc = _pm_task_svc(task_id, role="cell_pm")
|
||||
deps = _make_deps(task=task_svc)
|
||||
c = Choreographer(deps)
|
||||
|
||||
env = await c.i_will_plan(
|
||||
pm_id,
|
||||
task_id,
|
||||
plan="decompose work",
|
||||
rich_plan={
|
||||
"approach": _GOOD_APPROACH,
|
||||
"sub_tasks": [
|
||||
{"title": "Backend slice", "description": _LONG_SUBTASK_DESC}
|
||||
],
|
||||
},
|
||||
)
|
||||
body = env.as_dict()
|
||||
assert body["error"] == "incomplete_input", body
|
||||
assert "sub_tasks" in (body.get("missing") or []), body
|
||||
assert "description is too long" in str(body.get("field_hints", {})), body
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Test 9: a code-typed parent WITH sub_tasks is NOT rejected — PMs legitimately
|
||||
# plan code-typed parents into dev subtasks (the 2026-05-08 rule). The gate
|
||||
# must not ban sub_tasks on code tasks.
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_pm_can_plan_code_typed_parent_with_subtasks() -> None:
|
||||
"""A code-typed task planned by a PM with sub_tasks is legitimate
|
||||
decomposition (test_cell_pm_can_plan_code_typed_parent_via_i_will_plan
|
||||
establishes the 2026-05-08 rule). The cap/ceilings gate must NOT reject
|
||||
merely because task_type is 'code'."""
|
||||
pm_id = uuid4()
|
||||
task_id = uuid4()
|
||||
task_svc = AsyncMock()
|
||||
task_svc.get.return_value = MagicMock(
|
||||
id=task_id,
|
||||
status="pending",
|
||||
plan=None,
|
||||
assigned_to=None,
|
||||
task_type="code",
|
||||
parent_task_id=None,
|
||||
sequence=0,
|
||||
team="backend",
|
||||
commits=[],
|
||||
pr_number=None,
|
||||
branch_name=None,
|
||||
quick_context=None,
|
||||
)
|
||||
task_svc.agent_for.return_value = MagicMock(
|
||||
id=pm_id, role="cell_pm", team="backend", slug=None
|
||||
)
|
||||
task_svc.list_in_progress_for_agent.return_value = []
|
||||
task_svc.list_paused_for_agent.return_value = []
|
||||
task_svc.get_subtasks.return_value = []
|
||||
task_svc.session = MagicMock()
|
||||
task_svc.session.begin_nested = MagicMock(
|
||||
return_value=MagicMock(
|
||||
__aenter__=AsyncMock(return_value=None),
|
||||
__aexit__=AsyncMock(return_value=False),
|
||||
)
|
||||
)
|
||||
deps = _make_deps(task=task_svc)
|
||||
c = Choreographer(deps)
|
||||
|
||||
env = await c.i_will_plan(
|
||||
pm_id,
|
||||
task_id,
|
||||
plan="decompose the code parent",
|
||||
rich_plan={
|
||||
"approach": _GOOD_APPROACH,
|
||||
"sub_tasks": [
|
||||
{"title": "API slice", "description": _GOOD_SUBTASK_DESC},
|
||||
{"title": "Test slice", "description": _GOOD_SUBTASK_DESC},
|
||||
],
|
||||
},
|
||||
)
|
||||
body = env.as_dict()
|
||||
# The gate must not reject merely for task_type='code' with sub_tasks.
|
||||
assert body.get("error") != "incomplete_input", body
|
||||
|
||||
Reference in New Issue
Block a user