feat(tasks): task-content guardrails — structured plans + constraints split (#328)

* feat(tasks): task-content guardrails — structured plans + constraints split

Bound task PLANNING content the way journals/notes already are, fixing the
poor task quality flagged 2026-07-07 (degenerate roots, over-decomposed
leaves, descriptions bloated by an auto-attached conventions dump).

Phase A — plan/AC guardrails (no migration):
- _pm_sub_tasks_gate: cap sub_tasks at 7; per-subtask ceilings (title <=200,
  description <=600) enforced at both the Pydantic boundary and the gate.
  Dropped the min-2-roots and no-subtasks-on-code rules: both contradict the
  2026-05-08 rule (test_cell_pm_can_plan_code_typed_parent_via_i_will_plan)
  and break legitimate single-cell roots. Long comment in the gate explains.
- IWillPlanRequest: plan <=2000, approach <=800 (floor 150 kept), typed
  SubTaskCreate/RiskCreate/OpenQuestionCreate replacing loose list[dict].
- DelegateRequest + task_completeness: acceptance_criteria capped at 7 items,
  each <=200 chars. New FieldRule.MAX_LENGTH_LIST + _post_rule_reject helper
  (extracted to keep the gate under xenon B).
- Routes dump typed models to dicts for the existing rich_plan shaper.

Phase B — conventions split (migration 068):
- New nullable tasks.constraints Text column; _attach_baseline_constraints
  now writes the ## Constraints block there instead of appending to
  description, so description is the human-authored instruction only. The
  conventions still reach the agent independently at spawn via the ambient
  block, so agent correctness is unaffected.
- TaskResponse / Task model / panel Task type carry constraints; panel shows
  a read-only Constraints card. Field is optional on the TS type (backend
  returns null for flag-off / pre-migration rows).

Tests: 5 new gate unit tests, 7 schema tests, 3 AC policy tests, 3 e2e smoke
scenarios; 4 baseline-constraints integration tests updated. ruff/mypy/xenon
clean; 10026 unit+foundation+e2e green; panel typecheck clean.

Refs: plan breezy-imagining-kahn

* test(tasks): use typed SubTaskCreate instead of dict literals in plan tests

make quality runs mypy over tests/ (1079 files), not just roboco/ — the
four sites passing dict literals to the now-typed sub_tasks: list[SubTaskCreate]
field failed mypy. Construct SubTaskCreate directly; the typed model raising
ValidationError IS the boundary the rejection tests assert.

* fix(deps): drop unused python-jose — clears PYSEC-2026-1325 (ecdsa, no fix)

CI's pip-audit went red on a freshly-published advisory PYSEC-2026-1325
against ecdsa 0.19.2 (no fix published — 0.19.2 is the latest). ecdsa is a
transitive dep of python-jose, which is a DIRECT dep of roboco but is NOT
imported anywhere in roboco/ or tests/ (grep-verified). The actual JWT path
uses PyJWT (import jwt) + fastapi_users.jwt, not python-jose.

So python-jose is a dead dependency. Removing it (deletion over an
--ignore-vuln waiver) drops ecdsa + rsa + pyasn1 + their type stubs from the
lockfile, eliminating the CVE at the source. deptry roboco/ stays clean
(no missing-dep), mypy clean, auth + schema tests pass.

Master CI was green 9h before this PR's run, so the advisory published in
that window would red any run including master — this fix unblocks both.

* chore(prompts): regenerate verb tables for typed plan sub_tasks

Phase A's IWillPlanRequest schema change (sub_tasks/risks/open_questions from
loose list[dict] to typed SubTaskCreate/RiskCreate/OpenQuestionCreate) made
the auto-generated verb tables stale. Regenerated via
scripts/regenerate_verb_tables.py — the diff is purely the signature
reflection (list[str|str] -> list[SubTaskCreate], etc.). Required by the
foundation-check gate (Makefile:559).

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
This commit is contained in:
Renzo F
2026-07-08 02:01:23 +02:00
committed by GitHub
co-authored by Renn F
parent 92ab13bce0
commit 2a9d9e25d9
23 changed files with 789 additions and 139 deletions
+1 -1
View File
@@ -12,7 +12,7 @@
| `escalate_up` | `escalate_up(task_id: UUID, reason: str)` | | `escalate_up` | `escalate_up(task_id: UUID, reason: str)` |
| `give_me_work` | `give_me_work()` | | `give_me_work` | `give_me_work()` |
| `i_am_idle` | `i_am_idle()` | | `i_am_idle` | `i_am_idle()` |
| `i_will_plan` | `i_will_plan(task_id: UUID, plan: str, approach: str, sub_tasks: list[str | str] = PydanticUndefined, technical_considerations: list[str] = PydanticUndefined, risks: list[str | str] = PydanticUndefined, open_questions: list[str | str | bool] = PydanticUndefined)` | | `i_will_plan` | `i_will_plan(task_id: UUID, plan: str, approach: str, sub_tasks: list[SubTaskCreate] = PydanticUndefined, technical_considerations: list[str] = PydanticUndefined, risks: list[RiskCreate] = PydanticUndefined, open_questions: list[OpenQuestionCreate] = PydanticUndefined)` |
| `reassign` | `reassign(task_id: UUID, new_assignee: str)` | | `reassign` | `reassign(task_id: UUID, new_assignee: str)` |
| `request_changes` | `request_changes(task_id: UUID, issues: list[str])` | | `request_changes` | `request_changes(task_id: UUID, issues: list[str])` |
| `resume` | `resume(task_id: UUID)` | | `resume` | `resume(task_id: UUID)` |
+1 -1
View File
@@ -13,7 +13,7 @@
| `escalate_up` | `escalate_up(task_id: UUID, reason: str)` | | `escalate_up` | `escalate_up(task_id: UUID, reason: str)` |
| `give_me_work` | `give_me_work()` | | `give_me_work` | `give_me_work()` |
| `i_am_idle` | `i_am_idle()` | | `i_am_idle` | `i_am_idle()` |
| `i_will_plan` | `i_will_plan(task_id: UUID, plan: str, approach: str, sub_tasks: list[str | str] = PydanticUndefined, technical_considerations: list[str] = PydanticUndefined, risks: list[str | str] = PydanticUndefined, open_questions: list[str | str | bool] = PydanticUndefined)` | | `i_will_plan` | `i_will_plan(task_id: UUID, plan: str, approach: str, sub_tasks: list[SubTaskCreate] = PydanticUndefined, technical_considerations: list[str] = PydanticUndefined, risks: list[RiskCreate] = PydanticUndefined, open_questions: list[OpenQuestionCreate] = PydanticUndefined)` |
| `request_changes` | `request_changes(task_id: UUID, issues: list[str])` | | `request_changes` | `request_changes(task_id: UUID, issues: list[str])` |
| `resume` | `resume(task_id: UUID)` | | `resume` | `resume(task_id: UUID)` |
| `submit_root` | `submit_root(task_id: UUID, notes: str)` | | `submit_root` | `submit_root(task_id: UUID, notes: str)` |
+2 -2
View File
@@ -115,7 +115,7 @@ real tools live in their agent_sdk drivers, not role_config.
| `escalate_up` | `escalate_up(task_id: UUID, reason: str)` | | `escalate_up` | `escalate_up(task_id: UUID, reason: str)` |
| `give_me_work` | `give_me_work()` | | `give_me_work` | `give_me_work()` |
| `i_am_idle` | `i_am_idle()` | | `i_am_idle` | `i_am_idle()` |
| `i_will_plan` | `i_will_plan(task_id: UUID, plan: str, approach: str, sub_tasks: list[str | str] = PydanticUndefined, technical_considerations: list[str] = PydanticUndefined, risks: list[str | str] = PydanticUndefined, open_questions: list[str | str | bool] = PydanticUndefined)` | | `i_will_plan` | `i_will_plan(task_id: UUID, plan: str, approach: str, sub_tasks: list[SubTaskCreate] = PydanticUndefined, technical_considerations: list[str] = PydanticUndefined, risks: list[RiskCreate] = PydanticUndefined, open_questions: list[OpenQuestionCreate] = PydanticUndefined)` |
| `reassign` | `reassign(task_id: UUID, new_assignee: str)` | | `reassign` | `reassign(task_id: UUID, new_assignee: str)` |
| `request_changes` | `request_changes(task_id: UUID, issues: list[str])` | | `request_changes` | `request_changes(task_id: UUID, issues: list[str])` |
| `resume` | `resume(task_id: UUID)` | | `resume` | `resume(task_id: UUID)` |
@@ -152,7 +152,7 @@ real tools live in their agent_sdk drivers, not role_config.
| `escalate_up` | `escalate_up(task_id: UUID, reason: str)` | | `escalate_up` | `escalate_up(task_id: UUID, reason: str)` |
| `give_me_work` | `give_me_work()` | | `give_me_work` | `give_me_work()` |
| `i_am_idle` | `i_am_idle()` | | `i_am_idle` | `i_am_idle()` |
| `i_will_plan` | `i_will_plan(task_id: UUID, plan: str, approach: str, sub_tasks: list[str | str] = PydanticUndefined, technical_considerations: list[str] = PydanticUndefined, risks: list[str | str] = PydanticUndefined, open_questions: list[str | str | bool] = PydanticUndefined)` | | `i_will_plan` | `i_will_plan(task_id: UUID, plan: str, approach: str, sub_tasks: list[SubTaskCreate] = PydanticUndefined, technical_considerations: list[str] = PydanticUndefined, risks: list[RiskCreate] = PydanticUndefined, open_questions: list[OpenQuestionCreate] = PydanticUndefined)` |
| `request_changes` | `request_changes(task_id: UUID, issues: list[str])` | | `request_changes` | `request_changes(task_id: UUID, issues: list[str])` |
| `resume` | `resume(task_id: UUID)` | | `resume` | `resume(task_id: UUID)` |
| `submit_root` | `submit_root(task_id: UUID, notes: str)` | | `submit_root` | `submit_root(task_id: UUID, notes: str)` |
@@ -0,0 +1,41 @@
"""Add tasks.constraints — move the conventions dump out of description.
The 2026-07-07 task-quality defect: ``TaskService._attach_baseline_constraints``
appended a ``## Constraints`` block to ``tasks.description``, so descriptions
ran 3000-4250 chars dominated by an auto-attached conventions dump. The
conventions already reach the agent independently at spawn via
``ConventionsService.render_ambient_block`` -> ``compose_prompt(ambient=...)``,
so the stored block is redundant bloat. Move it to a dedicated nullable
``constraints`` Text column so ``description`` is the human-authored
instruction only and the panel can render the constraints as a read-only card.
Additive: nullable column, default NULL. Existing rows keep their bloated
descriptions (a backfill that regex-strips markdown from a Text column is
riskier than the win — new tasks get clean descriptions from B2); no data
change in the migration.
Revision ID: 068_tasks_constraints_column
Revises: 067_respawn_revisit_resets
Create Date: 2026-07-07
"""
from __future__ import annotations
import sqlalchemy as sa
from alembic import op
revision = "068_tasks_constraints_column"
down_revision = "067_respawn_revisit_resets"
branch_labels = None
depends_on = None
def upgrade() -> None:
op.add_column(
"tasks",
sa.Column("constraints", sa.Text(), nullable=True),
)
def downgrade() -> None:
op.drop_column("tasks", "constraints")
@@ -78,6 +78,7 @@ export function TaskDescription({ task }: TaskDescriptionProps) {
}; };
return ( return (
<>
<Card> <Card>
<CardHeader className="pb-3"> <CardHeader className="pb-3">
<div className="flex items-center justify-between"> <div className="flex items-center justify-between">
@@ -177,5 +178,22 @@ export function TaskDescription({ task }: TaskDescriptionProps) {
)} )}
</CardContent> </CardContent>
</Card> </Card>
{task.constraints ? (
<Card className="border-dashed">
<CardHeader className="pb-3">
<CardTitle className="text-base text-muted-foreground">
Constraints
</CardTitle>
</CardHeader>
<CardContent>
<p className="text-xs text-muted-foreground mb-3">
Architectural standard derived from the project conventions
read-only. Applies to every task in this project.
</p>
<Markdown>{task.constraints}</Markdown>
</CardContent>
</Card>
) : null}
</>
); );
} }
+4
View File
@@ -216,6 +216,10 @@ export interface Task {
id: string; id: string;
title: string; title: string;
description: string; description: string;
// Server-derived architectural constraints (project baseline conventions),
// moved out of description. Null when conventions are flag-off / none /
// pre-migration. Read-only on the panel — system-derived, not human-edited.
constraints?: string | null;
acceptance_criteria: string[]; acceptance_criteria: string[];
status: TaskStatus; status: TaskStatus;
priority: number; // 0=P0(highest), 1=P1, 2=P2, 3=P3(lowest) priority: number; // 0=P0(highest), 1=P1, 2=P2, 3=P3(lowest)
+1 -4
View File
@@ -35,7 +35,6 @@ dependencies = [
# Utilities # Utilities
"httpx", "httpx",
"python-multipart", "python-multipart",
"python-jose[cryptography]", # JWT
"passlib[bcrypt]", # Password hashing "passlib[bcrypt]", # Password hashing
"tenacity", # Retry logic "tenacity", # Retry logic
"structlog", # Structured logging "structlog", # Structured logging
@@ -83,7 +82,6 @@ dev = [
# Type Stubs # Type Stubs
"types-passlib", "types-passlib",
"types-python-jose",
"types-PyYAML", "types-PyYAML",
# Development # Development
@@ -397,8 +395,7 @@ DEP002 = [
"python-multipart", "python-multipart",
# Database migrations (CLI tool) # Database migrations (CLI tool)
"alembic", "alembic",
# Auth libraries (used via passlib[bcrypt], python-jose[cryptography]) # Auth libraries (used via passlib[bcrypt]; JWT via PyJWT + fastapi_users.jwt)
"python-jose",
"passlib", "passlib",
# LLM utilities (embeddings/token counting) # LLM utilities (embeddings/token counting)
"openai", "openai",
+3 -3
View File
@@ -71,10 +71,10 @@ async def i_will_plan(
body.plan, body.plan,
rich_plan={ rich_plan={
"approach": body.approach, "approach": body.approach,
"sub_tasks": body.sub_tasks, "sub_tasks": [s.model_dump() for s in body.sub_tasks],
"technical_considerations": body.technical_considerations, "technical_considerations": body.technical_considerations,
"risks": body.risks, "risks": [r.model_dump() for r in body.risks],
"open_questions": body.open_questions, "open_questions": [q.model_dump() for q in body.open_questions],
}, },
) )
return envelope_to_response(env, request) return envelope_to_response(env, request)
+7 -3
View File
@@ -71,10 +71,14 @@ async def i_will_plan(
body.plan, body.plan,
rich_plan={ rich_plan={
"approach": body.approach, "approach": body.approach,
"sub_tasks": body.sub_tasks, # Typed models -> dicts: _build_panel_shaped_plan filters
# isinstance(st, dict), so a SubTaskCreate instance would be
# dropped. Dump to the dict shape every downstream reader
# (_plan_subtasks, _normalize_sub_task, the gate) already uses.
"sub_tasks": [s.model_dump() for s in body.sub_tasks],
"technical_considerations": body.technical_considerations, "technical_considerations": body.technical_considerations,
"risks": body.risks, "risks": [r.model_dump() for r in body.risks],
"open_questions": body.open_questions, "open_questions": [q.model_dump() for q in body.open_questions],
}, },
) )
return envelope_to_response(env, request) return envelope_to_response(env, request)
+8 -1
View File
@@ -200,7 +200,9 @@ class TaskUpdate(BaseModel):
# Basic info # Basic info
title: str | None = Field(default=None, min_length=1, max_length=200) title: str | None = Field(default=None, min_length=1, max_length=200)
description: str | None = Field(default=None, min_length=20) description: str | None = Field(default=None, min_length=20)
acceptance_criteria: list[str] | None = Field(default=None, min_length=1) acceptance_criteria: list[str] | None = Field(
default=None, min_length=1, max_length=7
)
priority: int | None = Field(default=None, ge=0, le=3) priority: int | None = Field(default=None, ge=0, le=3)
sequence: int | None = Field(default=None, ge=0) # Order within siblings sequence: int | None = Field(default=None, ge=0) # Order within siblings
target_date: datetime | None = None target_date: datetime | None = None
@@ -288,6 +290,10 @@ class TaskResponse(BaseModel):
id: UUID id: UUID
title: str title: str
description: str description: str
# Server-derived architectural constraints (project baseline conventions),
# moved out of `description` (migration 068). None when conventions are
# flag-off, the project has none, or the row predates the column.
constraints: str | None = None
acceptance_criteria: list[str] acceptance_criteria: list[str]
# Status # Status
@@ -773,6 +779,7 @@ def task_to_response(task: "TaskTable") -> TaskResponse:
id=require_uuid(task.id), id=require_uuid(task.id),
title=task.title, title=task.title,
description=task.description, description=task.description,
constraints=getattr(task, "constraints", None),
acceptance_criteria=task.acceptance_criteria or [], acceptance_criteria=task.acceptance_criteria or [],
status=task.status, status=task.status,
priority=task.priority, priority=task.priority,
+57 -7
View File
@@ -20,6 +20,38 @@ from roboco.models.base import Complexity
# technical_considerations, covers_parent_criteria). # technical_considerations, covers_parent_criteria).
StrList = Annotated[list[str], BeforeValidator(coerce_str_list)] StrList = Annotated[list[str], BeforeValidator(coerce_str_list)]
# AC discipline (the 2026-07-07 task-quality defect: restated, over-long
# acceptance criteria). Mirrors roboco.foundation.policy.task_completeness
# so an over-long / over-count AC list is rejected at the request boundary.
_AC_MAX_ITEMS = 7
_AC_MAX_ITEM_CHARS = 200
class SubTaskCreate(BaseModel):
"""A PM sub_task — a delegate target AND a progress-checklist item.
Mirrors DelegateRequest title/description caps so an over-long sub_task
can't bloat the plan (the 2026-07-07 task-quality defect). The server
assigns id + order; callers supply title + description only.
"""
title: str = Field(..., min_length=1, max_length=200)
description: str = Field(..., min_length=20, max_length=600)
class RiskCreate(BaseModel):
"""A {risk, mitigation} entry — what could go wrong and how it's handled."""
risk: str = Field(..., min_length=1, max_length=300)
mitigation: str = Field(..., min_length=1, max_length=600)
class OpenQuestionCreate(BaseModel):
"""An open question the PM wants answered before/during work."""
question: str = Field(..., min_length=1, max_length=300)
answered: bool = False
class GiveMeWorkRequest(BaseModel): class GiveMeWorkRequest(BaseModel):
"""Empty request body — agent_id comes from header.""" """Empty request body — agent_id comes from header."""
@@ -240,7 +272,7 @@ class EscalateToCeoRequest(BaseModel):
class IWillPlanRequest(BaseModel): class IWillPlanRequest(BaseModel):
task_id: UUID task_id: UUID
plan: str = Field(..., min_length=1) plan: str = Field(..., min_length=1, max_length=2000)
# Pre-gateway parity: Approach is REQUIRED — agents # Pre-gateway parity: Approach is REQUIRED — agents
# could not transition claimed → in_progress without filling this in the # could not transition claimed → in_progress without filling this in the
# pre-gateway flow. The Plan tab depends on it; smoke run 3 confirmed # pre-gateway flow. The Plan tab depends on it; smoke run 3 confirmed
@@ -248,14 +280,16 @@ class IWillPlanRequest(BaseModel):
# min_length must match choreographer._impl._PM_APPROACH_MIN_LEN. Raised # min_length must match choreographer._impl._PM_APPROACH_MIN_LEN. Raised
# 20→150: a 20-char approach was a one-liner; the approach + # 20→150: a 20-char approach was a one-liner; the approach +
# sub_tasks are also the progress checklist, so they must be substantive. # sub_tasks are also the progress checklist, so they must be substantive.
approach: str = Field(..., min_length=150) # max_length caps the bloat defect (approach that ran to thousands of
sub_tasks: list[dict[str, str]] = Field( # chars restating the description). Must match the gate ceiling.
approach: str = Field(..., min_length=150, max_length=800)
sub_tasks: list[SubTaskCreate] = Field(
default_factory=list, default_factory=list,
description="List of {title, description} — server assigns id + order", description="List of {title, description} — server assigns id + order",
) )
technical_considerations: StrList = Field(default_factory=list) technical_considerations: StrList = Field(default_factory=list)
risks: list[dict[str, str]] = Field(default_factory=list) risks: list[RiskCreate] = Field(default_factory=list)
open_questions: list[dict[str, str | bool]] = Field(default_factory=list) open_questions: list[OpenQuestionCreate] = Field(default_factory=list)
class DelegateRequest(BaseModel): class DelegateRequest(BaseModel):
@@ -283,8 +317,24 @@ class DelegateRequest(BaseModel):
nature: str = Field(..., min_length=1) nature: str = Field(..., min_length=1)
estimated_complexity: Complexity estimated_complexity: Complexity
# acceptance_criteria is required and non-empty; downstream policy # acceptance_criteria is required and non-empty; downstream policy
# also denylist-checks each item against placeholder phrases. # also denylist-checks each item against placeholder phrases. The
acceptance_criteria: StrList = Field(..., min_length=1) # per-item cap (<=200 chars) + list cap (<=7) mirror the policy so an
# over-long / over-count AC list is rejected at the boundary (422), not
# only by the service-layer completeness check.
acceptance_criteria: StrList = Field(..., min_length=1, max_length=7)
@field_validator("acceptance_criteria")
@classmethod
def _ac_items_bounded(cls, v: list[str]) -> list[str]:
for i, item in enumerate(v):
if len(item.strip()) > _AC_MAX_ITEM_CHARS:
raise ValueError(
f"acceptance_criteria[{i}] is {len(item.strip())} chars "
f"(max {_AC_MAX_ITEM_CHARS}) — a criterion that long is a "
"restated description, not a verifiable outcome. Split it."
)
return v
# Optional per-subtask project override. When omitted, the choreographer # Optional per-subtask project override. When omitted, the choreographer
# resolves the project from the parent's Product map for this cell, then # resolves the project from the parent's Product map for this cell, then
# falls back to the parent's project. Plain optional field — no validator. # falls back to the parent's project. Plain optional field — no validator.
+6
View File
@@ -160,6 +160,12 @@ class TaskTable(Base):
) )
title: Mapped[str] = mapped_column(String(200), nullable=False) title: Mapped[str] = mapped_column(String(200), nullable=False)
description: Mapped[str] = mapped_column(Text, nullable=False) description: Mapped[str] = mapped_column(Text, nullable=False)
# Server-derived architectural constraints (project baseline conventions
# block), moved out of `description` (migration 068). Nullable: flag-off /
# no-conventions / pre-migration rows have none. The conventions also
# reach the agent at spawn via the ambient block — this column is for
# panel visibility, not agent correctness.
constraints: Mapped[str | None] = mapped_column(Text, nullable=True)
acceptance_criteria: Mapped[list[str]] = mapped_column( acceptance_criteria: Mapped[list[str]] = mapped_column(
ARRAY(String), nullable=False ARRAY(String), nullable=False
) )
+58 -14
View File
@@ -29,6 +29,7 @@ class FieldRule(StrEnum):
MIN_LENGTH = "min_length" MIN_LENGTH = "min_length"
NON_EMPTY_LIST = "non_empty_list" NON_EMPTY_LIST = "non_empty_list"
EXPLICITLY_DECLARED = "explicitly_declared" EXPLICITLY_DECLARED = "explicitly_declared"
MAX_LENGTH_LIST = "max_length_list"
@dataclass(frozen=True) @dataclass(frozen=True)
@@ -100,8 +101,11 @@ _HINT_ACCEPTANCE_CRITERIA = (
"non-empty list[str]; each item describes a verifiable outcome (e.g. " "non-empty list[str]; each item describes a verifiable outcome (e.g. "
"'returns 401 when token absent'). Do NOT use placeholder strings like " "'returns 401 when token absent'). Do NOT use placeholder strings like "
"'completed and reviewed by assignee' — that's a known evasion phrase the " "'completed and reviewed by assignee' — that's a known evasion phrase the "
"gateway rejects." "gateway rejects. Keep each item <= 200 chars and the list <= 7 items — a "
"300-char AC or a 12-item list is a restated description, not a criterion."
) )
_AC_MAX_ITEMS = 7
_AC_MAX_ITEM_CHARS = 200
_HINT_TASK_TYPE = ( _HINT_TASK_TYPE = (
"one of: code | documentation | research | planning | design | administrative" "one of: code | documentation | research | planning | design | administrative"
) )
@@ -129,6 +133,12 @@ TASK_AT_CREATE: CompletenessSpec = CompletenessSpec(
FieldRule.NON_EMPTY_LIST, FieldRule.NON_EMPTY_LIST,
hint=_HINT_ACCEPTANCE_CRITERIA, hint=_HINT_ACCEPTANCE_CRITERIA,
), ),
FieldRequirement(
"acceptance_criteria",
FieldRule.MAX_LENGTH_LIST,
_AC_MAX_ITEMS,
hint=_HINT_ACCEPTANCE_CRITERIA,
),
FieldRequirement( FieldRequirement(
"task_type", FieldRule.EXPLICITLY_DECLARED, hint=_HINT_TASK_TYPE "task_type", FieldRule.EXPLICITLY_DECLARED, hint=_HINT_TASK_TYPE
), ),
@@ -196,6 +206,14 @@ def _check_non_empty_list(value: Any) -> tuple[bool, str | None]:
return True, None return True, None
def _check_max_length_list(value: Any, maximum: int) -> tuple[bool, str | None]:
if not isinstance(value, list):
return False, f"must be a list of length <= {maximum}"
if len(value) > maximum:
return False, f"must have at most {maximum} items (got {len(value)})"
return True, None
def _check_field(req: FieldRequirement, value: Any) -> tuple[bool, str | None]: def _check_field(req: FieldRequirement, value: Any) -> tuple[bool, str | None]:
"""Return (passed, problem_description). problem_description is None on pass.""" """Return (passed, problem_description). problem_description is None on pass."""
if req.rule is FieldRule.EXPLICITLY_DECLARED: if req.rule is FieldRule.EXPLICITLY_DECLARED:
@@ -204,6 +222,8 @@ def _check_field(req: FieldRequirement, value: Any) -> tuple[bool, str | None]:
return _check_non_empty_string(value) return _check_non_empty_string(value)
if req.rule is FieldRule.MIN_LENGTH: if req.rule is FieldRule.MIN_LENGTH:
return _check_min_length(value, req.value or 0) return _check_min_length(value, req.value or 0)
if req.rule is FieldRule.MAX_LENGTH_LIST:
return _check_max_length_list(value, req.value or 0)
return _check_non_empty_list(value) return _check_non_empty_list(value)
@@ -217,6 +237,38 @@ def _matches_denylist_ac(items: Any) -> bool:
) )
def _overlong_ac_item(items: Any) -> tuple[int, str] | None:
"""First (index, raw) AC item exceeding the per-item char cap, else None."""
if not isinstance(items, list):
return None
for i, item in enumerate(items):
if isinstance(item, str) and len(item.strip()) > _AC_MAX_ITEM_CHARS:
return i, item
return None
def _post_rule_reject(field: str, value: Any) -> str | None:
"""Content-gate a field that already passed its structural rule.
Returns a hint prefix when the value is rejected, or None to pass. The
caller appends `req.hint`, so this returns only the rejection reason.
"""
if field == "acceptance_criteria":
if _matches_denylist_ac(value):
return "rejected: placeholder phrase from the legacy silent fallback. "
over = _overlong_ac_item(value)
if over is not None:
idx, item = over
return (
f"rejected: criterion #{idx + 1} is {len(item.strip())} chars "
f"(max {_AC_MAX_ITEM_CHARS}) — a criterion that long is a "
"restated description, not a verifiable outcome. Split it. "
)
if field == "description" and _matches_denylist_description(value):
return "rejected: placeholder/empty phrase. "
return None
def _matches_denylist_description(text: Any) -> bool: def _matches_denylist_description(text: Any) -> bool:
"""True if the description matches any denylist regex.""" """True if the description matches any denylist regex."""
if not isinstance(text, str): if not isinstance(text, str):
@@ -252,19 +304,11 @@ def check(spec: CompletenessSpec, task: Any) -> CompletenessResult:
field_hints[req.field] = req.hint field_hints[req.field] = req.hint
continue continue
# Denylist checks (post-rule). # Post-rule content checks (denylists / per-item caps).
if req.field == "acceptance_criteria" and _matches_denylist_ac(value): hint = _post_rule_reject(req.field, value)
missing.append("acceptance_criteria") if hint is not None:
field_hints["acceptance_criteria"] = ( missing.append(req.field)
"rejected: placeholder phrase from the legacy silent fallback. " field_hints[req.field] = hint + req.hint
+ req.hint
)
continue
if req.field == "description" and _matches_denylist_description(value):
missing.append("description")
field_hints["description"] = (
"rejected: placeholder/empty phrase. " + req.hint
)
continue continue
return CompletenessResult( return CompletenessResult(
+8 -1
View File
@@ -141,8 +141,15 @@ class Task(TimestampMixin):
id: UUID = Field(default_factory=uuid4, description="Unique task identifier") id: UUID = Field(default_factory=uuid4, description="Unique task identifier")
title: str = Field(..., min_length=1, max_length=200, description="Task title") title: str = Field(..., min_length=1, max_length=200, description="Task title")
description: str = Field(..., description="Detailed task description") description: str = Field(..., description="Detailed task description")
# Server-derived architectural constraints (the project's baseline
# conventions block), moved out of `description` so the description is the
# human-authored instruction only (2026-07-07 task-quality fix). The
# conventions ALSO reach the agent at spawn via the ambient block, so
# this field is for panel visibility, not agent correctness. Nullable:
# flag-off / no-conventions / pre-migration rows have none.
constraints: str | None = Field(default=None)
acceptance_criteria: list[str] = Field( acceptance_criteria: list[str] = Field(
..., min_length=1, description="How do we know it's done?" ..., min_length=1, max_length=7, description="How do we know it's done?"
) )
acceptance_criteria_ids: list[str] = Field( acceptance_criteria_ids: list[str] = Field(
default_factory=list, default_factory=list,
+51 -5
View File
@@ -77,13 +77,26 @@ _PM_APPROACH_MIN_LEN = 150
# description that actually says what the step does. # description that actually says what the step does.
_PM_SUBTASK_DESC_MIN_LEN = 60 _PM_SUBTASK_DESC_MIN_LEN = 60
# Per-sub_task ceilings — mirror IWillPlanRequest.SubTaskCreate so direct
# service-layer callers (MCP, fixtures) can't bypass the Pydantic boundary
# with a 5000-char description (the bloat defect: descriptions dominated by
# over-long sub-task prose).
_PM_SUBTASK_TITLE_MAX_LEN = 200
_PM_SUBTASK_DESC_MAX_LEN = 600
# Over-decomposition cap (the 2026-07-07 task-quality defect): a plan with
# more than this many sub_tasks should be split into sibling coordination
# tasks, not one giant plan.
_PM_SUBTASKS_MAX = 7
def _thin_subtask_hint(sub_tasks: list[Any]) -> str | None: def _thin_subtask_hint(sub_tasks: list[Any]) -> str | None:
"""Return a hint if any PM sub_task is title-only / thin. """Return a hint if any PM sub_task is title-only / thin / over-long.
Each sub_task is a delegate target AND a progress-checklist item, so Each sub_task is a delegate target AND a progress-checklist item, so
a title with no real description is not a plan. Returns None when a title with no real description is not a plan. Also caps title +
every sub_task carries a title and a substantive description. description length so a sub_task can't bloat the plan. Returns None
when every sub_task carries a title and a substantive description.
""" """
for i, st in enumerate(sub_tasks): for i, st in enumerate(sub_tasks):
if not isinstance(st, dict): if not isinstance(st, dict):
@@ -92,12 +105,24 @@ def _thin_subtask_hint(sub_tasks: list[Any]) -> str | None:
desc = str(st.get("description") or "").strip() desc = str(st.get("description") or "").strip()
if not title: if not title:
return f"sub_task #{i + 1} has no title." return f"sub_task #{i + 1} has no title."
if len(title) > _PM_SUBTASK_TITLE_MAX_LEN:
return (
f"sub_task #{i + 1} ('{title[:40]}') title is too long "
f"({len(title)} chars) — keep titles <= "
f"{_PM_SUBTASK_TITLE_MAX_LEN} chars."
)
if len(desc) < _PM_SUBTASK_DESC_MIN_LEN: if len(desc) < _PM_SUBTASK_DESC_MIN_LEN:
return ( return (
f"sub_task #{i + 1} ('{title[:40]}') description is too thin " f"sub_task #{i + 1} ('{title[:40]}') description is too thin "
f"({len(desc)} chars) — need >= {_PM_SUBTASK_DESC_MIN_LEN} " f"({len(desc)} chars) — need >= {_PM_SUBTASK_DESC_MIN_LEN} "
"characters describing what the step actually does." "characters describing what the step actually does."
) )
if len(desc) > _PM_SUBTASK_DESC_MAX_LEN:
return (
f"sub_task #{i + 1} ('{title[:40]}') description is too long "
f"({len(desc)} chars) — keep descriptions <= "
f"{_PM_SUBTASK_DESC_MAX_LEN} chars; split if larger."
)
return None return None
@@ -540,6 +565,24 @@ class Choreographer:
elif thin := _thin_subtask_hint(sub_tasks): elif thin := _thin_subtask_hint(sub_tasks):
missing.append("sub_tasks") missing.append("sub_tasks")
field_hints["sub_tasks"] = thin field_hints["sub_tasks"] = thin
elif len(sub_tasks) > _PM_SUBTASKS_MAX:
# Over-decomposition cap (the 2026-07-07 task-quality defect):
# a plan with >7 sub_tasks is one giant plan that should be split
# into sibling coordination tasks. Per-item title/description
# bounds are in _thin_subtask_hint above. We deliberately do NOT
# ban sub_tasks on code tasks (PMs legitimately plan code-typed
# parents into dev subtasks — the 2026-05-08 rule,
# test_cell_pm_can_plan_code_typed_parent_via_i_will_plan) nor
# require >=2 on roots (a single-cell root → one cell task is a
# legitimate pass-through, test_pm_can_plan_non_code_parent) —
# both judgment calls the gate can't make without false positives.
missing.append("sub_tasks")
field_hints["sub_tasks"] = (
f"too many sub_tasks ({len(sub_tasks)}) — decompose into "
f"at most {_PM_SUBTASKS_MAX}. If the work is genuinely "
"larger, split it into sibling coordination tasks instead "
"of one giant plan."
)
if not missing: if not missing:
return None return None
return await self._emit_rejection( return await self._emit_rejection(
@@ -549,8 +592,11 @@ class Choreographer:
remediate=( remediate=(
"re-issue i_will_plan(task_id, plan, approach, " "re-issue i_will_plan(task_id, plan, approach, "
"sub_tasks=[{'title': '...', 'description': '...'}, ...]) " "sub_tasks=[{'title': '...', 'description': '...'}, ...]) "
f"with approach >= {_PM_APPROACH_MIN_LEN} chars and every " f"with approach >= {_PM_APPROACH_MIN_LEN} and <= 800 "
f"sub_task description >= {_PM_SUBTASK_DESC_MIN_LEN} chars." f"chars, every sub_task description >= "
f"{_PM_SUBTASK_DESC_MIN_LEN} and <= "
f"{_PM_SUBTASK_DESC_MAX_LEN} chars, and at most "
f"{_PM_SUBTASKS_MAX} sub_tasks."
), ),
context_briefing=briefing, context_briefing=briefing,
).with_introspection(task=task, role=role_str), ).with_introspection(task=task, role=role_str),
+13 -8
View File
@@ -1124,25 +1124,30 @@ class TaskService(BaseService):
return task return task
async def _attach_baseline_constraints(self, task: TaskTable) -> None: async def _attach_baseline_constraints(self, task: TaskTable) -> None:
"""Append the project's block-rule constraints to the task description. """Write the project's block-rule constraints to ``task.constraints``.
Server-derived backstop (flag-gated): every project task carries the Server-derived backstop (flag-gated): every project task carries the
hard conventions even if nothing upstream added them the layer the hard conventions even if nothing upstream added them the layer the
design calls "can't be skipped". Idempotent and non-suppressible: it design calls "can't be skipped". Idempotent: appends only baseline
appends only baseline constraints not already present (dedup by exact constraints not already present in the column (dedup by exact string),
string), so a second pass adds nothing AND an agent-authored so a second pass adds nothing AND an agent-authored ``## Constraints``
``## Constraints`` section can never suppress the mandatory baseline. section in the description can never suppress the mandatory baseline
Best-effort: a failure never blocks task creation. (the column is the source of truth for the panel; the description is
the human-authored instruction only 2026-07-07 task-quality fix).
The conventions ALSO reach the agent at spawn via the ambient block
(``ConventionsService.render_ambient_block``), so this column is for
panel visibility, not agent correctness. Best-effort: a failure never
blocks task creation.
""" """
baseline = await self._project_baseline_constraints(task) baseline = await self._project_baseline_constraints(task)
if not baseline: if not baseline:
return return
existing = task.description or "" existing = task.constraints or ""
missing = [item for item in baseline if item not in existing] missing = [item for item in baseline if item not in existing]
if not missing: if not missing:
return return
section = "## Constraints\n" + "\n".join(f"- {item}" for item in missing) section = "## Constraints\n" + "\n".join(f"- {item}" for item in missing)
task.description = f"{existing}\n\n{section}" if existing else section task.constraints = f"{existing}\n\n{section}" if existing else section
await self.session.flush() await self.session.flush()
async def _project_baseline_constraints(self, task: TaskTable) -> list[str]: async def _project_baseline_constraints(self, task: TaskTable) -> list[str]:
+161
View File
@@ -0,0 +1,161 @@
"""Scenario: the PM plan-content guardrails reject over-decomposed plans.
The 2026-07-07 task-quality defects (root 79d686f0 decomposed into 1 subtask,
code leaf 55376b8a carrying a 5-subtask plan, descriptions bloated to 3000+
chars) were structural the planning verb barely guardrailed its content.
The fix adds ceilings (plan <= 2000, approach <= 800, sub-task title <= 200,
sub-task description <= 600) and an over-decomposition cap (>7 sub_tasks
rejected) at both the HTTP Pydantic boundary and the choreographer gate.
This scenario drives a MAIN_PM planning root through the real API and asserts
each guardrail surfaces a clean ``incomplete_input`` envelope with a
remediation hint the agent can act on not a 500, not a silent accept. The
gate is the load-bearing layer (direct service callers bypass Pydantic), so
the assertions land on the envelope, not the HTTP status.
"""
from __future__ import annotations
from typing import TYPE_CHECKING, Any
from tests.e2e_smoke.arcs import (
origin_branch,
seed_company,
seed_project,
seed_task,
set_branch_name,
)
from tests.e2e_smoke.harness import ScriptedAgent, expect_error
if TYPE_CHECKING:
from uuid import UUID
from tests.e2e_smoke.arcs import Company
from tests.e2e_smoke.harness import E2EStack
# Pydantic's IWillPlanRequest.approach enforces 150..800 chars at the HTTP
# boundary, so every i_will_plan call needs a compliant approach.
_APPROACH = (
"Plan and delegate the page-scoped refresh button work to the frontend "
"cell: land the provider/hook, add the navbar button, remove the inline "
"buttons, and route one planning subtask to fe-pm for delivery. "
"Sequenced strictly; no cross-cell dependencies for this slice."
)
_GOOD_SUB = {
"title": "Frontend cell: refresh button",
"description": (
"Delegate the navbar refresh button to fe-pm: land the provider/hook "
"and wire the click handler into the page, then open the leaf PR."
),
}
_PLAN = "Land the refresh button via the frontend cell."
def _seed_planning_root(
stack: E2EStack, company: Company
) -> tuple[ScriptedAgent, UUID]:
"""Seed a PENDING MAIN_PM planning root + its origin branch."""
from roboco.models import Team
from roboco.models.base import TaskStatus, TaskType
project_id, _project_slug = seed_project(stack, company)
main_pm = ScriptedAgent(stack, company.main_pm_id, "main-pm", "main_pm")
task_id = seed_task(
stack,
title="Root: page-scoped refresh button",
description="Frontend-only root: provider/hook + navbar button.",
acceptance_criteria=["the refresh button lands on master"],
task_type=TaskType.PLANNING,
team=Team.MAIN_PM,
project_id=project_id,
created_by=company.main_pm_id,
assigned_to=company.main_pm_id,
status=TaskStatus.PENDING,
)
branch = f"feature/main_pm/{str(task_id)[:8]}"
origin_branch(stack, branch, start="master")
set_branch_name(stack, task_id, branch)
return main_pm, task_id
def _plan_with(sub_tasks: list[dict[str, Any]]) -> dict[str, Any]:
return {
"plan": _PLAN,
"approach": _APPROACH,
"sub_tasks": sub_tasks,
}
def test_pm_plan_over_decomposition_cap_rejected(e2e_stack: E2EStack) -> None:
"""A plan with >7 sub_tasks is over-decomposition — the gate rejects it
with incomplete_input + a 'split into sibling coordination tasks' hint."""
stack = e2e_stack
company = seed_company(stack)
main_pm, task_id = _seed_planning_root(stack, company)
env = main_pm.flow(
"i_will_plan",
task_id=str(task_id),
plan=_PLAN,
approach=_APPROACH,
sub_tasks=[dict(_GOOD_SUB, title=f"Slice {i}") for i in range(8)],
)
body = expect_error(env, "incomplete_input", "8 sub_tasks rejected")
assert "sub_tasks" in (body.get("missing") or []), body
assert "at most 7" in str(body.get("field_hints", {})), body
def test_pm_plan_overlong_subtask_description_rejected(
e2e_stack: E2EStack,
) -> None:
"""A sub-task description >600 chars is the bloat defect — rejected at
the Pydantic boundary (422), so the envelope carries the validation
``detail`` rather than the gate's ``field_hints``. Both layers are the
guardrail working; this test pins the boundary layer."""
stack = e2e_stack
company = seed_company(stack)
main_pm, task_id = _seed_planning_root(stack, company)
bloated = dict(_GOOD_SUB, description="x" * 700)
env = main_pm.flow(
"i_will_plan",
task_id=str(task_id),
plan=_PLAN,
approach=_APPROACH,
sub_tasks=[bloated],
)
body = expect_error(env, "incomplete_input", "over-long subtask desc")
# Boundary 422: missing is [] but detail carries the Pydantic error
# naming sub_tasks + the 600-char cap.
detail = str(body.get("detail"))
assert "sub_tasks" in detail, body
assert "600" in detail, body
def test_pm_plan_valid_plan_passes_gate(e2e_stack: E2EStack) -> None:
"""A well-formed plan (2 sub_tasks, bounded fields) passes the guardrails
and transitions the root to in_progress the happy path stays green."""
stack = e2e_stack
company = seed_company(stack)
main_pm, task_id = _seed_planning_root(stack, company)
env = main_pm.flow(
"i_will_plan",
task_id=str(task_id),
plan=_PLAN,
approach=_APPROACH,
sub_tasks=[
_GOOD_SUB,
{
"title": "Frontend cell: remove inline buttons",
"description": (
"Delegate removal of the stale inline refresh buttons to "
"fe-pm so the navbar button is the single source of truth."
),
},
],
)
# The gate must not fire; the root moves to in_progress. Downstream may
# raise a different error (e.g. tracing_gap) but NOT incomplete_input.
body = env
assert body.get("error") != "incomplete_input", body
@@ -152,3 +152,34 @@ def test_fill_parent_from_active_task_does_not_overwrite_explicit() -> None:
payload = {"parent_task_id": "explicit-id"} payload = {"parent_task_id": "explicit-id"}
result = tc.fill_parent_from_active_task(payload, "active-id") result = tc.fill_parent_from_active_task(payload, "active-id")
assert result["parent_task_id"] == "explicit-id" assert result["parent_task_id"] == "explicit-id"
# ---------------------------------------------------------------------------
# AC discipline (2026-07-07 task-quality defect): per-item cap + max count.
# ---------------------------------------------------------------------------
def test_task_at_create_rejects_too_many_acceptance_criteria() -> None:
"""An AC list >7 items is over-decomposition of criteria."""
result = tc.check(
tc.TASK_AT_CREATE, _task(acceptance_criteria=[f"c{i}" for i in range(8)])
)
assert result.passed is False
assert "acceptance_criteria" in result.missing
def test_task_at_create_accepts_seven_acceptance_criteria() -> None:
"""Exactly 7 AC items is the cap — must pass."""
result = tc.check(
tc.TASK_AT_CREATE, _task(acceptance_criteria=[f"c{i}" for i in range(7)])
)
assert result.passed is True
def test_task_at_create_rejects_overlong_acceptance_criterion() -> None:
"""An AC item >200 chars is a restated description, not a criterion."""
long_ac = "x" * 201
result = tc.check(tc.TASK_AT_CREATE, _task(acceptance_criteria=[long_ac]))
assert result.passed is False
assert "acceptance_criteria" in result.missing
assert any("200" in h for h in result.field_hints.values())
@@ -68,9 +68,12 @@ async def test_baseline_attached_when_flag_on(
monkeypatch.setattr(settings, "conventions_enabled", True) monkeypatch.setattr(settings, "conventions_enabled", True)
agent, project = await _seed(db_session) agent, project = await _seed(db_session)
task = await TaskService(db_session).create(_req(agent, project, "Do the work")) task = await TaskService(db_session).create(_req(agent, project, "Do the work"))
assert task.description is not None # The conventions block lives in `constraints`, not `description` (the
assert "## Constraints" in task.description # 2026-07-07 fix: description is the human-authored instruction only).
assert "no lint suppressions" in task.description assert task.description == "Do the work"
assert task.constraints is not None
assert "## Constraints" in task.constraints
assert "no lint suppressions" in task.constraints
async def test_flag_off_attaches_nothing( async def test_flag_off_attaches_nothing(
@@ -80,20 +83,24 @@ async def test_flag_off_attaches_nothing(
agent, project = await _seed(db_session) agent, project = await _seed(db_session)
task = await TaskService(db_session).create(_req(agent, project, "Do the work")) task = await TaskService(db_session).create(_req(agent, project, "Do the work"))
assert task.description == "Do the work" assert task.description == "Do the work"
assert task.constraints is None
async def test_baseline_not_suppressed_by_agent_constraints_section( async def test_baseline_not_suppressed_by_agent_constraints_section(
db_session: AsyncSession, monkeypatch: pytest.MonkeyPatch db_session: AsyncSession, monkeypatch: pytest.MonkeyPatch
) -> None: ) -> None:
# An agent-authored ## Constraints section must NOT suppress the mandatory # An agent-authored ## Constraints section in the description must NOT
# server baseline — both are present. # suppress the mandatory server baseline — the description keeps the
# agent's block and `constraints` carries the server baseline (separate
# field, so neither suppresses the other).
monkeypatch.setattr(settings, "conventions_enabled", True) monkeypatch.setattr(settings, "conventions_enabled", True)
agent, project = await _seed(db_session) agent, project = await _seed(db_session)
seeded = "Do the work\n\n## Constraints\n- a task-specific note" seeded = "Do the work\n\n## Constraints\n- a task-specific note"
task = await TaskService(db_session).create(_req(agent, project, seeded)) task = await TaskService(db_session).create(_req(agent, project, seeded))
assert task.description is not None assert task.description is not None
assert "a task-specific note" in task.description assert "a task-specific note" in task.description
assert "no lint suppressions" in task.description assert task.constraints is not None
assert "no lint suppressions" in task.constraints
async def test_baseline_attach_is_idempotent( async def test_baseline_attach_is_idempotent(
@@ -103,8 +110,8 @@ async def test_baseline_attach_is_idempotent(
agent, project = await _seed(db_session) agent, project = await _seed(db_session)
svc = TaskService(db_session) svc = TaskService(db_session)
task = await svc.create(_req(agent, project, "Do the work")) task = await svc.create(_req(agent, project, "Do the work"))
before = task.description before = task.constraints
await svc._attach_baseline_constraints(task) await svc._attach_baseline_constraints(task)
assert task.description == before assert task.constraints == before
assert task.description is not None assert task.constraints is not None
assert task.description.count("no lint suppressions") == 1 assert task.constraints.count("no lint suppressions") == 1
@@ -13,7 +13,7 @@ from fastapi import FastAPI
from fastapi.testclient import TestClient from fastapi.testclient import TestClient
from roboco.api.deps import get_choreographer from roboco.api.deps import get_choreographer
from roboco.api.routes.v1.flow_main_pm import router from roboco.api.routes.v1.flow_main_pm import router
from roboco.api.schemas.v1.flow import IWillPlanRequest from roboco.api.schemas.v1.flow import IWillPlanRequest, SubTaskCreate
_AGENT_ID = "00000000-0000-0000-0004-000000000001" _AGENT_ID = "00000000-0000-0000-0004-000000000001"
_HEADERS = {"X-Agent-ID": _AGENT_ID, "X-Agent-Role": "main_pm"} _HEADERS = {"X-Agent-ID": _AGENT_ID, "X-Agent-Role": "main_pm"}
@@ -111,7 +111,9 @@ def test_i_will_plan_schema_accepts_rich_plan() -> None:
task_id=uuid4(), task_id=uuid4(),
plan="Route to backend", plan="Route to backend",
approach=_GOOD_APPROACH, approach=_GOOD_APPROACH,
sub_tasks=[{"title": "Backend slice", "description": _GOOD_SUBTASK_DESC}], sub_tasks=[
SubTaskCreate(title="Backend slice", description=_GOOD_SUBTASK_DESC)
],
risks=[], risks=[],
open_questions=[], open_questions=[],
) )
+99
View File
@@ -11,6 +11,7 @@ from roboco.api.schemas.v1.flow import (
DelegateRequest, DelegateRequest,
IWillPlanRequest, IWillPlanRequest,
IWillWorkOnRequest, IWillWorkOnRequest,
SubTaskCreate,
) )
from roboco.models.base import Complexity from roboco.models.base import Complexity
@@ -146,3 +147,101 @@ def test_strlist_drops_non_string_junk_instead_of_crashing() -> None:
technical_considerations=technical_considerations, technical_considerations=technical_considerations,
) )
assert req.technical_considerations == ["real note"] assert req.technical_considerations == ["real note"]
# ---------------------------------------------------------------------------
# Task-content guardrails (2026-07-07): ceilings on plan content + AC caps.
# ---------------------------------------------------------------------------
def test_i_will_plan_request_rejects_overlong_plan() -> None:
"""plan >2000 chars is rejected at the boundary — the bloat defect."""
with pytest.raises(ValidationError) as exc:
IWillPlanRequest(
task_id=uuid4(),
plan="x" * 2001,
approach="a" * 150,
)
assert "plan" in str(exc.value)
def test_i_will_plan_request_rejects_overlong_approach() -> None:
"""approach >800 chars is rejected — no ceiling was the bloat bug."""
with pytest.raises(ValidationError) as exc:
IWillPlanRequest(
task_id=uuid4(),
plan="plan",
approach="a" * 801,
)
assert "approach" in str(exc.value)
def test_i_will_plan_request_rejects_overlong_subtask_title() -> None:
with pytest.raises(ValidationError):
IWillPlanRequest(
task_id=uuid4(),
plan="plan",
approach="a" * 150,
sub_tasks=[SubTaskCreate(title="t" * 201, description="d" * 30)],
)
def test_i_will_plan_request_rejects_overlong_subtask_description() -> None:
with pytest.raises(ValidationError):
IWillPlanRequest(
task_id=uuid4(),
plan="plan",
approach="a" * 150,
sub_tasks=[SubTaskCreate(title="ok title", description="d" * 601)],
)
def test_i_will_plan_request_rejects_thin_subtask_description() -> None:
"""A sub_task description <20 chars fails the typed SubTaskCreate model."""
with pytest.raises(ValidationError):
IWillPlanRequest(
task_id=uuid4(),
plan="plan",
approach="a" * 150,
sub_tasks=[SubTaskCreate(title="ok title", description="too short")],
)
def test_delegate_request_rejects_overlong_acceptance_criterion() -> None:
"""An AC item >200 chars is rejected — a criterion that long is a restated
description, not a verifiable outcome."""
long_ac = "x" * 201
with pytest.raises(ValidationError) as exc:
DelegateRequest.model_validate(
{
"parent_task_id": uuid4(),
"title": "t",
"description": "add the new endpoint plus tests",
"assigned_to": "be-dev-1",
"team": "backend",
"task_type": "code",
"nature": "technical",
"estimated_complexity": "medium",
"acceptance_criteria": [long_ac],
}
)
assert "acceptance_criteria" in str(exc.value)
def test_delegate_request_rejects_too_many_acceptance_criteria() -> None:
"""An AC list >7 items is rejected — over-decomposition of criteria."""
with pytest.raises(ValidationError) as exc:
DelegateRequest.model_validate(
{
"parent_task_id": uuid4(),
"title": "t",
"description": "add the new endpoint plus tests",
"assigned_to": "be-dev-1",
"team": "backend",
"task_type": "code",
"nature": "technical",
"estimated_complexity": "medium",
"acceptance_criteria": [f"criterion {i}" for i in range(8)],
}
)
assert "acceptance_criteria" in str(exc.value)
@@ -32,6 +32,8 @@ _GOOD_SUBTASK_DESC = (
"README H1 leaving the rest untouched, commits with the task-id prefix, " "README H1 leaving the rest untouched, commits with the task-id prefix, "
"and opens the leaf PR for QA." "and opens the leaf PR for QA."
) )
_LONG_SUBTASK_DESC = "x" * 700 # exceeds _PM_SUBTASK_DESC_MAX_LEN (600)
_OVERLONG_TITLE = "t" * 250 # exceeds _PM_SUBTASK_TITLE_MAX_LEN (200)
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
@@ -429,3 +431,199 @@ async def test_pm_reentry_in_progress_short_circuits_before_gate() -> None:
"The gate is firing before _handle_pm_reentry — ordering bug not fixed." "The gate is firing before _handle_pm_reentry — ordering bug not fixed."
) )
assert body.get("status") == "in_progress", body assert body.get("status") == "in_progress", body
# ---------------------------------------------------------------------------
# Test 7: over-decomposition cap — >7 sub_tasks rejected (2026-07-07 defect)
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_pm_with_too_many_subtasks_gets_incomplete_input() -> None:
"""A plan with >7 sub_tasks is over-decomposition — rejected with a hint
to split into sibling coordination tasks instead of one giant plan."""
pm_id = uuid4()
task_id = uuid4()
task_svc = _pm_task_svc(task_id, role="cell_pm")
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_plan(
pm_id,
task_id,
plan="decompose work",
rich_plan={
"approach": _GOOD_APPROACH,
"sub_tasks": [
{"title": f"Slice {i}", "description": _GOOD_SUBTASK_DESC}
for i in range(8) # 8 > _PM_SUBTASKS_MAX (7)
],
},
)
body = env.as_dict()
assert body["error"] == "incomplete_input", body
assert "sub_tasks" in (body.get("missing") or []), body
assert "at most 7" in str(body.get("field_hints", {})), body
@pytest.mark.asyncio
async def test_pm_with_seven_subtasks_passes_cap() -> None:
"""Exactly 7 sub_tasks is the cap — must not be rejected by the cap."""
pm_id = uuid4()
task_id = uuid4()
task_svc = _pm_task_svc(task_id, role="cell_pm")
claimed = MagicMock(
id=task_id,
status="claimed",
plan=None,
assigned_to=pm_id,
task_type="planning",
)
started = MagicMock(
id=task_id,
status="in_progress",
plan={"text": "x"},
assigned_to=pm_id,
task_type="planning",
)
task_svc.claim.return_value = claimed
task_svc.set_plan.return_value = claimed
task_svc.start.return_value = started
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_plan(
pm_id,
task_id,
plan="decompose work",
rich_plan={
"approach": _GOOD_APPROACH,
"sub_tasks": [
{"title": f"Slice {i}", "description": _GOOD_SUBTASK_DESC}
for i in range(7) # exactly the cap
],
},
)
body = env.as_dict()
assert body.get("error") != "incomplete_input", body
# ---------------------------------------------------------------------------
# Test 8: per-sub_task ceilings — over-long title / description rejected
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_pm_with_overlong_subtask_title_gets_incomplete_input() -> None:
"""A sub_task title >200 chars is rejected — titles are single-line labels."""
pm_id = uuid4()
task_id = uuid4()
task_svc = _pm_task_svc(task_id, role="cell_pm")
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_plan(
pm_id,
task_id,
plan="decompose work",
rich_plan={
"approach": _GOOD_APPROACH,
"sub_tasks": [
{"title": _OVERLONG_TITLE, "description": _GOOD_SUBTASK_DESC}
],
},
)
body = env.as_dict()
assert body["error"] == "incomplete_input", body
assert "sub_tasks" in (body.get("missing") or []), body
assert "title is too long" in str(body.get("field_hints", {})), body
@pytest.mark.asyncio
async def test_pm_with_overlong_subtask_description_gets_incomplete_input() -> None:
"""A sub_task description >600 chars is rejected — over-long sub-task
prose is the bloat defect (descriptions dominated by one giant step)."""
pm_id = uuid4()
task_id = uuid4()
task_svc = _pm_task_svc(task_id, role="cell_pm")
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_plan(
pm_id,
task_id,
plan="decompose work",
rich_plan={
"approach": _GOOD_APPROACH,
"sub_tasks": [
{"title": "Backend slice", "description": _LONG_SUBTASK_DESC}
],
},
)
body = env.as_dict()
assert body["error"] == "incomplete_input", body
assert "sub_tasks" in (body.get("missing") or []), body
assert "description is too long" in str(body.get("field_hints", {})), body
# ---------------------------------------------------------------------------
# Test 9: a code-typed parent WITH sub_tasks is NOT rejected — PMs legitimately
# plan code-typed parents into dev subtasks (the 2026-05-08 rule). The gate
# must not ban sub_tasks on code tasks.
# ---------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_pm_can_plan_code_typed_parent_with_subtasks() -> None:
"""A code-typed task planned by a PM with sub_tasks is legitimate
decomposition (test_cell_pm_can_plan_code_typed_parent_via_i_will_plan
establishes the 2026-05-08 rule). The cap/ceilings gate must NOT reject
merely because task_type is 'code'."""
pm_id = uuid4()
task_id = uuid4()
task_svc = AsyncMock()
task_svc.get.return_value = MagicMock(
id=task_id,
status="pending",
plan=None,
assigned_to=None,
task_type="code",
parent_task_id=None,
sequence=0,
team="backend",
commits=[],
pr_number=None,
branch_name=None,
quick_context=None,
)
task_svc.agent_for.return_value = MagicMock(
id=pm_id, role="cell_pm", team="backend", slug=None
)
task_svc.list_in_progress_for_agent.return_value = []
task_svc.list_paused_for_agent.return_value = []
task_svc.get_subtasks.return_value = []
task_svc.session = MagicMock()
task_svc.session.begin_nested = MagicMock(
return_value=MagicMock(
__aenter__=AsyncMock(return_value=None),
__aexit__=AsyncMock(return_value=False),
)
)
deps = _make_deps(task=task_svc)
c = Choreographer(deps)
env = await c.i_will_plan(
pm_id,
task_id,
plan="decompose the code parent",
rich_plan={
"approach": _GOOD_APPROACH,
"sub_tasks": [
{"title": "API slice", "description": _GOOD_SUBTASK_DESC},
{"title": "Test slice", "description": _GOOD_SUBTASK_DESC},
],
},
)
body = env.as_dict()
# The gate must not reject merely for task_type='code' with sub_tasks.
assert body.get("error") != "incomplete_input", body
Generated
-77
View File
@@ -764,18 +764,6 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/a7/5f/ed01f9a3cdffbd5a008556fc7b2a08ddb1cc6ace7effa7340604b1d16699/docstring_parser-0.18.0-py3-none-any.whl", hash = "sha256:b3fcbed555c47d8479be0796ef7e19c2670d428d72e96da63f3a40122860374b", size = 22484, upload-time = "2026-04-14T04:09:18.638Z" }, { url = "https://files.pythonhosted.org/packages/a7/5f/ed01f9a3cdffbd5a008556fc7b2a08ddb1cc6ace7effa7340604b1d16699/docstring_parser-0.18.0-py3-none-any.whl", hash = "sha256:b3fcbed555c47d8479be0796ef7e19c2670d428d72e96da63f3a40122860374b", size = 22484, upload-time = "2026-04-14T04:09:18.638Z" },
] ]
[[package]]
name = "ecdsa"
version = "0.19.2"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "six" },
]
sdist = { url = "https://files.pythonhosted.org/packages/25/ca/8de7744cb3bc966c85430ca2d0fcaeea872507c6a4cf6e007f7fe269ed9d/ecdsa-0.19.2.tar.gz", hash = "sha256:62635b0ac1ca2e027f82122b5b81cb706edc38cd91c63dda28e4f3455a2bf930", size = 202432, upload-time = "2026-03-26T09:58:17.675Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/51/79/119091c98e2bf49e24ed9f3ae69f816d715d2904aefa6a2baa039a2ba0b0/ecdsa-0.19.2-py2.py3-none-any.whl", hash = "sha256:840f5dc5e375c68f36c1a7a5b9caad28f95daa65185c9253c0c08dd952bb7399", size = 150818, upload-time = "2026-03-26T09:58:15.808Z" },
]
[[package]] [[package]]
name = "email-validator" name = "email-validator"
version = "2.3.0" version = "2.3.0"
@@ -2113,15 +2101,6 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/9b/bf/7595e817906a29453ba4d99394e781b6fabe55d21f3c15d240f85dd06bb1/py_serializable-2.1.0-py3-none-any.whl", hash = "sha256:b56d5d686b5a03ba4f4db5e769dc32336e142fc3bd4d68a8c25579ebb0a67304", size = 23045, upload-time = "2025-07-21T09:56:46.848Z" }, { url = "https://files.pythonhosted.org/packages/9b/bf/7595e817906a29453ba4d99394e781b6fabe55d21f3c15d240f85dd06bb1/py_serializable-2.1.0-py3-none-any.whl", hash = "sha256:b56d5d686b5a03ba4f4db5e769dc32336e142fc3bd4d68a8c25579ebb0a67304", size = 23045, upload-time = "2025-07-21T09:56:46.848Z" },
] ]
[[package]]
name = "pyasn1"
version = "0.6.3"
source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/5c/5f/6583902b6f79b399c9c40674ac384fd9cd77805f9e6205075f828ef11fb2/pyasn1-0.6.3.tar.gz", hash = "sha256:697a8ecd6d98891189184ca1fa05d1bb00e2f84b5977c481452050549c8a72cf", size = 148685, upload-time = "2026-03-17T01:06:53.382Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/5d/a0/7d793dce3fa811fe047d6ae2431c672364b462850c6235ae306c0efd025f/pyasn1-0.6.3-py3-none-any.whl", hash = "sha256:a80184d120f0864a52a073acc6fc642847d0be408e7c7252f31390c0f4eadcde", size = 83997, upload-time = "2026-03-17T01:06:52.036Z" },
]
[[package]] [[package]]
name = "pycparser" name = "pycparser"
version = "3.0" version = "3.0"
@@ -2342,25 +2321,6 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/0b/d7/1959b9648791274998a9c3526f6d0ec8fd2233e4d4acce81bbae76b44b2a/python_dotenv-1.2.2-py3-none-any.whl", hash = "sha256:1d8214789a24de455a8b8bd8ae6fe3c6b69a5e3d64aa8a8e5d68e694bbcb285a", size = 22101, upload-time = "2026-03-01T16:00:25.09Z" }, { url = "https://files.pythonhosted.org/packages/0b/d7/1959b9648791274998a9c3526f6d0ec8fd2233e4d4acce81bbae76b44b2a/python_dotenv-1.2.2-py3-none-any.whl", hash = "sha256:1d8214789a24de455a8b8bd8ae6fe3c6b69a5e3d64aa8a8e5d68e694bbcb285a", size = 22101, upload-time = "2026-03-01T16:00:25.09Z" },
] ]
[[package]]
name = "python-jose"
version = "3.5.0"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "ecdsa" },
{ name = "pyasn1" },
{ name = "rsa" },
]
sdist = { url = "https://files.pythonhosted.org/packages/c6/77/3a1c9039db7124eb039772b935f2244fbb73fc8ee65b9acf2375da1c07bf/python_jose-3.5.0.tar.gz", hash = "sha256:fb4eaa44dbeb1c26dcc69e4bd7ec54a1cb8dd64d3b4d81ef08d90ff453f2b01b", size = 92726, upload-time = "2025-05-28T17:31:54.288Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/d9/c3/0bd11992072e6a1c513b16500a5d07f91a24017c5909b02c72c62d7ad024/python_jose-3.5.0-py2.py3-none-any.whl", hash = "sha256:abd1202f23d34dfad2c3d28cb8617b90acf34132c7afd60abd0b0b7d3cb55771", size = 34624, upload-time = "2025-05-28T17:31:52.802Z" },
]
[package.optional-dependencies]
cryptography = [
{ name = "cryptography" },
]
[[package]] [[package]]
name = "python-multipart" name = "python-multipart"
version = "0.0.32" version = "0.0.32"
@@ -2600,7 +2560,6 @@ dependencies = [
{ name = "pydantic" }, { name = "pydantic" },
{ name = "pydantic-settings" }, { name = "pydantic-settings" },
{ name = "pyjwt" }, { name = "pyjwt" },
{ name = "python-jose", extra = ["cryptography"] },
{ name = "python-multipart" }, { name = "python-multipart" },
{ name = "python-toon" }, { name = "python-toon" },
{ name = "pyyaml" }, { name = "pyyaml" },
@@ -2637,7 +2596,6 @@ dev = [
{ name = "rich" }, { name = "rich" },
{ name = "ruff" }, { name = "ruff" },
{ name = "types-passlib" }, { name = "types-passlib" },
{ name = "types-python-jose" },
{ name = "types-pyyaml" }, { name = "types-pyyaml" },
{ name = "vulture" }, { name = "vulture" },
{ name = "xenon" }, { name = "xenon" },
@@ -2685,7 +2643,6 @@ requires-dist = [
{ name = "pytest-asyncio", marker = "extra == 'dev'" }, { name = "pytest-asyncio", marker = "extra == 'dev'" },
{ name = "pytest-cov", marker = "extra == 'dev'" }, { name = "pytest-cov", marker = "extra == 'dev'" },
{ name = "pytest-xdist", marker = "extra == 'dev'" }, { name = "pytest-xdist", marker = "extra == 'dev'" },
{ name = "python-jose", extras = ["cryptography"] },
{ name = "python-multipart" }, { name = "python-multipart" },
{ name = "python-toon" }, { name = "python-toon" },
{ name = "pyyaml" }, { name = "pyyaml" },
@@ -2703,7 +2660,6 @@ requires-dist = [
{ name = "tree-sitter-python" }, { name = "tree-sitter-python" },
{ name = "tree-sitter-typescript" }, { name = "tree-sitter-typescript" },
{ name = "types-passlib", marker = "extra == 'dev'" }, { name = "types-passlib", marker = "extra == 'dev'" },
{ name = "types-python-jose", marker = "extra == 'dev'" },
{ name = "types-pyyaml", marker = "extra == 'dev'" }, { name = "types-pyyaml", marker = "extra == 'dev'" },
{ name = "uvicorn", extras = ["standard"] }, { name = "uvicorn", extras = ["standard"] },
{ name = "vulture", marker = "extra == 'dev'" }, { name = "vulture", marker = "extra == 'dev'" },
@@ -2771,18 +2727,6 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/a6/b2/9574d4d44f7760c2aa32d92a0a4f41698e33f5b204a0bf5c9758f52c79d5/rpds_py-2026.6.3-cp314-cp314t-win_amd64.whl", hash = "sha256:c7b9a2f8f4d8e90af72571d3d495deebdd7e3c75451f5b41719aee166e940fc2", size = 220600, upload-time = "2026-06-30T07:16:43.091Z" }, { url = "https://files.pythonhosted.org/packages/a6/b2/9574d4d44f7760c2aa32d92a0a4f41698e33f5b204a0bf5c9758f52c79d5/rpds_py-2026.6.3-cp314-cp314t-win_amd64.whl", hash = "sha256:c7b9a2f8f4d8e90af72571d3d495deebdd7e3c75451f5b41719aee166e940fc2", size = 220600, upload-time = "2026-06-30T07:16:43.091Z" },
] ]
[[package]]
name = "rsa"
version = "4.9.1"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "pyasn1" },
]
sdist = { url = "https://files.pythonhosted.org/packages/da/8a/22b7beea3ee0d44b1916c0c1cb0ee3af23b700b6da9f04991899d0c555d4/rsa-4.9.1.tar.gz", hash = "sha256:e7bdbfdb5497da4c07dfd35530e1a902659db6ff241e39d9953cad06ebd0ae75", size = 29034, upload-time = "2025-04-16T09:51:18.218Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/64/8d/0133e4eb4beed9e425d9a98ed6e081a55d195481b7632472be1af08d2f6b/rsa-4.9.1-py3-none-any.whl", hash = "sha256:68635866661c6836b8d39430f97a996acbd61bfa49406748ea243539fe239762", size = 34696, upload-time = "2025-04-16T09:51:17.142Z" },
]
[[package]] [[package]]
name = "ruff" name = "ruff"
version = "0.15.20" version = "0.15.20"
@@ -3110,27 +3054,6 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/14/6a/e9fc6a5b8f9a380a4a56b9f1e4dba5c6899561868017b17f6de382808b6f/types_passlib-1.7.7.20260211-py3-none-any.whl", hash = "sha256:c0f1ad440c513a6c07f333b28249530686056fd54a7b3ac6128ae31fd46305d3", size = 40457, upload-time = "2026-02-10T15:11:58.647Z" }, { url = "https://files.pythonhosted.org/packages/14/6a/e9fc6a5b8f9a380a4a56b9f1e4dba5c6899561868017b17f6de382808b6f/types_passlib-1.7.7.20260211-py3-none-any.whl", hash = "sha256:c0f1ad440c513a6c07f333b28249530686056fd54a7b3ac6128ae31fd46305d3", size = 40457, upload-time = "2026-02-10T15:11:58.647Z" },
] ]
[[package]]
name = "types-pyasn1"
version = "0.6.0.20260408"
source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/58/c0/02f897fc8543f64fa6b1ca6a30d388e37c4ec2f761f469a2d9a29b89cdef/types_pyasn1-0.6.0.20260408.tar.gz", hash = "sha256:32dc90927adbe504fd2eee83ae30cf5ef934e5db0d1d94886071fed47eb50c8c", size = 17312, upload-time = "2026-04-08T04:27:16.874Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/e1/a5/473e06d5aaec3730aab5a9d40c2044e673c927412c24bd7f3fa0df7e95d3/types_pyasn1-0.6.0.20260408-py3-none-any.whl", hash = "sha256:ee7fbd98bce61193c5d4f8f7812fa53cddc5b8cc5ceb9fcda6eea539947c6d6b", size = 24044, upload-time = "2026-04-08T04:27:16.002Z" },
]
[[package]]
name = "types-python-jose"
version = "3.5.0.20260408"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "types-pyasn1" },
]
sdist = { url = "https://files.pythonhosted.org/packages/a1/b1/87cdb410d22913df39d7aad8864f94a4f5ad8d507ee07556888fdbe55e19/types_python_jose-3.5.0.20260408.tar.gz", hash = "sha256:3f8dccdc327bfffea7a81084ea1cea722fa499f13c1d04f7978b491dd36e0cf1", size = 11989, upload-time = "2026-04-08T04:34:10.577Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/d3/83/df2b34e64f0a674935d718471cf10fb392a7e5bdb0e9e7c739885b62d274/types_python_jose-3.5.0.20260408-py3-none-any.whl", hash = "sha256:968d8a8eac1ff9da249d6335a2bb9f82288d59ba23afe91fcc2662eb9f485e2a", size = 14694, upload-time = "2026-04-08T04:34:09.747Z" },
]
[[package]] [[package]]
name = "types-pyyaml" name = "types-pyyaml"
version = "6.0.12.20260518" version = "6.0.12.20260518"