Files
roboco/tests/foundation/test_task_completeness.py
T
2a9d9e25d9 feat(tasks): task-content guardrails — structured plans + constraints split (#328)
* feat(tasks): task-content guardrails — structured plans + constraints split

Bound task PLANNING content the way journals/notes already are, fixing the
poor task quality flagged 2026-07-07 (degenerate roots, over-decomposed
leaves, descriptions bloated by an auto-attached conventions dump).

Phase A — plan/AC guardrails (no migration):
- _pm_sub_tasks_gate: cap sub_tasks at 7; per-subtask ceilings (title <=200,
  description <=600) enforced at both the Pydantic boundary and the gate.
  Dropped the min-2-roots and no-subtasks-on-code rules: both contradict the
  2026-05-08 rule (test_cell_pm_can_plan_code_typed_parent_via_i_will_plan)
  and break legitimate single-cell roots. Long comment in the gate explains.
- IWillPlanRequest: plan <=2000, approach <=800 (floor 150 kept), typed
  SubTaskCreate/RiskCreate/OpenQuestionCreate replacing loose list[dict].
- DelegateRequest + task_completeness: acceptance_criteria capped at 7 items,
  each <=200 chars. New FieldRule.MAX_LENGTH_LIST + _post_rule_reject helper
  (extracted to keep the gate under xenon B).
- Routes dump typed models to dicts for the existing rich_plan shaper.

Phase B — conventions split (migration 068):
- New nullable tasks.constraints Text column; _attach_baseline_constraints
  now writes the ## Constraints block there instead of appending to
  description, so description is the human-authored instruction only. The
  conventions still reach the agent independently at spawn via the ambient
  block, so agent correctness is unaffected.
- TaskResponse / Task model / panel Task type carry constraints; panel shows
  a read-only Constraints card. Field is optional on the TS type (backend
  returns null for flag-off / pre-migration rows).

Tests: 5 new gate unit tests, 7 schema tests, 3 AC policy tests, 3 e2e smoke
scenarios; 4 baseline-constraints integration tests updated. ruff/mypy/xenon
clean; 10026 unit+foundation+e2e green; panel typecheck clean.

Refs: plan breezy-imagining-kahn

* test(tasks): use typed SubTaskCreate instead of dict literals in plan tests

make quality runs mypy over tests/ (1079 files), not just roboco/ — the
four sites passing dict literals to the now-typed sub_tasks: list[SubTaskCreate]
field failed mypy. Construct SubTaskCreate directly; the typed model raising
ValidationError IS the boundary the rejection tests assert.

* fix(deps): drop unused python-jose — clears PYSEC-2026-1325 (ecdsa, no fix)

CI's pip-audit went red on a freshly-published advisory PYSEC-2026-1325
against ecdsa 0.19.2 (no fix published — 0.19.2 is the latest). ecdsa is a
transitive dep of python-jose, which is a DIRECT dep of roboco but is NOT
imported anywhere in roboco/ or tests/ (grep-verified). The actual JWT path
uses PyJWT (import jwt) + fastapi_users.jwt, not python-jose.

So python-jose is a dead dependency. Removing it (deletion over an
--ignore-vuln waiver) drops ecdsa + rsa + pyasn1 + their type stubs from the
lockfile, eliminating the CVE at the source. deptry roboco/ stays clean
(no missing-dep), mypy clean, auth + schema tests pass.

Master CI was green 9h before this PR's run, so the advisory published in
that window would red any run including master — this fix unblocks both.

* chore(prompts): regenerate verb tables for typed plan sub_tasks

Phase A's IWillPlanRequest schema change (sub_tasks/risks/open_questions from
loose list[dict] to typed SubTaskCreate/RiskCreate/OpenQuestionCreate) made
the auto-generated verb tables stale. Regenerated via
scripts/regenerate_verb_tables.py — the diff is purely the signature
reflection (list[str|str] -> list[SubTaskCreate], etc.). Required by the
foundation-check gate (Makefile:559).

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-07-08 02:01:23 +02:00

186 lines
6.6 KiB
Python

"""Tier 1 — task-completeness rules."""
from __future__ import annotations
from types import SimpleNamespace
from typing import Any
from roboco.foundation.policy import task_completeness as tc
MIN_HINT_LEN = 20
PARENT_PRIORITY_HIGH = 4 # parent task priority used for inheritance assertions
DEFAULT_PRIORITY_MEDIUM = 2 # fill_priority_from_parent default when no parent
def _task(**fields: Any) -> SimpleNamespace:
"""Build a SimpleNamespace mimicking a Task with the given fields."""
defaults = {
"title": "ok",
"description": "x" * 30,
"acceptance_criteria": ["criterion one"],
"task_type": "code",
"nature": "technical",
"estimated_complexity": "medium",
"team": "backend",
}
defaults.update(fields)
return SimpleNamespace(**defaults)
def test_task_at_create_passes_for_complete_task() -> None:
result = tc.check(tc.TASK_AT_CREATE, _task())
assert result.passed is True
assert result.missing == []
def test_task_at_create_rejects_empty_acceptance_criteria() -> None:
result = tc.check(tc.TASK_AT_CREATE, _task(acceptance_criteria=[]))
assert result.passed is False
assert "acceptance_criteria" in result.missing
def test_task_at_create_rejects_short_description() -> None:
result = tc.check(tc.TASK_AT_CREATE, _task(description="too short"))
assert result.passed is False
assert "description" in result.missing
def test_task_at_create_rejects_missing_nature() -> None:
result = tc.check(tc.TASK_AT_CREATE, _task(nature=None))
assert result.passed is False
assert "nature" in result.missing
def test_denylist_rejects_silent_fallback_phrase() -> None:
"""Specifically catch the deleted task.py:5061 placeholder."""
bad = _task(acceptance_criteria=["completed and reviewed by assignee"])
result = tc.check(tc.TASK_AT_CREATE, bad)
assert result.passed is False
assert "acceptance_criteria" in result.missing
assert any("placeholder" in h.lower() for h in result.field_hints.values())
def test_denylist_rejects_see_title_description() -> None:
result = tc.check(tc.TASK_AT_CREATE, _task(description="see title"))
assert result.passed is False
assert "description" in result.missing
def test_denylist_rejects_todo_description() -> None:
result = tc.check(tc.TASK_AT_CREATE, _task(description="TODO"))
assert result.passed is False
def test_completeness_error_carries_missing_and_hints() -> None:
err = tc.TaskCompletenessError(
missing=["acceptance_criteria"],
field_hints={"acceptance_criteria": "non-empty list"},
)
assert err.missing == ["acceptance_criteria"]
assert err.field_hints["acceptance_criteria"] == "non-empty list"
def test_field_hints_for_missing_fields_are_actionable() -> None:
"""Each hint must mention the field and what valid input looks like."""
result = tc.check(
tc.TASK_AT_CREATE,
_task(
description="",
acceptance_criteria=[],
nature=None,
),
)
for field in ("description", "acceptance_criteria", "nature"):
assert field in result.field_hints, f"no hint for {field}"
hint = result.field_hints[field]
assert len(hint) >= MIN_HINT_LEN, (
f"hint for {field} too short to be useful: {hint!r}"
)
def test_fill_team_from_assignee_resolves_dev_slug() -> None:
payload = {"assigned_to": "be-dev-1"}
result = tc.fill_team_from_assignee(payload)
assert result["team"] == "backend"
assert result["assigned_to"] == "be-dev-1"
def test_fill_team_from_assignee_does_not_overwrite_explicit_team() -> None:
payload = {"assigned_to": "be-dev-1", "team": "frontend"}
result = tc.fill_team_from_assignee(payload)
# Auto-fill never silently overwrites; team stays as caller passed it.
assert result["team"] == "frontend"
def test_fill_team_from_assignee_unknown_slug_returns_unchanged() -> None:
payload = {"assigned_to": "notreal-1"}
result = tc.fill_team_from_assignee(payload)
# Auto-fill is best-effort; unknown slug = no fill, downstream rejects.
assert "team" not in result
def test_fill_priority_from_parent_inherits() -> None:
payload: dict[str, Any] = {}
parent = SimpleNamespace(priority=PARENT_PRIORITY_HIGH)
result = tc.fill_priority_from_parent(payload, parent)
assert result["priority"] == PARENT_PRIORITY_HIGH
assert result["__priority_inherited"] is True
def test_fill_priority_from_parent_does_not_overwrite_explicit() -> None:
payload = {"priority": 1}
parent = SimpleNamespace(priority=PARENT_PRIORITY_HIGH)
result = tc.fill_priority_from_parent(payload, parent)
assert result["priority"] == 1
assert "__priority_inherited" not in result
def test_fill_priority_from_parent_no_parent_uses_medium_default() -> None:
payload: dict[str, Any] = {}
result = tc.fill_priority_from_parent(payload, None)
assert result["priority"] == DEFAULT_PRIORITY_MEDIUM
assert result["__priority_inherited"] is True
def test_fill_parent_from_active_task_sets_id() -> None:
payload: dict[str, Any] = {}
result = tc.fill_parent_from_active_task(payload, "task-id-123")
assert result["parent_task_id"] == "task-id-123"
def test_fill_parent_from_active_task_does_not_overwrite_explicit() -> None:
payload = {"parent_task_id": "explicit-id"}
result = tc.fill_parent_from_active_task(payload, "active-id")
assert result["parent_task_id"] == "explicit-id"
# ---------------------------------------------------------------------------
# AC discipline (2026-07-07 task-quality defect): per-item cap + max count.
# ---------------------------------------------------------------------------
def test_task_at_create_rejects_too_many_acceptance_criteria() -> None:
"""An AC list >7 items is over-decomposition of criteria."""
result = tc.check(
tc.TASK_AT_CREATE, _task(acceptance_criteria=[f"c{i}" for i in range(8)])
)
assert result.passed is False
assert "acceptance_criteria" in result.missing
def test_task_at_create_accepts_seven_acceptance_criteria() -> None:
"""Exactly 7 AC items is the cap — must pass."""
result = tc.check(
tc.TASK_AT_CREATE, _task(acceptance_criteria=[f"c{i}" for i in range(7)])
)
assert result.passed is True
def test_task_at_create_rejects_overlong_acceptance_criterion() -> None:
"""An AC item >200 chars is a restated description, not a criterion."""
long_ac = "x" * 201
result = tc.check(tc.TASK_AT_CREATE, _task(acceptance_criteria=[long_ac]))
assert result.passed is False
assert "acceptance_criteria" in result.missing
assert any("200" in h for h in result.field_hints.values())