mirror of
https://github.com/rennf93/roboco.git
synced 2026-08-03 07:23:24 +02:00
* [943d8c4d] Frontend data freshness and approval-queue reliability audit (#631)
* [233a8b0f] WebSocket reconnect message-loss audit and fix (#625)
* [233a8b0f] fix(panel): add REST catch-up to useNotificationStream on WS reconnect
connection.ts has no message buffering/replay, so a notification published
while the CEO bell's socket was down (disconnected/reconnecting) was lost
forever instead of merely delayed. Add a reconnect-triggered GET
/notifications?unread_only=true catch-up folded into the existing
notification_id dedup so a notification delivered both via catch-up and
live WS is never double-counted, and make clearMessages drop the held
catch-up batch too. use-a2a-live.ts and use-rate-limit-websocket.ts were
audited and already have working reconnect-triggered REST fallbacks
(verified via a2a/page.tsx, rate-limit-banner.tsx, usage-overview-panel.tsx
and their existing F083 tests) so no fix was needed there.
* [233a8b0f] docs(panel): add comprehensive WebSocket hooks reference and reconnect architecture guide
Add panel/docs/frontend/hooks.md with full API reference for useWebSocket, useNotificationStream (with new REST catch-up behavior), useAgentStream, useA2ALiveStream, and useConnectionStatus. Include examples, best practices, and testing guidance.
Add panel/docs/architecture/websocket-reconnect.md documenting the message-loss mitigation pattern: Strategy 1 (REST catch-up for events, used by useNotificationStream) and Strategy 2 (REST invalidation for state, used by A2A/rate-limit consumers), plus the dedup logic ensuring no notification is double-counted on reconnect.
---------
Co-authored-by: Frontend Developer 1 <fe-dev-1@roboco.tech>
Co-authored-by: Frontend Documenter <fe-doc@roboco.tech>
* [d5315683] fix(frontend): add distinct toast feedback for silently-swallowed x-post and release-proposal statuses, plus regression tests for all 4 approval queues (#626)
Co-authored-by: Frontend Developer 2 <fe-dev-2@roboco.tech>
* [cd953838] Data-hook null-guard audit and API client 429 retry-by-method fix (#630)
* [cd953838] fix(panel): gate 429 retry by HTTP method, add hook null-guard regression tests
* [cd953838] chore(conventions): waive test-fixture wrapper in hooks null-guard test
* [cd953838] docs(frontend): document API rate-limit retry behavior and null-guard audit results
Added `docs/frontend/api-rate-limiting.md` to document the 429 retry strategy: GET/PUT auto-retry, POST/PATCH/DELETE require X-Idempotency-Key header. Updated `docs/frontend/hooks.md` to confirm the data-hook null-guard audit found all hooks already have correct `enabled` guards and include a regression test suite for the board-review poll on/off behavior and enabled-guard assertions.
---------
Co-authored-by: Frontend Developer 1 <fe-dev-1@roboco.tech>
Co-authored-by: Frontend Documenter <fe-doc@roboco.tech>
---------
Co-authored-by: Frontend Developer 1 <fe-dev-1@roboco.tech>
Co-authored-by: Frontend Documenter <fe-doc@roboco.tech>
Co-authored-by: Frontend Developer 2 <fe-dev-2@roboco.tech>
* [4534c71a] Backend concurrency, state-machine, and engine audit (#634)
* [41de844a] fix(lifecycle): sync CLAIM_RULES with runtime + clear stale claimant on PM hand-off (#627)
Two confirmed state-machine gaps found while auditing lifecycle.py,
task_lifecycle.py, the _ESCALATABLE_TO_BLOCKED bypass, and every
_REVIEW_QUEUE_STATES entry point:
- lifecycle.py's CLAIM_RULES/claim-ActionSpec/StatusTransition table
did not grant CELL_PM/MAIN_PM re-claim of AWAITING_PM_REVIEW even
though task.py's runtime _ROLE_CLAIM_STATUSES already granted it
and claimed the spec agreed -- the two tables had silently drifted,
breaking i_will_plan re-claim on an awaiting_pm_review task.
- docs_complete's _maybe_advance_to_pm_review pre-assigns a specific
owning PM via assigned_to but left claimed_by/active_claimant_id
pointing at the outgoing documenter, unlike every sibling transition
into a review-queue state. A stale active_claimant_id makes
content_actions.py's _active_claim_violation wrongly reject the
newly-assigned PM's own content writes before it formally claims.
Reassign claimed_by + active_claimant_id to the owning PM alongside
assigned_to.
Adds a regression test asserting the documenter's stale claim does not
survive the docs_complete -> awaiting_pm_review hand-off.
Co-authored-by: Backend Developer 2 <be-dev-2@roboco.tech>
* [0c46f666] Engine dedup race + sequencing.py edge-case audit (#628)
* [0c46f666] fix(sequencing): dedup race audit + collision-edge fallback bug
Audited the list-open-then-originate dedup pattern across six engines:
RoadmapEngine, XEngine.run_cycle, DepUpdateEngine, and CIWatchEngine each
run inside exactly one sequential orchestrator-loop asyncio task (no other
call site invokes run_cycle), so they cannot race with themselves; their
in-cycle dedup sets/keys are correctly built before any commit. SelfHealEngine
is the same shape. VideoEngine.open_video_task is genuinely different: it is
reachable from the release-publish hook, the feature-spotlight hook, and the
on-demand POST /video/request route, so two overlapping calls for the same
occasion can both pass the "no open task yet" check before either commits.
Fixed by wrapping the check+insert in a short-lived Redis mutex (reusing
HeartbeatMutex) keyed by occasion, mirroring XPostService's existing
lock pattern, with a regression test proving only one of two concurrent
calls creates a task.
Verified ReleaseExecutor's half-landed retry path (release_commit_sha):
apply_version_bumps and write_changelog_entry both run as uncommitted
working-tree edits before commit_and_push's single `git add -A` + commit,
so a bumped-version-without-changelog state can never reach origin (and
therefore can never be observed by a fresh retry clone) - confirmed correct
with a real-git-repo regression test, no fix needed.
Fixed sequencing.py's dev_task_collision_edges: the `if edges: return edges`
short-circuit dropped the same-assignee-lane fallback entirely whenever ANY
surfaced sibling pair produced a collision edge, even for a completely
unrelated same-assignee pair with no declared surface. Now the fallback
always runs, skipping only pairs the analyzer already ordered (so the two
mechanisms can never disagree on direction for the same pair).
Verified sequencing.py rule 3 (all-shared batch generates no edges): correct
by inspection (_shared_last_edges skips every pair when both are shared) and
confirmed with a regression test - no fix needed.
* [0c46f666] docs(reference): concurrency audit summary - engine races, fixes, verified patterns
---------
Co-authored-by: Backend Developer 2 <be-dev-2@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>
* [8f7f167a] Redis mutex pre-lock write audit (#629)
* [8f7f167a] Redis mutex pre-lock write audit: add cross-session regression test for XPostService.approve
Audited x_post_service.py, video_post_service.py, release_proposal.py, and
heartbeat_mutex.py for the pre-lock DB-write anti-pattern (a session write
that happens before the SET NX / HeartbeatMutex acquire returns a token,
letting a losing racer's stale write clobber a winner's committed state).
XPostService.approve, VideoPostService.approve, and
ReleaseProposalService.approve/reject already implement the correct
validate-pure-pre-lock, apply-under-lock pattern (the XPostService fix
already shipped per CHANGELOG.md: "X edited_body write deferred into the
single-flight lock (M5)"). HeartbeatMutex holds no AsyncSession at all, so
the anti-pattern is structurally inapplicable there.
Adds a genuine cross-session concurrency regression test to
test_x_post_service.py (a real second DB connection, not an in-process
mock) mirroring VideoPostService's existing cross-session test, proving a
concurrently-committed post survives and the CEO's edited body never lands
on the just-posted row.
* [8f7f167a] Remove redundant inline comments flagged by QA in cross-session regression test
Both comments restated what the surrounding docstrings already say
explicitly, per QA findings F-dbadd8f0 (line 294) and F-27ac051e (line
631) — no behavior change, tests re-verified green against a sandbox
Postgres.
* [8f7f167a] Remove inline trailing comments flagged by QA (correct file this time)
QA findings F-e6f3e6a6 and F-24189858 cited tests/unit/services/
test_x_post_service.py:294 and :631 across 5 revision rounds, but that
file never contained the flagged comment text — a repo-wide grep for
the exact quoted strings shows both comments actually live in the
mirrored tests/unit/services/test_video_post_service.py file, in its
own cross-session concurrency regression tests (the caption-edit and
tiktok-skip tests). Removed both there:
- "# externally visible to the "concurrent" session below" on the
db_session.commit() call
- "# never attempted without credentials" on the tiktok_poster.calls
assertion
Both restated what the surrounding docstrings/test names already say;
no behavior change. Verified with the full make quality gate against a
sandbox Postgres/Redis: 13,717 passed, 94.41% coverage, clean except
one pre-existing unrelated failure in tests/unit/api/test_cloud_auth.py
::test_login_route_parses_oauth2_form_not_query_params, which connects
to the app's default localhost:5432 Postgres (not the db_session
sandbox fixture) and is unreachable in this sandboxed environment —
structurally unrelated to the auth subsystem this task never touches.
* [8f7f167a] Redis mutex pre-lock write audit (round 7): add cross-session regression tests for reject() lock protection
Round-7 QA findings F-7eb9fbcb, F-06f39a2e, and F-4d56e49b claim
XPostService.reject(), ReleaseProposalService.reject(), and
release_executor._await_proc() lack lock protection / a CancelledError
handler — but their cited line ranges (255-267, 429-454, 241-257)
describe a pre-fix, shorter version of these functions that predates
commit fb293a787d, already on this branch. At current HEAD:
- x_post_service.py reject() (lines 275-299) acquires _LOCK_PREFIX,
re-reads under the lock, applies markers.set_x_reject_reason() +
CANCELLED only inside the critical section, releases in finally.
- release_proposal.py reject() (lines 460-486) does the identical
dance with _RELEASE_LOCK_PREFIX.
- release_executor.py _await_proc() (lines 257-265) already has an
`except asyncio.CancelledError` block that kills + reaps the child
and re-raises, mirroring the TimeoutError handler, with an existing
dedicated regression test
(test_await_proc_kills_child_on_outer_cancellation).
The one genuine gap: neither reject() path had a cross-session
(real second DB connection, not an in-process mock) regression test
proving the in-lock re-read catches a concurrent approve/publish that
completes mid-lock-wait — only approve() had one. Added
test_reject_concurrent_approve_completes_during_lock_wait to both
test_x_post_service.py and test_release_proposal_status_guards.py,
mirroring the existing approve() cross-session test: a second engine
commits COMPLETED between reject's pre-lock read and lock acquisition,
and the test asserts the CANCELLED write / reject-reason marker never
lands on the just-completed row.
No production code changed — verified via 103 targeted tests green
against a sandbox Postgres/Redis, plus `make -o sync gate` clean.
* [8f7f167a] Regenerate stale lifecycle artifacts (restore auditor waive_finding)
foundation-check was the only failing gate: the committed lifecycle artifacts
were missing the auditor's waive_finding verb that the lifecycle source
defines, so make quality regenerated them and failed on the diff — nothing to
do with the mutex fix (which passes ruff/mypy/tests/coverage/bandit clean).
make lifecycle restores the drift; this is what the 8 revision rounds kept
missing.
---------
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* [d615f2e3] fix(tests): sync stale CLAIM_RULES pinning assertions with lifecycle.py (#635)
test_claim_rules_match_pre_gateway_table still asserted the pre-audit
two-member frozenset for CELL_PM/MAIN_PM claim rules. CLAIM_RULES in
lifecycle.py already grants both roles claim rights on
Status.AWAITING_PM_REVIEW (added by the state-machine exhaustiveness
audit) so a PM can re-claim its own review-queue task after a respawn.
Updated both assertions to include AWAITING_PM_REVIEW, matching the
actual dict. Grepped the repo for sibling stale copies of the old
literal; found none beyond this test.
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
* [3c4e7a35] fix(quality-gate): reflow CONCURRENCY_AUDIT.md and stub occasion lock in video tests (#636)
Root cause: PR #634's CI failed at the markdown-reflow check (make quality
Makefile:285) on CONCURRENCY_AUDIT.md — a hard-wrapped audit doc left over
from the merged "Engine dedup race + sequencing.py edge-case audit" unit
(PR #628). Fixed with `make reflow-docs` (the exact remedy the CI output
itself named).
Running the full local `make quality` (with a sandbox Postgres/Redis to
get past DB-gated skips) surfaced a second real regression from that same
PR #628 unit: it added a Redis-backed HeartbeatMutex occasion lock to
VideoEngine.open_video_task, but two pre-existing test files
(tests/unit/runtime/test_video_render_loop.py and
tests/integration/test_video_routes.py) call open_video_task without
stubbing that lock, so they failed closed against the suite's
deliberately-unreachable test Redis (_no_live_redis). Fixed by applying
the same lock-stub pattern tests/unit/services/test_video_engine.py
already uses for its own occasion-lock tests: an autouse HeartbeatMutex
stand-in fixture in test_video_render_loop.py, and wrapping the two
route-level video-request tests in test_video_routes.py with the file's
existing _LOCKED patch pair (already used by every other lock-dependent
test in that file).
The one remaining local failure,
test_cloud_auth.py::test_login_route_parses_oauth2_form_not_query_params,
is a pre-existing environment gap unrelated to this branch: it needs a
real Postgres reachable at localhost:5432 (which .github/workflows/ci.yml
provides as a service container) but this dev sandbox has no such binding
— confirmed unrelated to any of the four merged audit units.
make quality now passes clean: 13729 passed, 0 regressions, 94.49% coverage.
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
* [ddc8121f] regenerate lifecycle artifacts for awaiting_pm_review claim rules and waive_finding intent (#637)
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
---------
Co-authored-by: Backend Developer 2 <be-dev-2@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* fix(sequencing): drop lane-fallback edges that would cycle against analyzer edges
The dev-task collision fallback unioned the analyzer's authoritative edges
with same-assignee lane edges, deduping only the direct pair. A lane chain
through an unsurfaced middle sibling could still contradict an analyzer edge
transitively (the shared-last migration order inverts plain priority order),
closing a 3-cycle that made add_dependency raise ConflictError and wedged
every later delegate to that parent. Fallback edges are now accepted only
when they can't close a cycle against the edges already kept; a regression
test reproduces the exact scenario.
Also strip pre-merge cruft: remove the root CONCURRENCY_AUDIT.md working
report, delete the near-duplicate websocket-reconnect.md doc, fix the stale
a2a/page.tsx doc citation, correct the api-rate-limiting doc to state
idempotency-key retry is unimplemented, and fix two lifecycle.py comments
that referenced a guard function which never existed.
---------
Co-authored-by: Frontend Developer 1 <fe-dev-1@roboco.tech>
Co-authored-by: Frontend Documenter <fe-doc@roboco.tech>
Co-authored-by: Frontend Developer 2 <fe-dev-2@roboco.tech>
Co-authored-by: Backend Developer 2 <be-dev-2@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
1151 lines
42 KiB
Python
1151 lines
42 KiB
Python
"""Tier 1 — spec self-tests. Fast (no DB, no network)."""
|
|
|
|
from __future__ import annotations
|
|
|
|
from types import SimpleNamespace
|
|
from typing import Any, cast
|
|
from uuid import uuid4
|
|
|
|
import pytest
|
|
from roboco.foundation import _validate_lifecycle as _validate
|
|
from roboco.foundation._validate_lifecycle import reachable_from
|
|
from roboco.foundation.policy import lifecycle as spec
|
|
from roboco.foundation.policy.lifecycle import _INTENT_VERBS, IntentSpec
|
|
from roboco.models.base import TaskStatus as ModelTaskStatus
|
|
from roboco.models.base import TaskType as ModelTaskType
|
|
|
|
|
|
def test_role_enum_has_every_pre_gateway_role() -> None:
|
|
"""Every role from PERMISSIONS.md must be enumerated.
|
|
|
|
The canonical Role enum is now defined in `roboco.foundation.identity`
|
|
and re-exported here. It includes the 9 pre-gateway roles plus the
|
|
SYSTEM sentinel used for orchestrator-generated rows. The pre-gateway
|
|
PERMISSIONS.md is the historical canon — SYSTEM is the post-foundation
|
|
addition that doesn't appear in policy tables.
|
|
"""
|
|
expected = {
|
|
"developer",
|
|
"qa",
|
|
"documenter",
|
|
"cell_pm",
|
|
"main_pm",
|
|
"product_owner",
|
|
"head_marketing",
|
|
"auditor",
|
|
"pr_reviewer", # reviews inbound external/fork PRs (read-only)
|
|
"prompter", # post-gateway intake role (human-only, drafts tasks)
|
|
"secretary", # CEO's chief-of-staff (human-only, gated CEO authority)
|
|
"ceo",
|
|
"system",
|
|
}
|
|
actual = {r.value for r in spec.Role}
|
|
assert actual == expected, f"Role enum drift: {actual ^ expected}"
|
|
|
|
|
|
def test_status_enum_has_every_pre_gateway_status() -> None:
|
|
"""Every status from STATUS_TRANSITIONS.md must be enumerated."""
|
|
expected = {
|
|
"backlog",
|
|
"pending",
|
|
"claimed",
|
|
"in_progress",
|
|
"blocked",
|
|
"paused",
|
|
"verifying",
|
|
"awaiting_qa",
|
|
"needs_revision",
|
|
"awaiting_documentation",
|
|
"awaiting_pr_review",
|
|
"awaiting_pm_review",
|
|
"awaiting_ceo_approval",
|
|
"completed",
|
|
"cancelled",
|
|
}
|
|
actual = {s.value for s in spec.Status}
|
|
assert actual == expected, f"Status enum drift: {actual ^ expected}"
|
|
|
|
|
|
def test_task_type_enum_matches_models() -> None:
|
|
"""The spec's TaskType must match the existing models.base.TaskType.
|
|
|
|
If the existing model adds/removes a type, the spec must be updated
|
|
in lockstep — that's the entire point of this module.
|
|
"""
|
|
spec_values = {t.value for t in spec.TaskType}
|
|
model_values = {t.value for t in ModelTaskType}
|
|
assert spec_values == model_values, (
|
|
f"TaskType drift between lifecycle.spec and models.base: "
|
|
f"{spec_values ^ model_values}"
|
|
)
|
|
|
|
|
|
def test_status_enum_matches_models() -> None:
|
|
"""The spec's Status must match models.base.TaskStatus — the ORM column
|
|
type and the lifecycle map must not drift. TaskType has this guard; Status
|
|
did not, so adding/renaming a status in one enum only wedged silently."""
|
|
spec_values = {s.value for s in spec.Status}
|
|
model_values = {s.value for s in ModelTaskStatus}
|
|
assert spec_values == model_values, (
|
|
f"Status drift between lifecycle.spec and models.base: "
|
|
f"{spec_values ^ model_values}"
|
|
)
|
|
|
|
|
|
def test_status_enum_parity_validator_passes_on_real_spec() -> None:
|
|
"""The import-time parity validator must agree with the real enums."""
|
|
_validate._check_status_enum_parity() # no raise
|
|
|
|
|
|
def test_status_coverage_rejects_stray_string_target() -> None:
|
|
"""A transition referencing a non-Status target string must fail the
|
|
coverage validator — the old check was a tautology (STATUS_GRAPH keys every
|
|
Status by construction) and let stray-string targets through."""
|
|
fake = (
|
|
SimpleNamespace(
|
|
source=spec.Status.PENDING, target="bogus_state", triggered_by_action="x"
|
|
),
|
|
)
|
|
original = spec._STATUS_TRANSITIONS
|
|
spec._STATUS_TRANSITIONS = cast("Any", fake)
|
|
try:
|
|
with pytest.raises(_validate.LifecycleSpecError, match="non-Status"):
|
|
_validate._check_status_enum_coverage()
|
|
finally:
|
|
spec._STATUS_TRANSITIONS = original
|
|
|
|
|
|
def test_status_coverage_rejects_orphan_non_terminal_source() -> None:
|
|
"""A non-terminal status that is the source of no transition (an orphan
|
|
state) must fail the coverage validator — the cancel fan-out made the old
|
|
'is a key in STATUS_GRAPH' check structurally always-true."""
|
|
original = spec._STATUS_TRANSITIONS
|
|
spec._STATUS_TRANSITIONS = tuple(
|
|
t for t in original if t.source is not spec.Status.PAUSED
|
|
)
|
|
try:
|
|
with pytest.raises(
|
|
_validate.LifecycleSpecError, match="no outgoing transition"
|
|
):
|
|
_validate._check_status_enum_coverage()
|
|
finally:
|
|
spec._STATUS_TRANSITIONS = original
|
|
|
|
|
|
def test_terminal_exit_requires_a_completed_path() -> None:
|
|
"""Every non-terminal status must reach COMPLETED specifically — the cancel
|
|
fan-out made the old {COMPLETED, CANCELLED} check trivial, so a status whose
|
|
sole exit was cancel passed the guard with no real forward completion path."""
|
|
original = spec.STATUS_GRAPH
|
|
fake = dict(original)
|
|
fake[spec.Status.PAUSED] = frozenset({spec.Status.CANCELLED})
|
|
spec.STATUS_GRAPH = fake
|
|
try:
|
|
with pytest.raises(_validate.LifecycleSpecError, match="no path to COMPLETED"):
|
|
_validate._check_terminal_exits()
|
|
finally:
|
|
spec.STATUS_GRAPH = original
|
|
|
|
|
|
def test_decision_allow_has_no_rejection_kind() -> None:
|
|
d = spec.Decision.allow()
|
|
assert d.allowed is True
|
|
assert d.rejection_kind is None
|
|
assert d.message is None
|
|
assert d.missing == []
|
|
assert d.remediate is None
|
|
|
|
|
|
def test_decision_reject_requires_rejection_kind() -> None:
|
|
d = spec.Decision.reject(
|
|
kind="not_authorized",
|
|
message="role 'developer' may not call delegate",
|
|
remediate="only PMs delegate; call give_me_work() instead",
|
|
)
|
|
assert d.allowed is False
|
|
assert d.rejection_kind == "not_authorized"
|
|
assert d.message == "role 'developer' may not call delegate"
|
|
assert d.remediate == "only PMs delegate; call give_me_work() instead"
|
|
|
|
|
|
def test_decision_tracing_gap_carries_missing_list() -> None:
|
|
d = spec.Decision.tracing_gap(
|
|
missing=["plan", "journal:decision"],
|
|
remediate="provide plan and a journal:decision entry",
|
|
)
|
|
assert d.allowed is False
|
|
assert d.rejection_kind == "tracing_gap"
|
|
assert d.missing == ["plan", "journal:decision"]
|
|
assert d.remediate == "provide plan and a journal:decision entry"
|
|
|
|
|
|
def test_decision_tracing_gap_defensively_copies_missing() -> None:
|
|
"""tracing_gap must isolate the stored list from the caller's source."""
|
|
src = ["plan"]
|
|
d = spec.Decision.tracing_gap(missing=src, remediate="r")
|
|
src.append("mutated")
|
|
assert d.missing == ["plan"]
|
|
|
|
|
|
def test_decision_invariants_enforced_at_construction() -> None:
|
|
"""allowed=True ⇒ rejection_kind None; allowed=False ⇒ kind set."""
|
|
with pytest.raises(ValueError, match="allowed=True requires rejection_kind=None"):
|
|
spec.Decision(
|
|
allowed=True,
|
|
rejection_kind="not_authorized",
|
|
message="x",
|
|
missing=[],
|
|
remediate="x",
|
|
)
|
|
with pytest.raises(ValueError, match="allowed=False requires rejection_kind"):
|
|
spec.Decision(
|
|
allowed=False,
|
|
rejection_kind=None,
|
|
message="x",
|
|
missing=[],
|
|
remediate="x",
|
|
)
|
|
|
|
|
|
def test_decision_invariant_rejects_allowed_with_missing_or_remediate() -> None:
|
|
"""allowed=True with missing or remediate set raises (Fix 1 lock-in)."""
|
|
with pytest.raises(
|
|
ValueError, match="allowed=True requires missing=\\[\\] and remediate=None"
|
|
):
|
|
spec.Decision(
|
|
allowed=True,
|
|
rejection_kind=None,
|
|
message=None,
|
|
missing=["plan"],
|
|
remediate=None,
|
|
)
|
|
with pytest.raises(
|
|
ValueError,
|
|
match="allowed=True requires missing=\\[\\] and remediate=None",
|
|
):
|
|
spec.Decision(
|
|
allowed=True,
|
|
rejection_kind=None,
|
|
message=None,
|
|
missing=[],
|
|
remediate="oops",
|
|
)
|
|
|
|
|
|
def test_precondition_check_returns_bool() -> None:
|
|
"""A Precondition.check() is the gate-table evaluator."""
|
|
p = spec.Precondition(
|
|
key="commits>=1",
|
|
check=lambda task, _agent, _ctx: bool(getattr(task, "commits", None)),
|
|
remediate="commit at least once before opening a PR",
|
|
missing_token="commits>=1",
|
|
)
|
|
|
|
task_with = SimpleNamespace(commits=["abc"])
|
|
task_without = SimpleNamespace(commits=[])
|
|
assert p.check(task_with, None, None) is True
|
|
assert p.check(task_without, None, None) is False
|
|
|
|
|
|
def test_action_spec_holds_role_status_and_precondition_data() -> None:
|
|
a = spec.ActionSpec(
|
|
name="claim",
|
|
allowed_roles=frozenset({spec.Role.DEVELOPER}),
|
|
source_statuses=frozenset({spec.Status.PENDING, spec.Status.NEEDS_REVISION}),
|
|
target_status=spec.Status.CLAIMED,
|
|
allowed_task_types=None,
|
|
preconditions=(),
|
|
self_review_block=False,
|
|
needs_team_match=True,
|
|
)
|
|
assert a.name == "claim"
|
|
assert spec.Role.DEVELOPER in a.allowed_roles
|
|
assert a.target_status == spec.Status.CLAIMED
|
|
|
|
|
|
def test_intent_spec_composes_atomic_actions() -> None:
|
|
i = spec.IntentSpec(
|
|
name="i_will_work_on",
|
|
allowed_roles=frozenset({spec.Role.DEVELOPER}),
|
|
description="Claim a task and start work on it.",
|
|
composes=("claim", "set_plan", "start"),
|
|
extra_preconditions=(),
|
|
side_effects=(),
|
|
next_hint=lambda _t: "edit + commit, then open_pr",
|
|
)
|
|
assert i.composes == ("claim", "set_plan", "start")
|
|
assert i.next_hint(None) == "edit + commit, then open_pr"
|
|
|
|
|
|
def test_status_transition_carries_role_constraint_optional() -> None:
|
|
t = spec.StatusTransition(
|
|
source=spec.Status.AWAITING_QA,
|
|
target=spec.Status.AWAITING_DOCUMENTATION,
|
|
triggered_by_action="qa_pass",
|
|
role_constraint=frozenset({spec.Role.QA}),
|
|
)
|
|
assert t.source == spec.Status.AWAITING_QA
|
|
assert t.target == spec.Status.AWAITING_DOCUMENTATION
|
|
assert t.triggered_by_action == "qa_pass"
|
|
assert t.role_constraint == frozenset({spec.Role.QA})
|
|
|
|
|
|
def test_status_transitions_includes_dev_path() -> None:
|
|
"""The dev happy path: pending → claimed → in_progress → verifying → awaiting_qa."""
|
|
sources = {(t.source, t.target) for t in spec._STATUS_TRANSITIONS}
|
|
assert (spec.Status.PENDING, spec.Status.CLAIMED) in sources
|
|
assert (spec.Status.CLAIMED, spec.Status.IN_PROGRESS) in sources
|
|
assert (spec.Status.IN_PROGRESS, spec.Status.VERIFYING) in sources
|
|
assert (spec.Status.VERIFYING, spec.Status.AWAITING_QA) in sources
|
|
|
|
|
|
def test_status_transitions_includes_qa_paths() -> None:
|
|
sources = {(t.source, t.target) for t in spec._STATUS_TRANSITIONS}
|
|
assert (spec.Status.AWAITING_QA, spec.Status.CLAIMED) in sources # QA claims
|
|
assert (spec.Status.AWAITING_QA, spec.Status.AWAITING_DOCUMENTATION) in sources
|
|
assert (spec.Status.AWAITING_QA, spec.Status.NEEDS_REVISION) in sources
|
|
|
|
|
|
def test_status_transitions_includes_ceo_paths() -> None:
|
|
sources = {(t.source, t.target) for t in spec._STATUS_TRANSITIONS}
|
|
assert (spec.Status.AWAITING_PM_REVIEW, spec.Status.COMPLETED) in sources
|
|
assert (
|
|
spec.Status.AWAITING_PM_REVIEW,
|
|
spec.Status.AWAITING_CEO_APPROVAL,
|
|
) in sources
|
|
assert (spec.Status.AWAITING_CEO_APPROVAL, spec.Status.COMPLETED) in sources
|
|
assert (spec.Status.AWAITING_CEO_APPROVAL, spec.Status.NEEDS_REVISION) in sources
|
|
# #100: a branchless coordination root rejected by the CEO routes to PENDING
|
|
# (Main PM re-plans) — the edge is in the spec so the audited privileged
|
|
# override that applies it can't be wedged by future admin-override tightening.
|
|
assert (spec.Status.AWAITING_CEO_APPROVAL, spec.Status.PENDING) in sources
|
|
# A blocked task the PM cannot resolve can also be surfaced to the CEO.
|
|
assert (spec.Status.BLOCKED, spec.Status.AWAITING_CEO_APPROVAL) in sources
|
|
|
|
|
|
def test_status_transitions_includes_block_pause_paths() -> None:
|
|
sources = {(t.source, t.target) for t in spec._STATUS_TRANSITIONS}
|
|
assert (spec.Status.IN_PROGRESS, spec.Status.BLOCKED) in sources
|
|
assert (spec.Status.IN_PROGRESS, spec.Status.PAUSED) in sources
|
|
assert (spec.Status.BLOCKED, spec.Status.IN_PROGRESS) in sources
|
|
assert (spec.Status.PAUSED, spec.Status.IN_PROGRESS) in sources
|
|
|
|
|
|
def test_every_non_terminal_status_can_be_cancelled() -> None:
|
|
"""PERMISSIONS.md says PM/CEO can cancel from any state."""
|
|
cancellable = {
|
|
t.source for t in spec._STATUS_TRANSITIONS if t.target == spec.Status.CANCELLED
|
|
}
|
|
non_terminal = set(spec.Status) - {spec.Status.COMPLETED, spec.Status.CANCELLED}
|
|
assert non_terminal <= cancellable, (
|
|
f"Statuses missing a cancel transition: {non_terminal - cancellable}"
|
|
)
|
|
|
|
|
|
def test_status_graph_lookup_returns_targets() -> None:
|
|
"""STATUS_GRAPH is a quick `source -> {targets}` lookup."""
|
|
assert spec.Status.CLAIMED in spec.STATUS_GRAPH[spec.Status.PENDING]
|
|
assert spec.Status.AWAITING_QA in spec.STATUS_GRAPH[spec.Status.VERIFYING]
|
|
assert spec.STATUS_GRAPH[spec.Status.COMPLETED] == frozenset()
|
|
|
|
|
|
def test_status_transitions_role_constraints_match_canon() -> None:
|
|
"""role_constraint must encode the per-row role gates from
|
|
PERMISSIONS.md / STATUS_TRANSITIONS.md exactly. Tests that look only
|
|
at (source, target) pairs miss role-typo regressions; this test
|
|
pins the gates explicitly.
|
|
"""
|
|
by_pair = {
|
|
(t.source, t.target, t.triggered_by_action): t.role_constraint
|
|
for t in spec._STATUS_TRANSITIONS
|
|
}
|
|
# QA is the only role that can claim awaiting_qa
|
|
assert by_pair[
|
|
(spec.Status.AWAITING_QA, spec.Status.CLAIMED, "claim")
|
|
] == frozenset({spec.Role.QA})
|
|
# Documenter is the only role that can claim awaiting_documentation
|
|
assert by_pair[
|
|
(spec.Status.AWAITING_DOCUMENTATION, spec.Status.CLAIMED, "claim")
|
|
] == frozenset({spec.Role.DOCUMENTER})
|
|
# qa_pass / qa_fail: QA only
|
|
assert by_pair[
|
|
(spec.Status.AWAITING_QA, spec.Status.AWAITING_DOCUMENTATION, "qa_pass")
|
|
] == frozenset({spec.Role.QA})
|
|
assert by_pair[
|
|
(spec.Status.AWAITING_QA, spec.Status.NEEDS_REVISION, "qa_fail")
|
|
] == frozenset({spec.Role.QA})
|
|
# docs_complete: documenter only
|
|
assert by_pair[
|
|
(
|
|
spec.Status.AWAITING_DOCUMENTATION,
|
|
spec.Status.AWAITING_PM_REVIEW,
|
|
"docs_complete",
|
|
)
|
|
] == frozenset({spec.Role.DOCUMENTER})
|
|
# PM complete: cell + main PM (not board, not CEO)
|
|
assert by_pair[
|
|
(spec.Status.AWAITING_PM_REVIEW, spec.Status.COMPLETED, "complete")
|
|
] == frozenset({spec.Role.CELL_PM, spec.Role.MAIN_PM})
|
|
# escalate_to_ceo: main_pm + product_owner + head_marketing — from a
|
|
# completed review and from a blocked task, same role gate.
|
|
escalate_roles = frozenset(
|
|
{
|
|
spec.Role.MAIN_PM,
|
|
spec.Role.PRODUCT_OWNER,
|
|
spec.Role.HEAD_MARKETING,
|
|
}
|
|
)
|
|
assert (
|
|
by_pair[
|
|
(
|
|
spec.Status.AWAITING_PM_REVIEW,
|
|
spec.Status.AWAITING_CEO_APPROVAL,
|
|
"escalate_to_ceo",
|
|
)
|
|
]
|
|
== escalate_roles
|
|
)
|
|
assert (
|
|
by_pair[
|
|
(
|
|
spec.Status.BLOCKED,
|
|
spec.Status.AWAITING_CEO_APPROVAL,
|
|
"escalate_to_ceo",
|
|
)
|
|
]
|
|
== escalate_roles
|
|
)
|
|
# CEO actions: CEO only
|
|
assert by_pair[
|
|
(spec.Status.AWAITING_CEO_APPROVAL, spec.Status.COMPLETED, "ceo_approve")
|
|
] == frozenset({spec.Role.CEO})
|
|
assert by_pair[
|
|
(spec.Status.AWAITING_CEO_APPROVAL, spec.Status.NEEDS_REVISION, "ceo_reject")
|
|
] == frozenset({spec.Role.CEO})
|
|
# Cancel: PM + CEO from any non-terminal status EXCEPT the CEO approval
|
|
# queue — cancelling a task the CEO is reviewing is the CEO's call, so
|
|
# awaiting_ceo_approval -> cancelled is gated to CEO only (a PM cancelling
|
|
# it would bypass the human CEO gate).
|
|
cancel_constraint = frozenset({spec.Role.CELL_PM, spec.Role.MAIN_PM, spec.Role.CEO})
|
|
for src in spec.Status:
|
|
if src in (spec.Status.COMPLETED, spec.Status.CANCELLED):
|
|
continue
|
|
expected = (
|
|
frozenset({spec.Role.CEO})
|
|
if src is spec.Status.AWAITING_CEO_APPROVAL
|
|
else cancel_constraint
|
|
)
|
|
assert by_pair[(src, spec.Status.CANCELLED, "cancel")] == expected, (
|
|
f"cancel from {src.value} has wrong role_constraint"
|
|
)
|
|
|
|
|
|
def test_atomic_action_table_has_pre_gateway_actions() -> None:
|
|
"""Every task tool from PERMISSIONS.md must have an ActionSpec."""
|
|
expected = {
|
|
"activate",
|
|
"claim",
|
|
"start",
|
|
"set_plan",
|
|
"block",
|
|
"unblock",
|
|
"pause",
|
|
"resume",
|
|
"submit_verification",
|
|
"submit_qa",
|
|
"qa_pass",
|
|
"qa_fail",
|
|
"docs_complete",
|
|
"complete",
|
|
"submit_pm_review",
|
|
"escalate_to_ceo",
|
|
"ceo_approve",
|
|
"ceo_reject",
|
|
"cancel",
|
|
"create_subtask",
|
|
}
|
|
assert expected <= set(spec._ATOMIC_ACTIONS), (
|
|
f"Missing ActionSpec entries: {expected - set(spec._ATOMIC_ACTIONS)}"
|
|
)
|
|
|
|
|
|
def test_claim_action_allows_developer_from_pending() -> None:
|
|
a = spec._ATOMIC_ACTIONS["claim"]
|
|
assert spec.Role.DEVELOPER in a.allowed_roles
|
|
assert spec.Status.PENDING in a.source_statuses
|
|
assert a.target_status == spec.Status.CLAIMED
|
|
|
|
|
|
def test_qa_pass_self_review_blocks() -> None:
|
|
"""A QA cannot qa_pass a task they themselves committed to."""
|
|
assert spec._ATOMIC_ACTIONS["qa_pass"].self_review_block is True
|
|
assert spec._ATOMIC_ACTIONS["qa_fail"].self_review_block is True
|
|
assert spec._ATOMIC_ACTIONS["docs_complete"].self_review_block is True
|
|
|
|
|
|
def test_claim_rules_match_pre_gateway_table() -> None:
|
|
"""PERMISSIONS.md "What Each Role Can Claim From" — exact match.
|
|
|
|
PMs claim from PENDING and NEEDS_REVISION — the latter to recover a rejected
|
|
coordination task (pr_fail / qa_fail / ceo_reject) by re-planning and
|
|
re-delegating fixes (scoped by give_me_work routing, which offers only the
|
|
caller's own assigned tasks). BACKLOG → PENDING is a separate `activate`
|
|
action (strict transitions; no implicit activate-on-claim).
|
|
"""
|
|
assert spec.CLAIM_RULES[spec.Role.DEVELOPER] == frozenset(
|
|
{spec.Status.PENDING, spec.Status.NEEDS_REVISION}
|
|
)
|
|
assert spec.CLAIM_RULES[spec.Role.QA] == frozenset({spec.Status.AWAITING_QA})
|
|
assert spec.CLAIM_RULES[spec.Role.DOCUMENTER] == frozenset(
|
|
{spec.Status.PENDING, spec.Status.AWAITING_DOCUMENTATION}
|
|
)
|
|
assert spec.CLAIM_RULES[spec.Role.CELL_PM] == frozenset(
|
|
{
|
|
spec.Status.PENDING,
|
|
spec.Status.NEEDS_REVISION,
|
|
spec.Status.AWAITING_PM_REVIEW,
|
|
}
|
|
)
|
|
assert spec.CLAIM_RULES[spec.Role.MAIN_PM] == frozenset(
|
|
{
|
|
spec.Status.PENDING,
|
|
spec.Status.NEEDS_REVISION,
|
|
spec.Status.AWAITING_PM_REVIEW,
|
|
}
|
|
)
|
|
|
|
|
|
def test_team_rules_pin_team_for_seeded_agents() -> None:
|
|
assert spec.ROLE_TEAM_RULES["be-dev-1"] == "backend"
|
|
assert spec.ROLE_TEAM_RULES["be-pm"] == "backend"
|
|
assert spec.ROLE_TEAM_RULES["fe-qa"] == "frontend"
|
|
assert spec.ROLE_TEAM_RULES["main-pm"] is None # cross-cell
|
|
|
|
|
|
def test_intent_verbs_table_has_every_gateway_verb() -> None:
|
|
"""Every gateway intent verb must have an IntentSpec."""
|
|
expected = {
|
|
"give_me_work",
|
|
"i_will_work_on",
|
|
"i_will_plan",
|
|
"delegate",
|
|
"open_pr",
|
|
"i_am_done",
|
|
"i_am_blocked",
|
|
"unclaim",
|
|
"resume",
|
|
"i_am_idle",
|
|
"claim_review",
|
|
"pass_review",
|
|
"fail_review",
|
|
"claim_doc_task",
|
|
"i_documented",
|
|
"complete",
|
|
"escalate_up",
|
|
"escalate_to_ceo",
|
|
"submit_up",
|
|
"unblock",
|
|
"triage",
|
|
"triage_all",
|
|
}
|
|
assert expected <= set(spec._INTENT_VERBS), (
|
|
f"Missing IntentSpec entries: {expected - set(spec._INTENT_VERBS)}"
|
|
)
|
|
|
|
|
|
def test_i_will_work_on_composes_claim_set_plan_start() -> None:
|
|
iv = spec._INTENT_VERBS["i_will_work_on"]
|
|
assert iv.composes == ("claim", "set_plan", "start")
|
|
assert spec.Role.DEVELOPER in iv.allowed_roles
|
|
|
|
|
|
def test_i_will_plan_composes_claim_set_plan_start() -> None:
|
|
"""PMs use i_will_plan; the composition mirrors i_will_work_on."""
|
|
iv = spec._INTENT_VERBS["i_will_plan"]
|
|
assert iv.composes == ("claim", "set_plan", "start")
|
|
assert iv.allowed_roles == frozenset({spec.Role.CELL_PM, spec.Role.MAIN_PM})
|
|
|
|
|
|
def test_i_am_done_composes_submit_verification_then_submit_qa() -> None:
|
|
iv = spec._INTENT_VERBS["i_am_done"]
|
|
assert iv.composes == ("submit_verification", "submit_qa")
|
|
|
|
|
|
def test_open_pr_has_git_side_effects() -> None:
|
|
"""open_pr is a side-effect-only verb (no DB transition)."""
|
|
iv = spec._INTENT_VERBS["open_pr"]
|
|
assert "push_branch" in iv.side_effects
|
|
assert "create_pr" in iv.side_effects
|
|
assert iv.composes == () # pure side effect verb
|
|
|
|
|
|
def test_delegate_composes_create_subtask() -> None:
|
|
iv = spec._INTENT_VERBS["delegate"]
|
|
assert iv.composes == ("create_subtask",)
|
|
assert iv.allowed_roles == frozenset({spec.Role.CELL_PM, spec.Role.MAIN_PM})
|
|
|
|
|
|
_STUB_TASK_DEFAULTS: dict[str, Any] = {
|
|
"status": "pending",
|
|
"task_type": "code",
|
|
"commits": [],
|
|
"plan": None,
|
|
"assigned_to": None,
|
|
"pr_number": None,
|
|
}
|
|
|
|
|
|
def _stub_task(**overrides: Any) -> SimpleNamespace:
|
|
fields = {**_STUB_TASK_DEFAULTS, **overrides}
|
|
fields["commits"] = fields["commits"] or []
|
|
return SimpleNamespace(**fields)
|
|
|
|
|
|
def test_can_claim_developer_pending_allowed() -> None:
|
|
d = spec.can_claim(spec.Role.DEVELOPER, _stub_task(status="pending"))
|
|
assert d.allowed is True
|
|
|
|
|
|
def test_can_claim_developer_completed_rejected() -> None:
|
|
d = spec.can_claim(spec.Role.DEVELOPER, _stub_task(status="completed"))
|
|
assert d.allowed is False
|
|
assert d.rejection_kind == "invalid_state"
|
|
|
|
|
|
def test_can_claim_developer_awaiting_qa_rejected() -> None:
|
|
"""Devs cannot claim awaiting_qa - that's QA's path."""
|
|
d = spec.can_claim(spec.Role.DEVELOPER, _stub_task(status="awaiting_qa"))
|
|
assert d.allowed is False
|
|
assert d.rejection_kind == "not_authorized"
|
|
|
|
|
|
def test_can_invoke_intent_developer_can_call_i_will_work_on() -> None:
|
|
d = spec.can_invoke_intent(
|
|
spec.Role.DEVELOPER,
|
|
"i_will_work_on",
|
|
_stub_task(status="pending"),
|
|
context=spec.Context(plan="my plan"),
|
|
)
|
|
assert d.allowed is True
|
|
|
|
|
|
def test_can_invoke_intent_pm_cannot_call_i_will_work_on() -> None:
|
|
"""PMs use i_will_plan; i_will_work_on is dev-only."""
|
|
d = spec.can_invoke_intent(
|
|
spec.Role.CELL_PM,
|
|
"i_will_work_on",
|
|
_stub_task(status="pending"),
|
|
context=spec.Context(plan="x"),
|
|
)
|
|
assert d.allowed is False
|
|
assert d.rejection_kind == "not_authorized"
|
|
|
|
|
|
def test_can_invoke_intent_developer_open_pr_no_commits_tracing_gap() -> None:
|
|
"""open_pr requires >=1 commit. Without one -> tracing_gap."""
|
|
d = spec.can_invoke_intent(
|
|
spec.Role.DEVELOPER,
|
|
"open_pr",
|
|
_stub_task(status="in_progress", commits=[]),
|
|
context=spec.Context(),
|
|
)
|
|
assert d.allowed is False
|
|
assert d.rejection_kind == "tracing_gap"
|
|
assert "commits>=1" in d.missing
|
|
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
# open_pr must enforce the PR-open state gate (parity with the HTTP path)
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
|
|
def _owned_task(**overrides: Any) -> SimpleNamespace:
|
|
"""A task owned by ``actor`` with commits and no prior PR — only the state
|
|
gate can fail, isolating the PR-open-state precondition."""
|
|
actor = overrides.pop("actor_id", uuid4())
|
|
return _stub_task(
|
|
assigned_to=actor,
|
|
commits=["abc123"],
|
|
pr_number=None,
|
|
**overrides,
|
|
)
|
|
|
|
|
|
def test_open_pr_rejected_on_claimed_task() -> None:
|
|
"""``open_pr`` must be rejected from ``claimed`` — only ``in_progress`` may
|
|
open a PR (mirrors the HTTP path's ``_assert_pr_create_allowed``)."""
|
|
actor = uuid4()
|
|
d = spec.can_invoke_intent(
|
|
spec.Role.DEVELOPER,
|
|
"open_pr",
|
|
_owned_task(status="claimed", actor_id=actor),
|
|
context=spec.Context(actor_id=actor),
|
|
)
|
|
assert d.allowed is False
|
|
assert d.rejection_kind == "invalid_state"
|
|
|
|
|
|
def test_open_pr_rejected_on_paused_task() -> None:
|
|
actor = uuid4()
|
|
d = spec.can_invoke_intent(
|
|
spec.Role.DEVELOPER,
|
|
"open_pr",
|
|
_owned_task(status="paused", actor_id=actor),
|
|
context=spec.Context(actor_id=actor),
|
|
)
|
|
assert d.allowed is False
|
|
assert d.rejection_kind == "invalid_state"
|
|
|
|
|
|
def test_open_pr_rejected_on_blocked_task() -> None:
|
|
actor = uuid4()
|
|
d = spec.can_invoke_intent(
|
|
spec.Role.DEVELOPER,
|
|
"open_pr",
|
|
_owned_task(status="blocked", actor_id=actor),
|
|
context=spec.Context(actor_id=actor),
|
|
)
|
|
assert d.allowed is False
|
|
assert d.rejection_kind == "invalid_state"
|
|
|
|
|
|
def test_open_pr_rejected_on_completed_task() -> None:
|
|
"""A completed task is terminal — opening a PR on it is nonsensical."""
|
|
actor = uuid4()
|
|
d = spec.can_invoke_intent(
|
|
spec.Role.DEVELOPER,
|
|
"open_pr",
|
|
_owned_task(status="completed", actor_id=actor),
|
|
context=spec.Context(actor_id=actor),
|
|
)
|
|
assert d.allowed is False
|
|
assert d.rejection_kind == "invalid_state"
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
"status",
|
|
[
|
|
"in_progress",
|
|
"verifying",
|
|
"awaiting_qa",
|
|
"awaiting_documentation",
|
|
"needs_revision",
|
|
],
|
|
)
|
|
def test_open_pr_allowed_in_pr_open_states(status: str) -> None:
|
|
"""Regression guard: every PR-open-eligible state still lets the owner open
|
|
a PR — the new state gate must not over-restrict the legitimate path."""
|
|
actor = uuid4()
|
|
d = spec.can_invoke_intent(
|
|
spec.Role.DEVELOPER,
|
|
"open_pr",
|
|
_owned_task(status=status, actor_id=actor),
|
|
context=spec.Context(actor_id=actor),
|
|
)
|
|
assert d.allowed is True, f"open_pr should be allowed from {status}"
|
|
|
|
|
|
def test_open_pr_state_gate_takes_priority_over_unowned() -> None:
|
|
"""A non-owner in a wrong state: ownership (not_authorized) is checked
|
|
before state, mirroring the HTTP path's assignee-first ordering."""
|
|
d = spec.can_invoke_intent(
|
|
spec.Role.DEVELOPER,
|
|
"open_pr",
|
|
_owned_task(status="claimed", actor_id=uuid4()),
|
|
context=spec.Context(actor_id=uuid4()), # different actor -> not owner
|
|
)
|
|
assert d.allowed is False
|
|
assert d.rejection_kind == "not_authorized"
|
|
|
|
|
|
def test_escalate_up_rejected_on_completed_task() -> None:
|
|
"""A PM must not resurrect a COMPLETED task via ``escalate_up`` — the spec
|
|
gate rejects terminal tasks before the journal:decision write fires."""
|
|
d = spec.can_invoke_intent(
|
|
spec.Role.CELL_PM,
|
|
"escalate_up",
|
|
_stub_task(status="completed"),
|
|
context=spec.Context(notes="stuck on something"),
|
|
)
|
|
assert d.allowed is False
|
|
assert d.rejection_kind == "invalid_state"
|
|
|
|
|
|
def test_escalate_up_rejected_on_cancelled_task() -> None:
|
|
"""Cancelled is terminal — escalate_up must not resurrect it either."""
|
|
d = spec.can_invoke_intent(
|
|
spec.Role.MAIN_PM,
|
|
"escalate_up",
|
|
_stub_task(status="cancelled"),
|
|
context=spec.Context(notes="stuck on something"),
|
|
)
|
|
assert d.allowed is False
|
|
assert d.rejection_kind == "invalid_state"
|
|
|
|
|
|
def test_escalate_up_allowed_on_blocked_task() -> None:
|
|
"""The terminal guard must not over-restrict — BLOCKED is the natural
|
|
escalation source and must still be allowed."""
|
|
d = spec.can_invoke_intent(
|
|
spec.Role.CELL_PM,
|
|
"escalate_up",
|
|
_stub_task(status="blocked"),
|
|
context=spec.Context(notes="stuck on something"),
|
|
)
|
|
assert d.allowed is True
|
|
|
|
|
|
def test_valid_next_verbs_developer_in_progress_includes_open_pr_and_i_am_done() -> (
|
|
None
|
|
):
|
|
verbs = spec.valid_next_verbs(spec.Role.DEVELOPER, _stub_task(status="in_progress"))
|
|
assert "open_pr" in verbs
|
|
assert "i_am_done" in verbs
|
|
assert "i_am_blocked" in verbs
|
|
|
|
|
|
def test_valid_next_verbs_pm_pending_includes_i_will_plan() -> None:
|
|
# A PM i_will_plan's a PLANNING task (coordination), not a code task — the
|
|
# PM/code claim carve-out (Fix 2) removes i_will_plan from a pending code
|
|
# task's verb set, so the legitimate path is exercised with task_type=planning.
|
|
verbs = spec.valid_next_verbs(
|
|
spec.Role.CELL_PM, _stub_task(status="pending", task_type="planning")
|
|
)
|
|
assert "i_will_plan" in verbs
|
|
|
|
|
|
def test_composed_actions_for_returns_intent_composition() -> None:
|
|
assert spec.composed_actions_for("i_will_work_on") == ("claim", "set_plan", "start")
|
|
assert spec.composed_actions_for("open_pr") == ()
|
|
|
|
|
|
def test_intents_for_role_returns_role_scoped_verbs() -> None:
|
|
dev_verbs = spec.intents_for_role(spec.Role.DEVELOPER)
|
|
assert "i_will_work_on" in dev_verbs
|
|
assert "open_pr" in dev_verbs
|
|
assert "i_am_done" in dev_verbs
|
|
assert "delegate" not in dev_verbs # PM only
|
|
assert "claim_review" not in dev_verbs # QA only
|
|
|
|
|
|
def test_status_after_returns_target_status() -> None:
|
|
assert spec.status_after("claim", spec.Status.PENDING) == spec.Status.CLAIMED
|
|
assert (
|
|
spec.status_after("submit_qa", spec.Status.VERIFYING) == spec.Status.AWAITING_QA
|
|
)
|
|
assert (
|
|
spec.status_after("set_plan", spec.Status.IN_PROGRESS) is None
|
|
) # no transition
|
|
|
|
|
|
def test_can_invoke_intent_open_pr_passes_when_owner_with_commits() -> None:
|
|
"""Green path for open_pr: owner + commits + no prior PR → allow."""
|
|
owner_id = uuid4()
|
|
task = _stub_task(
|
|
status="in_progress",
|
|
commits=["abc"],
|
|
pr_number=None,
|
|
assigned_to=owner_id,
|
|
)
|
|
d = spec.can_invoke_intent(
|
|
spec.Role.DEVELOPER,
|
|
"open_pr",
|
|
task,
|
|
context=spec.Context(actor_id=owner_id),
|
|
)
|
|
assert d.allowed is True, f"expected allow, got {d}"
|
|
|
|
|
|
def test_can_invoke_intent_open_pr_rejects_non_owner() -> None:
|
|
"""Non-owner trying open_pr → not_authorized (PRECONDITION_OWNERSHIP)."""
|
|
owner_id = uuid4()
|
|
intruder_id = uuid4()
|
|
task = _stub_task(
|
|
status="in_progress",
|
|
commits=["abc"],
|
|
pr_number=None,
|
|
assigned_to=owner_id,
|
|
)
|
|
d = spec.can_invoke_intent(
|
|
spec.Role.DEVELOPER,
|
|
"open_pr",
|
|
task,
|
|
context=spec.Context(actor_id=intruder_id),
|
|
)
|
|
assert d.allowed is False
|
|
assert d.rejection_kind == "not_authorized"
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Task 8 — self-consistency validators (`_validate.py`)
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def test_validators_pass_on_real_spec() -> None:
|
|
"""Importing roboco.foundation.policy.lifecycle must not raise —
|
|
module-level import IS the test. We additionally call the runner
|
|
directly so a future refactor that detaches it from import doesn't
|
|
silently skip the gate.
|
|
"""
|
|
_validate.run_all_lifecycle_validators()
|
|
|
|
|
|
def test_every_status_reachable_from_pending() -> None:
|
|
"""Reachability — except CANCELLED is its own thing and BACKLOG predates pending."""
|
|
reachable = reachable_from(spec.Status.PENDING)
|
|
expected_reachable = set(spec.Status) - {spec.Status.BACKLOG, spec.Status.CANCELLED}
|
|
assert expected_reachable <= reachable, (
|
|
f"Unreachable from pending: {expected_reachable - reachable}"
|
|
)
|
|
|
|
|
|
def test_every_intent_verb_composes_known_actions() -> None:
|
|
"""Every IntentSpec.composes must reference declared atomic actions."""
|
|
for name, iv in spec._INTENT_VERBS.items():
|
|
for action_name in iv.composes:
|
|
assert action_name in spec._ATOMIC_ACTIONS, (
|
|
f"Intent '{name}' composes unknown action '{action_name}'"
|
|
)
|
|
|
|
|
|
def test_self_review_symmetry() -> None:
|
|
"""If qa_pass blocks, qa_fail and docs_complete must too."""
|
|
qp = spec._ATOMIC_ACTIONS["qa_pass"].self_review_block
|
|
qf = spec._ATOMIC_ACTIONS["qa_fail"].self_review_block
|
|
dc = spec._ATOMIC_ACTIONS["docs_complete"].self_review_block
|
|
assert qp == qf == dc, (
|
|
"self_review_block asymmetry between qa_pass/qa_fail/docs_complete"
|
|
)
|
|
|
|
|
|
def test_run_all_validators_raises_on_unknown_intent_action(
|
|
monkeypatch: pytest.MonkeyPatch,
|
|
) -> None:
|
|
"""If an IntentSpec.composes references a non-existent action, the
|
|
validator must raise LifecycleSpecError. Pins the gate's actual
|
|
behavior — without this test, refactors that move run_all_validators()
|
|
out of the import path could silently disable the gate.
|
|
"""
|
|
iv = _INTENT_VERBS["delegate"]
|
|
broken = IntentSpec(
|
|
name=iv.name,
|
|
allowed_roles=iv.allowed_roles,
|
|
description=iv.description,
|
|
composes=("create_subtask", "ZZZ_FAKE_ACTION_DOES_NOT_EXIST"),
|
|
extra_preconditions=iv.extra_preconditions,
|
|
side_effects=iv.side_effects,
|
|
next_hint=iv.next_hint,
|
|
)
|
|
patched_intents = dict(_INTENT_VERBS)
|
|
patched_intents["delegate"] = broken
|
|
monkeypatch.setattr(
|
|
"roboco.foundation.policy.lifecycle._INTENT_VERBS", patched_intents
|
|
)
|
|
with pytest.raises(_validate.LifecycleSpecError, match="ZZZ_FAKE_ACTION"):
|
|
_validate.run_all_lifecycle_validators()
|
|
|
|
|
|
def test_next_hint_pr_fail_main_pm_root_steers_to_redelegate() -> None:
|
|
"""A ``pr_fail`` on a Main-PM branch-bearing root must steer the Main PM to
|
|
re-delegate the fixes, NOT re-submit the unchanged root. The root is an
|
|
assembled cell→root / root→master PR — coordination, not the Main PM's own
|
|
code — so re-submitting it is the 2026-06-27 infinite ``pr_fail`` loop."""
|
|
t = SimpleNamespace(team=spec.Team.MAIN_PM, branch_name="feature/main_pm/c80e19ff")
|
|
hint = _INTENT_VERBS["pr_fail"].next_hint(t)
|
|
assert "re-delegate" in hint
|
|
assert "do NOT re-submit" in hint
|
|
|
|
|
|
def test_next_hint_pr_fail_cell_dev_keeps_dev_revise() -> None:
|
|
"""A cell / dev task is revised in place by its dev, so ``pr_fail`` keeps the
|
|
dev-revise hint (the cell→root PR carries that dev's own code)."""
|
|
t = SimpleNamespace(team=spec.Team.BACKEND, branch_name="feature/backend/abc12345")
|
|
hint = _INTENT_VERBS["pr_fail"].next_hint(t)
|
|
assert hint == "idle - dev will revise and re-submit"
|
|
|
|
|
|
def test_next_hint_pr_fail_branchless_main_pm_keeps_dev_revise() -> None:
|
|
"""A branchless Main-PM umbrella (no ``branch_name``) assembles no PR of its
|
|
own, so the gate never lands a ``pr_fail`` on it — but defensively it keeps
|
|
the dev-revise hint rather than the re-delegate steer."""
|
|
t = SimpleNamespace(team=spec.Team.MAIN_PM, branch_name=None)
|
|
hint = _INTENT_VERBS["pr_fail"].next_hint(t)
|
|
assert hint == "idle - dev will revise and re-submit"
|
|
|
|
|
|
def test_unmigrated_is_pinned() -> None:
|
|
"""The known-debt set; remove an entry once that consumer is migrated."""
|
|
assert (
|
|
frozenset(
|
|
{
|
|
"enforcement.task_lifecycle._LEGACY_OPERATIONAL_EDGES",
|
|
"enforcement.task_lifecycle._LEGACY_ROLE_GATES",
|
|
}
|
|
)
|
|
== spec.UNMIGRATED
|
|
)
|
|
|
|
|
|
# --- PM/code claim invariant (Fix 2 + bug-1): the claim gate does NOT block a
|
|
# PM claiming a code task. A PM's only claim verb is i_will_plan, and planning a
|
|
# code-typed PARENT (to decompose + delegate the code) is legitimate (bug-1:
|
|
# scoping pm_cannot_execute_code to i_will_plan deadlocked the slice). Execution
|
|
# is blocked at the intent level — i_will_work_on is _DEV_ROLES only. The
|
|
# create/delegate guards (pm_cannot_own_code) block a PM from being ASSIGNED a
|
|
# fresh code task; the needs_revision carve-out (a PM resolving review issues
|
|
# directly / recovering a rejected coordination task) is naturally allowed
|
|
# because PMs claim NEEDS_REVISION. These pin that the claim gate does not
|
|
# regress bug-1.
|
|
|
|
|
|
def _claim_task(*, status: str, task_type: str) -> Any:
|
|
return SimpleNamespace(status=status, task_type=task_type)
|
|
|
|
|
|
def test_claim_allows_cell_pm_claiming_code_from_pending() -> None:
|
|
"""bug-1: a cell PM i_will_plan-ing a code-typed parent (PENDING) to plan +
|
|
delegate the code MUST be allowed — rejecting it deadlocks the slice."""
|
|
t = _claim_task(status="pending", task_type="code")
|
|
d = spec.can_invoke_action(spec.Role.CELL_PM, "claim", t)
|
|
assert d.allowed, d.message
|
|
|
|
|
|
def test_claim_allows_cell_pm_claiming_code_from_needs_revision() -> None:
|
|
"""Carve-out: a PM may take a code task in needs_revision to resolve the
|
|
review/QA issues directly / recover a rejected coordination task."""
|
|
t = _claim_task(status="needs_revision", task_type="code")
|
|
d = spec.can_invoke_action(spec.Role.CELL_PM, "claim", t)
|
|
assert d.allowed, d.message
|
|
|
|
|
|
def test_claim_allows_main_pm_claiming_code_from_needs_revision() -> None:
|
|
"""The same carve-out holds for the Main PM (coordination-recovery path)."""
|
|
t = _claim_task(status="needs_revision", task_type="code")
|
|
d = spec.can_invoke_action(spec.Role.MAIN_PM, "claim", t)
|
|
assert d.allowed, d.message
|
|
|
|
|
|
def test_claim_allows_main_pm_claiming_code_from_pending() -> None:
|
|
"""bug-1 parity: a Main PM planning a code-typed parent (PENDING) is allowed
|
|
for the same reason as the cell PM — execution is blocked at i_will_work_on,
|
|
not at the claim gate."""
|
|
t = _claim_task(status="pending", task_type="code")
|
|
d = spec.can_invoke_action(spec.Role.MAIN_PM, "claim", t)
|
|
assert d.allowed, d.message
|
|
|
|
|
|
def test_claim_allows_pm_claiming_planning_from_pending() -> None:
|
|
"""A PM claiming a planning task is the legitimate coordination path."""
|
|
t = _claim_task(status="pending", task_type="planning")
|
|
assert spec.can_invoke_action(spec.Role.CELL_PM, "claim", t).allowed
|
|
assert spec.can_invoke_action(spec.Role.MAIN_PM, "claim", t).allowed
|
|
|
|
|
|
def test_claim_allows_developer_claiming_code_from_pending() -> None:
|
|
"""A developer claiming fresh code is unaffected (the PM invariant is
|
|
enforced at create/delegate + i_will_work_on, not the claim gate)."""
|
|
t = _claim_task(status="pending", task_type="code")
|
|
assert spec.can_invoke_action(spec.Role.DEVELOPER, "claim", t).allowed
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Edge cases — logical-gap element sweep (2026-06-30)
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def test_claim_pr_review_rejected_on_gate_task_points_to_claim_gate_review() -> None:
|
|
"""claim_pr_review is for an inbound external-PR task in PENDING only. An
|
|
awaiting_pr_review gate task must be rejected (and remediation must point
|
|
the reviewer at claim_gate_review), not silently accepted by the spec gate."""
|
|
d = spec.can_invoke_intent(
|
|
spec.Role.PR_REVIEWER,
|
|
"claim_pr_review",
|
|
_stub_task(status="awaiting_pr_review"),
|
|
)
|
|
assert d.allowed is False
|
|
assert d.rejection_kind == "invalid_state"
|
|
assert "claim_gate_review" in (d.remediate or "")
|
|
|
|
|
|
def test_claim_pr_review_allowed_on_pending_external_review() -> None:
|
|
"""Green path: a pending external-PR review task is claimable."""
|
|
d = spec.can_invoke_intent(
|
|
spec.Role.PR_REVIEWER,
|
|
"claim_pr_review",
|
|
_stub_task(status="pending"),
|
|
)
|
|
assert d.allowed is True
|
|
|
|
|
|
def test_needs_team_match_rejects_cross_team_claim_when_agent_team_supplied() -> None:
|
|
"""needs_team_match was a dead spec field; when the caller supplies the
|
|
agent's team via Context, the spec gate must enforce it (a backend dev
|
|
cannot claim a frontend task)."""
|
|
d = spec.can_invoke_action(
|
|
spec.Role.DEVELOPER,
|
|
"claim",
|
|
_stub_task(status="pending", team="frontend"),
|
|
context=spec.Context(agent_team="backend"),
|
|
)
|
|
assert d.allowed is False
|
|
assert d.rejection_kind == "not_authorized"
|
|
|
|
|
|
def test_needs_team_match_allows_same_team_claim() -> None:
|
|
d = spec.can_invoke_action(
|
|
spec.Role.DEVELOPER,
|
|
"claim",
|
|
_stub_task(status="pending", team="backend"),
|
|
context=spec.Context(agent_team="backend"),
|
|
)
|
|
assert d.allowed is True
|
|
|
|
|
|
def test_needs_team_match_defers_when_agent_team_absent() -> None:
|
|
"""Backward compat: without agent_team in Context, the spec gate stays
|
|
permissive (the service layer still enforces team-match)."""
|
|
d = spec.can_invoke_action(
|
|
spec.Role.DEVELOPER,
|
|
"claim",
|
|
_stub_task(status="pending", team="frontend"),
|
|
)
|
|
assert d.allowed is True
|
|
|
|
|
|
def test_valid_next_verbs_omits_claim_review_when_qa_not_in_awaiting_qa() -> None:
|
|
"""valid_next_verbs must apply claim-rule narrowing for empty-compose
|
|
claim verbs; a QA reviewer on a COMPLETED task must not be told
|
|
claim_review is callable."""
|
|
verbs = spec.valid_next_verbs(spec.Role.QA, _stub_task(status="completed"))
|
|
assert "claim_review" not in verbs
|
|
|
|
|
|
def test_valid_next_verbs_includes_claim_review_for_qa_in_awaiting_qa() -> None:
|
|
verbs = spec.valid_next_verbs(spec.Role.QA, _stub_task(status="awaiting_qa"))
|
|
assert "claim_review" in verbs
|
|
|
|
|
|
def test_pr_reviewer_has_unclaim_release_verb() -> None:
|
|
"""A PR reviewer who cannot finish a review must have a self-release
|
|
verb (unclaim), not wedge the lane until the stale-claim reaper."""
|
|
assert "unclaim" in spec.intents_for_role(spec.Role.PR_REVIEWER)
|
|
|
|
|
|
def test_unclaim_allowed_for_pr_reviewer() -> None:
|
|
d = spec.can_invoke_intent(
|
|
spec.Role.PR_REVIEWER,
|
|
"unclaim",
|
|
_stub_task(status="awaiting_pr_review"),
|
|
)
|
|
assert d.allowed is True
|
|
|
|
|
|
def test_complete_intent_declares_no_inverted_pr_merge_side_effect() -> None:
|
|
"""complete's IntentSpec must not declare a trailing pr_merge side_effect:
|
|
TaskService.complete asserts the PR is already merged, so the merge runs
|
|
FIRST (choreographer verb body owns the ordering). The spec must match
|
|
reality, not lie about a complete-then-merge composition."""
|
|
iv = spec._INTENT_VERBS["complete"]
|
|
assert iv.composes == ("complete",)
|
|
assert iv.side_effects == ()
|