mirror of
https://github.com/rennf93/roboco.git
synced 2026-08-03 07:23:24 +02:00
* feat(sandbox): throwaway per-agent Postgres/Redis sandbox containers
Orchestrator-provisioned sibling containers per agent spawn
(SandboxProvisioner, roboco/runtime/sandbox.py). Per-project opt-in via
projects.sandbox_services (migration 057); master switch
ROBOCO_SANDBOX_DB_ENABLED, default-off, armed in the NAS compose only.
When active, ROBOCO_TEST_DB_* / ROBOCO_TEST_REDIS_* point at the sandbox
and the prod-creds gate-env injection is suppressed (sandbox replaces,
never coexists). Sandbox lifetime tracks the agent container: teardown at
every removal path, orphan janitor at startup + each reaper tick with a
grace window for mid-flight spawns. The pre-spawn stale-clear spares the
just-provisioned sandbox; provision pre-clears stale same-named
containers from a crash-missed teardown.
Panel: per-project sandbox-service switches in the edit dialog + feature
flag card entry.
* docs: CLAUDE.md entry for the sandboxed dev DB/Redis subsystem
* feat(security): isolate prod Postgres/Redis from agent containers (roboco_data network)
Second user-defined bridge roboco_data carries postgres+redis only; the
orchestrator is multi-homed (default + data). Spawned agents and their
sandbox sidecars stay on roboco_default and can no longer resolve or
reach roboco-postgres:5432 / roboco-redis:6379 (redis has no auth —
membership is its only containment). Normal bridge, so host-published
ports (15432/16379) keep working. Applied to both build composes and
the registry compose; docker-compose.yml re-synced byte-identical with
docker-compose.yaml (it had drifted by the sandbox flag block).
ROBOCO_DB_NETWORK_ISOLATED (config default false, armed alongside the
topology) suppresses the legacy _append_gate_env prod-creds injection:
under isolation those creds dead-end, and unreachable creds are worse
than none. DB-needing projects opt into sandbox_services instead. The
flag is deliberately not a panel feature flag - it must travel with the
compose networks: stanzas.
Preserved by construction: agent<->agent A2A and orchestrator->agent SDK
polls on :9000, MCP->orchestrator on :8000, ollama reachability, docker
exec/inspect (daemon socket), host port publishing.
* feat(panel): full mobile responsiveness pass
Shared primitives: useIsMobile (useSyncExternalStore, hydration-safe,
memoized matchMedia subscribe), ResponsiveTable table->card switch below
md (single subtree mounted, no duplicated interactive rows), scrollable
snap TabsList in the base primitive (justify-center-safe so the first
tab stays reachable on overflow), persistent md:hidden bottom tab bar
(Overview/Tasks/Kanban/Chat, safe-area padded).
Applied: card lists for tasks/projects/products/work-sessions/sessions
+ the three raw metrics tables; CEO approval queue / release proposal /
playbook review action rows stack on narrow; command-center reorders
approvals above the fold on mobile; task-header metadata wraps;
Communications + A2A become URL-driven single-pane drill-downs below lg
(fixes the unconstrained-height ScrollArea bug) with dvh heights;
recharts label density/radius adapts via useIsMobile; git diff viewer
gets mobile font + wrap toggle; vh->dvh sweep; chat composers get
safe-area-inset padding; dashboard main p-4 md:p-6 + pb-20 for the bar.
Verified at 375px on the built app: bottom bar, drawer, approval-first
overview, swipeable kanban tab strip. All gates green (eslint, tsc,
vitest 249, next build 24/24 routes).
* feat(auth): cloud auth via FastAPI Users (default-off, single-user cookie session)
ROBOCO_CLOUD_AUTH_ENABLED (default off) lets the panel/API be exposed
beyond localhost without changing the CEO's local no-login flow while
off — get_agent_context and the WS gate are byte-for-byte unchanged in
off-mode. On: header-trust dies for humans — any agent-role claim (ceo
or a privileged PM/board role) with no valid HMAC token or session
cookie is 401, closing the header-spoof hole on the host-published
:8000 port for every role. The agent-fleet HMAC path and the system
self-PATCH keep working unmodified in both modes.
Single seeded CEO user (migration 058 users table, UserTable), no
registration router — idempotent env-driven upsert at startup by PK.
Cookie transport (httponly/secure/samesite=lax) + a JWTStrategy bound
to a fingerprint of the current password hash (rotating the password
invalidates every prior session). Sliding 30-day session: every
authenticated request re-mints the cookie, so an active session never
expires — no unexpected logouts.
Panel: (auth)/login page + proxy.ts (Next 16 rename of middleware; probes
/auth/status over the docker-internal URL, fails open to off) gate the
dashboard; client.ts gets withCredentials + 401->/login. nginx unchanged.
Review hardening: broadened the on-mode rejection from ceo-only to every
non-CEO role without a valid token (was only closed when
ROBOCO_AGENT_AUTH_REQUIRED was also armed); Next-16 proxy.ts rename to
clear the middleware deprecation warning.
* feat(x): RoboCo X account engine — HoM drafts, per-post CEO approval (default-off)
ROBOCO_X_ENGINE_ENABLED (default off, inert without creds). Mirrors the
ReleaseManagerEngine held-artifact shape: XEngine drafts a post when a
release publishes (via a draft_release_post seam on ReleaseProposalService
.approve) and drafts replies to meaningful mentions (dedicated poll loop,
x_seen_mentions dedup ledger, per-cycle/open caps). Drafting is
local-model-only, clamped to 280 chars. Nothing auto-posts — every tweet
is a held task (source x_post/x_reply, confirmed_by_human=False,
Secretary-owned, dispatcher-skipped) the CEO edits/approves/rejects in a
panel queue.
The four OAuth 1.0a secrets live Fernet-encrypted in a singleton
x_credentials row (migration 059, all-or-nothing, API returns only
has_credentials); decryption is server-side, agents never hold creds or
egress. Hand-rolled OAuth 1.0a HMAC-SHA1 signer, no new dependency;
NullXClient makes the unconfigured path a graceful no-op.
XPostService.approve (CEO-only) is the sole caller of post_tweet.
Review hardening: closed a double-post race — the approve path now
re-reads committed task state inside the Redis lock and commits COMPLETED
before releasing, so a concurrent approve that acquires the lock after the
winner released can't re-post (SET-NX is non-waiting, and the route-level
commit landed after the lock dropped). Added a regression test.
* feat(roadmap): board roadmap engine — PO proposes themed cycles, CEO approves per-item (default-off)
ROBOCO_ROADMAP_ENGINE_ENABLED (default off). Weekly, RoadmapEngine opens
ONE held exploration task (source=board_roadmap, confirmed_by_human=False,
Product-Owner-assigned), deduped to one open cycle. A dedicated one-shot
_dispatch_roadmap_exploration spawns the PO solo (not the two-reviewer
board path, which would also spawn HoM + fire Approve-&-Start). The PO
explores read-only (git/KB/metrics/releases/charter/web) and makes one
propose_roadmap call (PO-only content verb) authoring a themed cycle —
goal + 3-7 item drafts — persisted as a roadmap_cycle marker (no table,
no migration; head stays 059).
The CEO acts per-item in the panel roadmap queue: approve materializes a
BACKLOG task (source=roadmap, no assignee — never auto-starts), reject
records a reason; all-items-terminal completes the exploration task.
RoadmapService is idempotent per item. Dispatchers skip board_roadmap.
Includes a real SQLAlchemy dirty-check fix (deep-copy the JSON marker
before mutating, or the in-place edit + reassign compares equal to its
own baseline and the UPDATE is skipped).
Review hardening: create_task_from_draft now honors a draft-declared
source only from a {prompter, roadmap} whitelist — drafts are
LLM-authored, so an unbounded source could impersonate a privileged
origin (release_manager would even wedge that engine's dedup).
* chore(release): 0.17.0
Wave 3 — six default-off subsystems: sandboxed dev DB/Redis, prod
Postgres/Redis network isolation, full mobile UI pass, cloud auth
(FastAPI Users), the RoboCo X account engine, and the board roadmap
engine. Plus the waves 1+2 work already on master since 0.16.0.
Version bumped across the canonical set (config.py, __init__.py,
pyproject.toml, panel/package.json, uv.lock); CHANGELOG [Unreleased]
cut to [0.17.0]; docs/map delta added.
Compose: every optional feature armed :-true in the NAS composes, OFF
in the user-facing registry compose. Two opt-in exceptions default off
(CLOUD_AUTH — needs email/password/secret + TLS, would otherwise fail
startup; ROUTING_STRICT — fail-closed spawning). DB_NETWORK_ISOLATED
stays on in both (coupled to the roboco_data topology).
* chore(compose): arm cloud_auth + routing_strict ON in the NAS composes
Every feature defaults ON in the NAS composes per policy — these two
were wrongly left off. Both keep the ${VAR:-true} form so the operator
controls the real runtime via .env: cloud auth needs
ROBOCO_CLOUD_AUTH_EMAIL/_PASSWORD/_SECRET + TLS set there before a boot
(else startup fails loud), and routing_strict is fail-closed. Registry
compose keeps both off.
* fix(ci): reflow board.md prose (quality gate) + document v0.17.0 env creds
The roadmap section added hard-wrapped prose that failed the markdown
prose gate; reflowed (token-invariant). Also brought .env.example
current: cloud auth (now armed — needs SECRET or startup fails), routing
strict, the X engine (panel-entered OAuth), and web research.
* fix(ci): reduce cyclomatic complexity of five wave-3 blocks (xenon gate)
The wave-3 subagents introduced C-rank functions the CI xenon gate
rejects (my per-item reviews ran ruff/mypy/pytest but not xenon):
- sandbox.janitor_sweep -> extract _list_labeled_sandboxes /
_list_live_agent_containers / _prune_grace
- x_client.fetch_mentions -> extract _parse_mention_items
- x_engine.run_cycle -> extract _process_mentions
- orchestrator._dispatch_pm_work -> extract the source-skip into a
MODULE-level _is_held_ceo_source (module, not method, so the
wholesale-mocked dispatcher unit tests exercise the real logic)
- auth/seed.ensure_seed_user -> extract _apply_seed_updates (module avg -> A)
Behavior-preserving; full suite green (11902), xenon clean.
* fix(ci): declare pyjwt + fastapi-users-db-sqlalchemy as direct deps (deptry)
The cloud-auth code imports jwt and fastapi_users_db_sqlalchemy directly
but they were only transitive deps (via fastapi-users), which deptry
(quality gate, DEP003) rejects. Declared explicitly; deptry roboco/ clean.
Missed originally because local make quality stopped at earlier gates
before reaching deptry.
* feat(x): gate mention replies behind ROBOCO_X_REPLIES_ENABLED (default off)
Per CEO decision: the X engine should only post about releases by
default. Reading mentions needs a paid X API tier, so the mention-reply
half is now a deliberate opt-in on top of release posting.
New default-off flag x_replies_enabled gates the mentions poll loop
(_x_mentions_poll_loop) and XEngine.run_cycle; release-post drafting
(the release-proposal approve hook) is unaffected and still runs when
x_engine_enabled + credentials are set. Added to FEATURE_FLAGS + the
panel card. Tests: release posting works with replies off; run_cycle +
the poll loop are no-ops with replies off.
* fix: 401 only redirects to /login when cloud auth is on; panel-token strips .env quotes
Two bugs that together dead-ended login in secure mode:
- client.ts redirected to /login on ANY 401, so a mismatched panel
token (header-trust/secure mode, cloud auth off) bounced the user to a
login page whose backend route isn't mounted -> 404. Now it probes
/auth/status (bare fetch, no interceptor re-entry) and only redirects
when cloud_auth_enabled.
- make panel-token read the .env secret with grep|cut without stripping
surrounding quotes, so a quoted ROBOCO_AGENT_AUTH_SECRET produced a
token signed with the quotes included — which never verifies against
the orchestrator (docker-compose/pydantic unquote the secret). Now
strips surrounding single/double quotes.
* fix: git-log 500 on '|' in commit message; X queue shows an empty state
- GET /api/git/log 500'd (ValueError: Invalid isoformat) when a commit
SUBJECT contained a '|' (e.g. the 'curl|sh' lockdown commit): the
fixed '|' field delimiter let the subject's pipe shift the split so
author+date collapsed into one field. Switched to \x1f (Unit
Separator), which can't appear in commit content. Regression test with
a piped subject.
- The X Post Queue returned null when empty, so there was no visible
place for the X drafts. It now renders a discoverable empty state
pointing at Settings -> X credentials.
* docs: bring docs/rag + docs/map current for v0.17.0 (waves 1-3)
Agent-facing RAG corpus and codebase map updated for every feature in
the 0.17.0 span, code-verified:
- wave 3: sandbox DB, DB network isolation, cloud auth, X engine
(+ x_replies_enabled sub-flag), board roadmap engine — new RAG
architecture pages + role/tool/config-reference updates; new symbols,
migrations 057-059, panel surfaces, and the get_agent_context
dual-path across the map slices.
- waves 1-2: A2A live view + switchboard, prompter memory
(search_past_tasks), Secretary edit access + PM-lighter scope, the
PR-gate auto-submit turn cut (ROBOCO_PR_GATE_AUTO_SUBMIT_ENABLED).
- correctness fix: api-routes-schemas.md no longer claims the A2A admin
routes are reachable by any authenticated agent — they carry a
_require_ceo gate (wave 2c).
docs/internal, _front.md deltas, and the frozen _complete_map.md
snapshot untouched.
* fix(rag): atomic upsert for indexed-doc tracking (kills e2e segfault)
The indexed-document tracking write used check-then-insert in two paths
(IndexedDocumentRepository.upsert_batch and the file-source
_upsert_doc_record). Under concurrent indexing both callers saw no row
and both inserted, so the second violated uq_indexed_doc_source and
poisoned its transaction — surfacing in CI as the intermittent
_checkin_failed SIGSEGV on the failed connection's pool checkin.
Both paths now use INSERT ... ON CONFLICT DO UPDATE against the
constraint: coalesce keeps an existing title/preview when the new value
is empty (matching the old guards) and metadata is jsonb-merged. The
batch dedupes within itself first (ON CONFLICT can't touch a row twice
in one statement). expire_all after the Core upsert keeps same-session
ORM reads consistent with the merged DB row.
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
1334 lines
48 KiB
Python
1334 lines
48 KiB
Python
"""Unit tests for PrompterService.
|
|
|
|
Covers the live-intake draft → task flow (``create_task_from_draft`` /
|
|
``confirm_live_draft`` + the enum/priority/team coercion) and the pure
|
|
draft/description helpers. DB-backed tests use an in-memory async session via
|
|
conftest fixtures.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
from datetime import UTC, datetime, timedelta
|
|
from types import SimpleNamespace
|
|
from typing import TYPE_CHECKING, Any, cast
|
|
from uuid import UUID, uuid4
|
|
|
|
import pytest
|
|
|
|
if TYPE_CHECKING:
|
|
from sqlalchemy.ext.asyncio import AsyncSession
|
|
from roboco.db.tables import (
|
|
AgentTable,
|
|
ProductTable,
|
|
ProjectTable,
|
|
TaskTable,
|
|
)
|
|
from roboco.models.base import (
|
|
AgentRole,
|
|
AgentStatus,
|
|
Complexity,
|
|
TaskNature,
|
|
TaskStatus,
|
|
TaskType,
|
|
Team,
|
|
)
|
|
from roboco.seeds.initial_data import AGENT_UUIDS
|
|
from roboco.services import prompter as prompter_module
|
|
from roboco.services.base import ServiceError, ValidationError
|
|
from roboco.services.prompter import (
|
|
_HISTORY_DIGEST_PER_PROJECT_LIMIT,
|
|
_HISTORY_TITLE_EXCERPT_CAP,
|
|
PrompterService,
|
|
_cell_teams,
|
|
_clean_list,
|
|
_draft_cell_map,
|
|
_task_activity_date,
|
|
_title_excerpt,
|
|
build_history_digest,
|
|
compact_task_rows,
|
|
compose_description,
|
|
derive_scale,
|
|
get_prompter_service,
|
|
history_digest_layer,
|
|
parse_readiness,
|
|
)
|
|
|
|
# =============================================================================
|
|
# Pure function tests (no DB)
|
|
# =============================================================================
|
|
|
|
|
|
def test_parse_readiness_extracts_and_strips_tag() -> None:
|
|
content = (
|
|
"Here is my question about scope.\n\n"
|
|
'```roboco-meta\n{"covered": ["objective", "scope"], '
|
|
'"ready": true, "scale": "multi"}\n```'
|
|
)
|
|
clean, tag = parse_readiness(content)
|
|
assert clean == "Here is my question about scope."
|
|
assert tag is not None
|
|
assert tag.ready is True
|
|
assert tag.scale == "multi"
|
|
assert tag.covered == ["objective", "scope"]
|
|
# The control block must not leak into the user-visible text.
|
|
assert "roboco-meta" not in clean
|
|
|
|
|
|
def test_parse_readiness_absent_block_is_not_ready() -> None:
|
|
clean, tag = parse_readiness("Just a plain reply, no control block.")
|
|
assert clean == "Just a plain reply, no control block."
|
|
assert tag is None
|
|
|
|
|
|
def test_parse_readiness_malformed_json_is_graceful() -> None:
|
|
content = "Reply text.\n```roboco-meta\n{not valid json]\n```"
|
|
clean, tag = parse_readiness(content)
|
|
assert "roboco-meta" not in clean
|
|
assert clean == "Reply text."
|
|
assert tag is None
|
|
|
|
|
|
def test_parse_readiness_uses_last_block() -> None:
|
|
content = (
|
|
'```roboco-meta\n{"ready": false, "scale": "single"}\n```\n'
|
|
"Final answer.\n"
|
|
'```roboco-meta\n{"ready": true, "scale": "multi"}\n```'
|
|
)
|
|
clean, tag = parse_readiness(content)
|
|
assert tag is not None
|
|
assert tag.ready is True
|
|
assert tag.scale == "multi"
|
|
assert "roboco-meta" not in clean
|
|
|
|
|
|
def test_derive_scale_single_vs_multi() -> None:
|
|
assert derive_scale([{"team": "backend"}]) == "single"
|
|
assert derive_scale([{"team": "backend"}, {"team": "frontend"}]) == "multi"
|
|
# Non-cell teams (e.g. main_pm) do not count toward cell breadth.
|
|
assert derive_scale([{"team": "backend"}, {"team": "main_pm"}]) == "single"
|
|
assert derive_scale([]) == "single"
|
|
|
|
|
|
# -----------------------------------------------------------------------------
|
|
# the_work shape tolerance — the intake agent is an LLM and sometimes emits
|
|
# the_work as a list of bare team-name strings ("backend") instead of the
|
|
# documented {team, summary, items} objects. Every consumer must tolerate that
|
|
# without raising (regression: preview-batch used to 500 with
|
|
# "'str' object has no attribute 'get'").
|
|
# -----------------------------------------------------------------------------
|
|
|
|
|
|
def test_cell_teams_tolerates_bare_string_entries() -> None:
|
|
# The LLM emitted the_work as a list of team names, not objects.
|
|
assert _cell_teams(["backend", "frontend", "backend"]) == ["backend", "frontend"]
|
|
# A bare string that isn't a cell is skipped, just like a non-cell dict.
|
|
assert _cell_teams(["backend", "main_pm"]) == ["backend"]
|
|
assert _cell_teams(["nonsense"]) == []
|
|
|
|
|
|
def test_lead_cell_team_tolerates_bare_string_entries() -> None:
|
|
draft = {"the_work": ["frontend", "backend"]}
|
|
assert PrompterService._lead_cell_team(draft, default=Team.BACKEND) is Team.FRONTEND
|
|
# First valid cell wins; an invalid bare string is skipped.
|
|
draft = {"the_work": ["nonsense", "ux_ui"]}
|
|
assert PrompterService._lead_cell_team(draft, default=Team.BACKEND) is Team.UX_UI
|
|
|
|
|
|
def test_derive_scale_tolerates_bare_string_entries() -> None:
|
|
assert derive_scale(["backend"]) == "single"
|
|
assert derive_scale(["backend", "frontend"]) == "multi"
|
|
|
|
|
|
def test_compose_description_renders_bare_string_work_entries() -> None:
|
|
draft = {
|
|
"objective": "Fix the intake batch preview.",
|
|
"the_work": ["backend", "frontend"],
|
|
"acceptance_criteria": ["Preview no longer 500s"],
|
|
}
|
|
md = compose_description(draft)
|
|
# Each bare string renders as a cell heading; multi-cell gets the board-led line.
|
|
assert "## The Work" in md
|
|
assert "**Backend**" in md
|
|
assert "**Frontend**" in md
|
|
assert "Board-led" in md
|
|
|
|
|
|
def test_compose_description_single_cell_markdown() -> None:
|
|
draft = {
|
|
"objective": "Let humans track token usage.",
|
|
"what_this_builds": ["A usage panel on the Metrics page"],
|
|
"the_work": [
|
|
{
|
|
"team": "frontend",
|
|
"summary": "Render the usage panel",
|
|
"items": ["Add the chart", "Wire the API"],
|
|
}
|
|
],
|
|
"notes": ["Reuse the existing Metrics layout"],
|
|
"acceptance_criteria": ["Panel shows totals", "Panel filters by range"],
|
|
}
|
|
md = compose_description(draft)
|
|
assert "## Objective" in md
|
|
assert "## What This Builds" in md
|
|
assert "## The Work" in md
|
|
assert "**Frontend** — Render the usage panel" in md
|
|
assert "## Notes" in md
|
|
assert "## Success Criteria" in md
|
|
assert "- Panel shows totals" in md
|
|
# Single-cell tasks get no board-led lead line.
|
|
assert "Board-led" not in md
|
|
|
|
|
|
def test_compose_description_multi_cell_has_board_led_lead() -> None:
|
|
draft = {
|
|
"objective": "Ship the Prompter.",
|
|
"the_work": [
|
|
{"team": "backend", "summary": "Chat endpoint", "items": []},
|
|
{"team": "frontend", "summary": "Chat UI", "items": []},
|
|
{"team": "ux_ui", "summary": "Interaction design", "items": []},
|
|
],
|
|
"acceptance_criteria": ["It works end to end"],
|
|
}
|
|
md = compose_description(draft)
|
|
assert "Board-led" in md
|
|
assert "**Backend**" in md
|
|
assert "**UX/UI**" in md
|
|
|
|
|
|
def test_compose_description_falls_back_to_provided_description() -> None:
|
|
# Sparse structured fields → fall back to a model-provided description.
|
|
draft = {"description": "A perfectly adequate fallback description here."}
|
|
md = compose_description(draft)
|
|
assert md == "A perfectly adequate fallback description here."
|
|
|
|
|
|
def test_lead_cell_team_prefers_the_work_cell() -> None:
|
|
draft = {"the_work": [{"team": "frontend"}], "team": "backend"}
|
|
assert PrompterService._lead_cell_team(draft, default=Team.BACKEND) is Team.FRONTEND
|
|
# Empty the_work falls back to the provided default.
|
|
assert PrompterService._lead_cell_team({}, default=Team.BACKEND) is Team.BACKEND
|
|
|
|
|
|
def test_lead_cell_team_skips_invalid_cell_names() -> None:
|
|
# An off-enum cell name is skipped, not raised on; falls through to a valid one.
|
|
draft = {"the_work": [{"team": "nonsense"}, {"team": "frontend"}]}
|
|
assert PrompterService._lead_cell_team(draft, default=Team.BACKEND) is Team.FRONTEND
|
|
|
|
|
|
def test_coerce_draft_enums_defaults_invalid_values() -> None:
|
|
# Regression: the LLM emits off-enum values (e.g. task_type="feature"). The
|
|
# confirm must coerce to defaults, never raise — a bad enum guess must not
|
|
# 400 the launch and force the agent to self-correct in-chat.
|
|
draft = {
|
|
"team": "backend",
|
|
"task_type": "feature", # not a valid TaskType
|
|
"nature": "bogus", # not a valid TaskNature
|
|
"estimated_complexity": "enormous", # not a valid Complexity
|
|
}
|
|
team, task_type, nature, complexity = PrompterService._coerce_draft_enums(draft)
|
|
assert team is Team.BACKEND
|
|
assert task_type is TaskType.CODE
|
|
assert nature is TaskNature.TECHNICAL
|
|
assert complexity is Complexity.MEDIUM
|
|
|
|
|
|
def test_coerce_priority_maps_words_clamps_and_defaults() -> None:
|
|
# Regression: priority is the one non-enum field the agent guesses, and it
|
|
# guesses a word ("high") as often as a number — int("high") used to 500.
|
|
# word/number -> expected priority int (0=urgent .. 3=low).
|
|
cases: dict[object, int] = {
|
|
"urgent": 0,
|
|
"high": 1,
|
|
"medium": 2,
|
|
"low": 3,
|
|
1: 1,
|
|
"3": 3,
|
|
99: 3, # clamped into range
|
|
"nonsense": 2, # unrecognized -> default medium
|
|
None: 2, # missing -> default medium
|
|
}
|
|
for value, expected in cases.items():
|
|
assert PrompterService._coerce_priority(value) == expected
|
|
|
|
|
|
def test_coerce_draft_enums_keeps_valid_and_derives_missing_team() -> None:
|
|
# Valid values pass through; a missing team is derived from the_work.
|
|
draft = {
|
|
"task_type": "documentation",
|
|
"nature": "technical",
|
|
"estimated_complexity": "medium",
|
|
"the_work": [{"team": "frontend"}],
|
|
}
|
|
team, task_type, nature, complexity = PrompterService._coerce_draft_enums(draft)
|
|
assert team is Team.FRONTEND
|
|
assert task_type is TaskType.DOCUMENTATION
|
|
assert nature is TaskNature.TECHNICAL
|
|
assert complexity is Complexity.MEDIUM
|
|
|
|
|
|
# =============================================================================
|
|
# Factory
|
|
# =============================================================================
|
|
|
|
|
|
def test_get_prompter_service_no_db() -> None:
|
|
service = get_prompter_service()
|
|
assert isinstance(service, PrompterService)
|
|
assert service._db is None
|
|
|
|
|
|
def test_get_prompter_service_raises_without_db_for_session_methods() -> None:
|
|
service = get_prompter_service()
|
|
with pytest.raises(ServiceError, match="DB session"):
|
|
_ = service._session
|
|
|
|
|
|
# =============================================================================
|
|
# DB-backed: assignee routing + confirm_live_draft
|
|
# =============================================================================
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_assignee_is_board_distinguishes_roles(db_session: Any) -> None:
|
|
"""Drives product team routing: a board reviewer keeps the root on the board.
|
|
|
|
A product confirmed via "Board review & Start" is assigned to a board
|
|
reviewer and must stay team=board so the CEO's Approve & Start gate appears;
|
|
one assigned to main-pm (or a cell dev) is not a board task.
|
|
"""
|
|
service = get_prompter_service(db=db_session)
|
|
|
|
def _agent(role: AgentRole) -> AgentTable:
|
|
return AgentTable(
|
|
id=uuid4(),
|
|
name="A",
|
|
slug=f"a-{uuid4().hex[:8]}",
|
|
role=role,
|
|
team=None,
|
|
status=AgentStatus.ACTIVE,
|
|
model_config={},
|
|
system_prompt="x",
|
|
capabilities=[],
|
|
permissions={},
|
|
metrics={},
|
|
)
|
|
|
|
po = _agent(AgentRole.PRODUCT_OWNER)
|
|
hom = _agent(AgentRole.HEAD_MARKETING)
|
|
dev = _agent(AgentRole.DEVELOPER)
|
|
db_session.add_all([po, hom, dev])
|
|
await db_session.flush()
|
|
|
|
assert await service._assignee_is_board(cast("UUID", po.id)) is True
|
|
assert await service._assignee_is_board(cast("UUID", hom.id)) is True
|
|
assert await service._assignee_is_board(cast("UUID", dev.id)) is False
|
|
# Unknown id is not a board agent — defensive, must not raise.
|
|
assert await service._assignee_is_board(uuid4()) is False
|
|
|
|
|
|
async def _seed_project_and_ceo(db_session: Any) -> tuple[UUID, UUID]:
|
|
"""Seed a system agent + project + CEO; return (project_id, ceo_id).
|
|
|
|
Returns plain ``UUID``s (not the ORM rows) so callers pass real uuids to the
|
|
service — no casting the ORM ``.id`` column type at the call site.
|
|
"""
|
|
system_id, project_id, ceo_id = uuid4(), uuid4(), uuid4()
|
|
system = AgentTable(
|
|
id=system_id,
|
|
name="System",
|
|
slug=f"system-{uuid4().hex[:8]}",
|
|
role=AgentRole.SYSTEM,
|
|
team=None,
|
|
status=AgentStatus.ACTIVE,
|
|
model_config={},
|
|
system_prompt="system",
|
|
capabilities=[],
|
|
permissions={},
|
|
metrics={},
|
|
)
|
|
db_session.add(system)
|
|
await db_session.flush()
|
|
project = ProjectTable(
|
|
id=project_id,
|
|
name="Intake Test Project",
|
|
slug=f"intake-{uuid4().hex[:8]}",
|
|
git_url="https://github.com/example/intake.git",
|
|
default_branch="main",
|
|
protected_branches=["main"],
|
|
assigned_cell=Team.BACKEND,
|
|
created_by=system_id,
|
|
is_active=True,
|
|
)
|
|
ceo = AgentTable(
|
|
id=ceo_id,
|
|
name="CEO",
|
|
slug=f"ceo-{uuid4().hex[:8]}",
|
|
role=AgentRole.CEO,
|
|
team=None,
|
|
status=AgentStatus.ACTIVE,
|
|
model_config={},
|
|
system_prompt="ceo",
|
|
capabilities=[],
|
|
permissions={},
|
|
metrics={},
|
|
)
|
|
db_session.add_all([project, ceo])
|
|
await db_session.flush()
|
|
# The "& Start" routes assign the draft to a fixed board/PM agent
|
|
# (product-owner for "Board review", main-pm for "Approve & Start"); those
|
|
# rows must exist for the assigned_to FK. merge() is idempotent, so this is
|
|
# safe whether or not another test already committed them on the shared DB.
|
|
for slug, role, team in (
|
|
("product-owner", AgentRole.PRODUCT_OWNER, None),
|
|
("main-pm", AgentRole.MAIN_PM, Team.MAIN_PM),
|
|
):
|
|
await db_session.merge(
|
|
AgentTable(
|
|
id=UUID(AGENT_UUIDS[slug]),
|
|
name=slug,
|
|
slug=slug,
|
|
role=role,
|
|
team=team,
|
|
status=AgentStatus.ACTIVE,
|
|
model_config={},
|
|
system_prompt=slug,
|
|
capabilities=[],
|
|
permissions={},
|
|
metrics={},
|
|
)
|
|
)
|
|
await db_session.flush()
|
|
return project_id, ceo_id
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_confirm_live_draft_board_route_assigns_po(db_session: Any) -> None:
|
|
""" "Board review & Start" (default route) → PENDING, assigned to the Product
|
|
Owner so the orchestrator fires the PO + HoM review."""
|
|
project_id, ceo_id = await _seed_project_and_ceo(db_session)
|
|
service = get_prompter_service(db=db_session)
|
|
|
|
draft = {
|
|
"title": "Add token metrics",
|
|
"objective": "See token usage at a glance.",
|
|
"acceptance_criteria": ["Dashboard shows total tokens"],
|
|
"team": "backend",
|
|
"the_work": [
|
|
{"team": "backend", "summary": "instrument", "items": ["count tokens"]}
|
|
],
|
|
}
|
|
task_id = await service.confirm_live_draft(draft, ceo_id, project_id=project_id)
|
|
|
|
row = await db_session.get(TaskTable, task_id)
|
|
assert row is not None
|
|
assert row.status == TaskStatus.PENDING # "& Start" — started now
|
|
assert row.assigned_to == UUID(AGENT_UUIDS["product-owner"]) # board review
|
|
assert row.source == "prompter"
|
|
assert row.confirmed_by_human is True
|
|
assert row.team == Team.BACKEND # lead cell from the_work
|
|
assert row.created_by == ceo_id
|
|
assert row.nature is not None and row.task_type is not None
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_confirm_live_draft_main_pm_route_assigns_main_pm(
|
|
db_session: Any,
|
|
) -> None:
|
|
""" "Approve & Start" (route="main_pm") → PENDING, assigned to the Main PM."""
|
|
project_id, ceo_id = await _seed_project_and_ceo(db_session)
|
|
service = get_prompter_service(db=db_session)
|
|
draft = {
|
|
"title": "Quick fix",
|
|
"acceptance_criteria": ["done"],
|
|
"team": "backend",
|
|
}
|
|
task_id = await service.confirm_live_draft(
|
|
draft, ceo_id, project_id=project_id, route="main_pm"
|
|
)
|
|
row = await db_session.get(TaskTable, task_id)
|
|
assert row.status == TaskStatus.PENDING
|
|
assert row.assigned_to == UUID(AGENT_UUIDS["main-pm"])
|
|
# A PM coordinates — a code task handed to the Main PM is coerced to
|
|
# planning (the PM/code invariant; the draft's team=backend is honored but
|
|
# the type is retyped so the combo never persists).
|
|
assert row.task_type == TaskType.PLANNING
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_confirm_live_draft_product_routes_to_main_pm(db_session: Any) -> None:
|
|
"""A product-scoped draft via the "Approve & Start" path is a Main-PM root.
|
|
|
|
The board path (the ``route="board"`` default) keeps the root at
|
|
``team=board`` until the CEO approves; the Main-PM path is selected
|
|
explicitly with ``route="main_pm"``.
|
|
"""
|
|
_project_id, ceo_id = await _seed_project_and_ceo(db_session)
|
|
product_id = uuid4()
|
|
product = ProductTable(
|
|
id=product_id,
|
|
name="Intake Product",
|
|
slug=f"prod-{uuid4().hex[:8]}",
|
|
description="x",
|
|
created_by=ceo_id,
|
|
)
|
|
db_session.add(product)
|
|
await db_session.flush()
|
|
|
|
service = get_prompter_service(db=db_session)
|
|
draft = {
|
|
"title": "Board-led feature",
|
|
"acceptance_criteria": ["works end to end"],
|
|
"team": "backend",
|
|
}
|
|
task_id = await service.confirm_live_draft(
|
|
draft, ceo_id, product_id=product_id, route="main_pm"
|
|
)
|
|
row = await db_session.get(TaskTable, task_id)
|
|
assert row.team == Team.MAIN_PM
|
|
assert row.product_id == product_id
|
|
assert row.project_id is None
|
|
# A Main-PM coordination root is never code — intake coerces code->planning
|
|
# so main_pm + code can never coexist (the 2026-06-27 meltdown shape).
|
|
assert row.task_type == TaskType.PLANNING
|
|
|
|
|
|
# =============================================================================
|
|
# MegaTask: confirm_live_batch (umbrella + sequenced root-subtasks)
|
|
# =============================================================================
|
|
|
|
|
|
async def _seed_second_project(db_session: Any, ceo_id: UUID) -> UUID:
|
|
"""Seed a second project so a MegaTask can span multiple repos."""
|
|
project_id = uuid4()
|
|
db_session.add(
|
|
ProjectTable(
|
|
id=project_id,
|
|
name="Intake Test Project 2",
|
|
slug=f"intake2-{uuid4().hex[:8]}",
|
|
git_url="https://github.com/example/intake2.git",
|
|
default_branch="main",
|
|
protected_branches=["main"],
|
|
assigned_cell=Team.FRONTEND,
|
|
created_by=ceo_id,
|
|
is_active=True,
|
|
)
|
|
)
|
|
await db_session.flush()
|
|
return project_id
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_confirm_live_batch_builds_umbrella_and_sequenced_subtasks(
|
|
db_session: Any,
|
|
) -> None:
|
|
"""A MegaTask creates one branchless umbrella + N root-subtasks across many
|
|
projects, with the collision-derived dependency edges wired so the
|
|
dependency-gate runs the waves in order."""
|
|
project1, ceo_id = await _seed_project_and_ceo(db_session)
|
|
project2 = await _seed_second_project(db_session, ceo_id)
|
|
service = get_prompter_service(db=db_session)
|
|
|
|
# A & B both add a migration → serial chain A→B (the migration rule orders
|
|
# them by priority then index). C is an independent frontend task in another
|
|
# project, so it runs in parallel with A in wave 0.
|
|
drafts: list[dict[str, Any]] = [
|
|
{
|
|
"title": "A: add table",
|
|
"acceptance_criteria": ["a"],
|
|
"team": "backend",
|
|
"project_id": str(project1),
|
|
"intends_to_touch": ["roboco/services/foo.py"],
|
|
"adds_migration": True,
|
|
},
|
|
{
|
|
"title": "B: extend table",
|
|
"acceptance_criteria": ["b"],
|
|
"team": "backend",
|
|
"project_id": str(project1),
|
|
"intends_to_touch": ["roboco/services/bar.py"],
|
|
"adds_migration": True,
|
|
},
|
|
{
|
|
"title": "C: frontend widget",
|
|
"acceptance_criteria": ["c"],
|
|
"team": "frontend",
|
|
"project_id": str(project2),
|
|
"intends_to_touch": ["panel/src/widget.tsx"],
|
|
},
|
|
]
|
|
result = await service.confirm_live_batch(
|
|
"Three things",
|
|
drafts,
|
|
ceo_id,
|
|
project_ids=[project1, project2],
|
|
route="main_pm",
|
|
)
|
|
|
|
# A (migration) and C (independent) run in wave 0; B chains after A.
|
|
assert result["waves"] == [[0, 2], [1]]
|
|
ids = result["root_subtask_ids"]
|
|
assert len(ids) == len(drafts)
|
|
|
|
umbrella_id = UUID(result["umbrella_task_id"])
|
|
umbrella = await db_session.get(TaskTable, umbrella_id)
|
|
assert umbrella.batch_id is not None
|
|
assert umbrella.parent_task_id is None
|
|
assert umbrella.project_id is None and umbrella.product_id is None
|
|
assert umbrella.team == Team.MAIN_PM
|
|
assert umbrella.status == TaskStatus.PENDING
|
|
assert umbrella.branch_name is None # branchless
|
|
# A Main-PM coordination root is never code — the umbrella is planning-typed.
|
|
assert umbrella.task_type == TaskType.PLANNING
|
|
|
|
a, b, c = [await db_session.get(TaskTable, UUID(sid)) for sid in ids]
|
|
for sub in (a, b, c):
|
|
assert sub.parent_task_id == umbrella_id
|
|
assert sub.batch_id == umbrella.batch_id
|
|
assert sub.team == Team.MAIN_PM
|
|
assert sub.status == TaskStatus.PENDING
|
|
# Each root-subtask is a Main-PM coordination root: code->planning coerced
|
|
# at intake so main_pm + code can never coexist (the 2026-06-27 meltdown
|
|
# shape). It still gets its own branch + PR + submit_root + pr_review gate
|
|
# — the gate is branch-keyed, not task_type-keyed.
|
|
assert sub.task_type == TaskType.PLANNING
|
|
assert a.project_id == project1
|
|
assert b.project_id == project1
|
|
assert c.project_id == project2
|
|
# sequence = wave index: A and C in wave 0, B in wave 1.
|
|
assert (a.sequence, b.sequence, c.sequence) == (0, 1, 0)
|
|
# Dependency wiring: B waits on A; C is independent.
|
|
assert UUID(ids[0]) in b.dependency_ids
|
|
assert c.dependency_ids == []
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_confirm_live_batch_board_route_holds_subtasks_in_backlog(
|
|
db_session: Any,
|
|
) -> None:
|
|
"""The "board" route sends the umbrella to the Product Owner for batch review
|
|
and holds the root-subtasks in BACKLOG until the umbrella is approved."""
|
|
project1, ceo_id = await _seed_project_and_ceo(db_session)
|
|
project2 = await _seed_second_project(db_session, ceo_id)
|
|
service = get_prompter_service(db=db_session)
|
|
drafts = [
|
|
{
|
|
"title": "One",
|
|
"acceptance_criteria": ["x"],
|
|
"team": "backend",
|
|
"project_id": str(project1),
|
|
},
|
|
{
|
|
"title": "Two",
|
|
"acceptance_criteria": ["y"],
|
|
"team": "frontend",
|
|
"project_id": str(project2),
|
|
},
|
|
]
|
|
result = await service.confirm_live_batch(
|
|
"Two repos", drafts, ceo_id, project_ids=[project1, project2], route="board"
|
|
)
|
|
|
|
umbrella = await db_session.get(TaskTable, UUID(result["umbrella_task_id"]))
|
|
assert umbrella.team == Team.BOARD
|
|
assert umbrella.assigned_to == UUID(AGENT_UUIDS["product-owner"])
|
|
assert umbrella.status == TaskStatus.PENDING
|
|
sub = await db_session.get(TaskTable, UUID(result["root_subtask_ids"][0]))
|
|
assert sub.status == TaskStatus.BACKLOG # held until batch review approves
|
|
assert sub.team == Team.BOARD
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_confirm_live_batch_rejects_empty(db_session: Any) -> None:
|
|
_project1, ceo_id = await _seed_project_and_ceo(db_session)
|
|
service = get_prompter_service(db=db_session)
|
|
with pytest.raises(ValidationError):
|
|
await service.confirm_live_batch(
|
|
"Empty", [], ceo_id, project_ids=[uuid4(), uuid4()]
|
|
)
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_confirm_live_batch_rejects_draft_outside_scope(db_session: Any) -> None:
|
|
"""A draft targeting a project NOT in the scoped project_ids is refused — the
|
|
intake agent only read the scoped repos."""
|
|
project1, ceo_id = await _seed_project_and_ceo(db_session)
|
|
project2 = await _seed_second_project(db_session, ceo_id)
|
|
service = get_prompter_service(db=db_session)
|
|
outside = uuid4() # never in scope
|
|
drafts = [
|
|
{"title": "A", "acceptance_criteria": ["a"], "project_id": str(project1)},
|
|
{"title": "B", "acceptance_criteria": ["b"], "project_id": str(outside)},
|
|
]
|
|
with pytest.raises(ValidationError, match="outside this MegaTask"):
|
|
await service.confirm_live_batch(
|
|
"Scoped", drafts, ceo_id, project_ids=[project1, project2], route="main_pm"
|
|
)
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_confirm_live_batch_rejects_single_project(db_session: Any) -> None:
|
|
"""A degenerate batch whose drafts all target one project is not a MegaTask."""
|
|
project1, ceo_id = await _seed_project_and_ceo(db_session)
|
|
project2 = await _seed_second_project(db_session, ceo_id)
|
|
service = get_prompter_service(db=db_session)
|
|
drafts = [
|
|
{"title": "A", "acceptance_criteria": ["a"], "project_id": str(project1)},
|
|
{"title": "B", "acceptance_criteria": ["b"], "project_id": str(project1)},
|
|
]
|
|
with pytest.raises(ValidationError, match="at least two distinct projects"):
|
|
await service.confirm_live_batch(
|
|
"One repo",
|
|
drafts,
|
|
ceo_id,
|
|
project_ids=[project1, project2],
|
|
route="main_pm",
|
|
)
|
|
|
|
|
|
def test_preview_batch_computes_waves_without_creating() -> None:
|
|
"""preview_batch is pure: it returns the same waves confirm would wire, with
|
|
no DB session and no task creation."""
|
|
service = get_prompter_service() # no db — pure compute
|
|
drafts: list[dict[str, Any]] = [
|
|
{"title": "A", "adds_migration": True, "intends_to_touch": ["a.py"]},
|
|
{"title": "B", "adds_migration": True, "intends_to_touch": ["b.py"]},
|
|
{"title": "C", "intends_to_touch": ["c.py"]},
|
|
]
|
|
result = service.preview_batch(drafts)
|
|
# A & B chain on the migration rule; C is independent → [[0, 2], [1]].
|
|
assert result["waves"] == [[0, 2], [1]]
|
|
assert isinstance(result["warnings"], list)
|
|
|
|
|
|
def test_preview_batch_honours_declared_depends_on() -> None:
|
|
"""B1b: a draft's declared depends_on becomes a real edge even when the
|
|
collision surfaces are disjoint (the live S6 break: declared waves were
|
|
dropped because the intends_to_touch globs didn't overlap)."""
|
|
service = get_prompter_service()
|
|
drafts: list[dict[str, Any]] = [
|
|
{"title": "A", "intends_to_touch": ["a.py"]},
|
|
{"title": "B", "intends_to_touch": ["b.py"], "depends_on": [0]},
|
|
]
|
|
result = service.preview_batch(drafts)
|
|
assert result["waves"] == [[0], [1]]
|
|
|
|
|
|
def test_preview_batch_coerces_string_declared_indices() -> None:
|
|
"""The LLM sometimes emits depends_on indices as strings ("0")."""
|
|
service = get_prompter_service()
|
|
drafts: list[dict[str, Any]] = [
|
|
{"title": "A", "intends_to_touch": ["a.py"]},
|
|
{"title": "B", "intends_to_touch": ["b.py"], "depends_on": ["0"]},
|
|
]
|
|
result = service.preview_batch(drafts)
|
|
assert result["waves"] == [[0], [1]]
|
|
|
|
|
|
def test_preview_batch_rejects_out_of_range_declared_dep() -> None:
|
|
service = get_prompter_service()
|
|
drafts: list[dict[str, Any]] = [
|
|
{"title": "A", "intends_to_touch": ["a.py"], "depends_on": [9]},
|
|
]
|
|
with pytest.raises(ValidationError):
|
|
service.preview_batch(drafts)
|
|
|
|
|
|
def test_preview_batch_rejects_empty() -> None:
|
|
service = get_prompter_service()
|
|
with pytest.raises(ValidationError):
|
|
service.preview_batch([])
|
|
|
|
|
|
def test_preview_batch_tolerates_bare_string_the_work() -> None:
|
|
"""Regression: the LLM sometimes emits the_work as bare team-name strings.
|
|
preview_batch must not 500 on that shape (it did: 'str' has no 'get')."""
|
|
service = get_prompter_service()
|
|
drafts: list[dict[str, Any]] = [
|
|
{
|
|
"title": "A",
|
|
"project_id": str(uuid4()),
|
|
"the_work": ["backend"],
|
|
"intends_to_touch": ["a.py"],
|
|
},
|
|
{
|
|
"title": "B",
|
|
"project_id": str(uuid4()),
|
|
"the_work": ["backend", "frontend"],
|
|
"intends_to_touch": ["b.py"],
|
|
},
|
|
]
|
|
result = service.preview_batch(drafts)
|
|
assert isinstance(result["waves"], list)
|
|
assert isinstance(result["warnings"], list)
|
|
|
|
|
|
# =============================================================================
|
|
# Per-cell project map (multi-cell MegaTask root-subtask seam) — pure helpers
|
|
# =============================================================================
|
|
|
|
|
|
def _work(team: str, project_id: UUID | None) -> dict[str, Any]:
|
|
entry: dict[str, Any] = {"team": team, "summary": "s", "items": ["x"]}
|
|
if project_id is not None:
|
|
entry["project_id"] = str(project_id)
|
|
return entry
|
|
|
|
|
|
def test_draft_cell_map_collects_per_cell_projects_in_order() -> None:
|
|
"""A multi-cell draft yields one (team, project_id) per the_work entry,
|
|
in the_work order, de-duped by team."""
|
|
be_proj, fe_proj = uuid4(), uuid4()
|
|
draft = {
|
|
"the_work": [
|
|
_work("backend", be_proj),
|
|
_work("frontend", fe_proj),
|
|
]
|
|
}
|
|
assert _draft_cell_map(draft) == [(Team.BACKEND, be_proj), (Team.FRONTEND, fe_proj)]
|
|
|
|
|
|
def test_draft_cell_map_dedupes_repeated_team_keeping_first() -> None:
|
|
"""Two entries for the same cell (LLM noise) keep the first mapping — a
|
|
task_cell_projects row is unique per (task, team)."""
|
|
first, second = uuid4(), uuid4()
|
|
draft = {
|
|
"the_work": [
|
|
_work("backend", first),
|
|
_work("backend", second),
|
|
]
|
|
}
|
|
assert _draft_cell_map(draft) == [(Team.BACKEND, first)]
|
|
|
|
|
|
def test_draft_cell_map_skips_entries_without_project_id() -> None:
|
|
"""An entry with no project_id (single-cell legacy or a bare team string) is
|
|
skipped — the draft then falls back to its top-level project_id."""
|
|
be_proj = uuid4()
|
|
draft = {
|
|
"the_work": [
|
|
_work("backend", be_proj),
|
|
{"team": "frontend", "summary": "s", "items": []}, # no project_id
|
|
]
|
|
}
|
|
assert _draft_cell_map(draft) == [(Team.BACKEND, be_proj)]
|
|
|
|
|
|
def test_draft_cell_map_empty_when_no_entry_has_project_id() -> None:
|
|
"""A legacy single-cell draft (top-level project_id, bare-string the_work)
|
|
yields an empty map — the caller falls back to the top-level project_id."""
|
|
assert _draft_cell_map({"the_work": ["backend", "frontend"]}) == []
|
|
assert _draft_cell_map({"the_work": [{"team": "backend"}]}) == []
|
|
|
|
|
|
def test_draft_cell_map_skips_off_enum_teams_but_rejects_bad_uuids() -> None:
|
|
"""Off-enum team names are skipped (the intake agent is an LLM and can emit
|
|
a non-cell team), and an entry with no project_id is skipped (legacy
|
|
single-cell). But a present-but-malformed project_id is a hard error —
|
|
silently dropping it would collapse a 2-cell map to 1-cell and mis-route the
|
|
draft as a single-project task (#58)."""
|
|
good = uuid4()
|
|
draft = {
|
|
"the_work": [
|
|
_work("backend", good),
|
|
{"team": "marketing", "project_id": str(uuid4())}, # not a cell
|
|
_work("frontend", None), # missing project_id — skipped
|
|
]
|
|
}
|
|
assert _draft_cell_map(draft) == [(Team.BACKEND, good)]
|
|
|
|
bad = {
|
|
"the_work": [
|
|
_work("backend", good),
|
|
{"team": "ux_ui", "project_id": "not-a-uuid"}, # malformed — reject
|
|
]
|
|
}
|
|
with pytest.raises(ValidationError, match="Invalid project_id"):
|
|
_draft_cell_map(bad)
|
|
|
|
|
|
def test_validate_batch_scope_accepts_single_multi_cell_draft() -> None:
|
|
"""One 2-cell draft already spans ≥2 distinct projects → valid MegaTask."""
|
|
be_proj, fe_proj = uuid4(), uuid4()
|
|
drafts = [
|
|
{
|
|
"title": "S1",
|
|
"acceptance_criteria": ["a"],
|
|
"the_work": [
|
|
_work("backend", be_proj),
|
|
_work("frontend", fe_proj),
|
|
],
|
|
}
|
|
]
|
|
# Must not raise: 2 distinct projects across the one draft's cells.
|
|
PrompterService._validate_batch_scope(drafts, [be_proj, fe_proj])
|
|
|
|
|
|
def test_validate_batch_scope_rejects_out_of_scope_per_cell_project() -> None:
|
|
"""A per-cell project_id outside the scoped set is refused."""
|
|
in_scope, out_of_scope = uuid4(), uuid4()
|
|
drafts = [
|
|
{
|
|
"title": "S1",
|
|
"acceptance_criteria": ["a"],
|
|
"the_work": [
|
|
_work("backend", in_scope),
|
|
_work("frontend", out_of_scope),
|
|
],
|
|
}
|
|
]
|
|
with pytest.raises(ValidationError, match="outside this MegaTask"):
|
|
PrompterService._validate_batch_scope(drafts, [in_scope, uuid4()])
|
|
|
|
|
|
def test_validate_batch_scope_rejects_draft_with_no_project() -> None:
|
|
"""A draft with neither a per-cell map nor a top-level project_id is refused."""
|
|
drafts = [
|
|
{
|
|
"title": "S1",
|
|
"acceptance_criteria": ["a"],
|
|
"the_work": [_work("backend", None), _work("frontend", None)],
|
|
}
|
|
]
|
|
with pytest.raises(ValidationError, match="has no project"):
|
|
PrompterService._validate_batch_scope(drafts, [uuid4(), uuid4()])
|
|
|
|
|
|
def test_validate_batch_scope_distinct_count_spans_all_cells() -> None:
|
|
"""The ≥2 minimum counts distinct projects across ALL drafts' cells, not per
|
|
draft. Two single-cell drafts on the same project still fail (degenerate)."""
|
|
only = uuid4()
|
|
drafts = [
|
|
{
|
|
"title": "A",
|
|
"acceptance_criteria": ["a"],
|
|
"the_work": [_work("backend", only)],
|
|
},
|
|
{
|
|
"title": "B",
|
|
"acceptance_criteria": ["b"],
|
|
"the_work": [_work("frontend", only)], # same project, different cell
|
|
},
|
|
]
|
|
with pytest.raises(ValidationError, match="at least two distinct projects"):
|
|
PrompterService._validate_batch_scope(drafts, [only, uuid4()])
|
|
|
|
|
|
def test_validate_batch_scope_legacy_single_cell_drafts_still_work() -> None:
|
|
"""Back-compat: drafts using a top-level project_id (no the_work map) still
|
|
validate against the scope and the ≥2 distinct minimum."""
|
|
p1, p2 = uuid4(), uuid4()
|
|
drafts = [
|
|
{"title": "A", "acceptance_criteria": ["a"], "project_id": str(p1)},
|
|
{"title": "B", "acceptance_criteria": ["b"], "project_id": str(p2)},
|
|
]
|
|
PrompterService._validate_batch_scope(drafts, [p1, p2])
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_resolve_owning_team_multi_cell_map_routes_to_main_pm() -> None:
|
|
"""A multi-cell ad-hoc map is a coordination root (mirrors a product root), so
|
|
it routes to the Main PM — never the lead cell (a cell PM can't delegate
|
|
cross-cell; that would deadlock the fan-out). No DB access on this branch."""
|
|
service = get_prompter_service() # no db — the cell-map branch never reads it
|
|
be_proj, fe_proj = uuid4(), uuid4()
|
|
draft = {
|
|
"the_work": [
|
|
_work("backend", be_proj),
|
|
_work("frontend", fe_proj),
|
|
]
|
|
}
|
|
team = await service._resolve_owning_team(
|
|
draft,
|
|
resolved_product_id=None,
|
|
resolved_assigned_to=None,
|
|
team_override=None,
|
|
default_lead=Team.BACKEND,
|
|
)
|
|
assert team is Team.MAIN_PM
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_resolve_owning_team_single_cell_still_routes_to_lead_cell() -> None:
|
|
"""A single-cell project draft (no product, no multi-cell map) keeps its
|
|
legacy owner: the lead cell."""
|
|
service = get_prompter_service()
|
|
draft = {"the_work": [_work("backend", uuid4())]}
|
|
team = await service._resolve_owning_team(
|
|
draft,
|
|
resolved_product_id=None,
|
|
resolved_assigned_to=None,
|
|
team_override=None,
|
|
default_lead=Team.BACKEND,
|
|
)
|
|
assert team is Team.BACKEND
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_resolve_owning_team_product_with_cell_map_stays_board(
|
|
monkeypatch: pytest.MonkeyPatch,
|
|
) -> None:
|
|
"""#160: a product draft that also carries a ≥2-cell the_work map is still a
|
|
product root — on the board-review path it stays team=board, not forced to
|
|
Main PM (which would strand it past the CEO Approve & Start gate)."""
|
|
service = get_prompter_service()
|
|
be_proj, fe_proj = uuid4(), uuid4()
|
|
draft = {"the_work": [_work("backend", be_proj), _work("frontend", fe_proj)]}
|
|
product_id = uuid4()
|
|
|
|
async def _is_board(_agent_id: UUID) -> bool:
|
|
return True
|
|
|
|
monkeypatch.setattr(service, "_assignee_is_board", _is_board)
|
|
team = await service._resolve_owning_team(
|
|
draft,
|
|
resolved_product_id=product_id,
|
|
resolved_assigned_to=uuid4(),
|
|
team_override=None,
|
|
default_lead=Team.BACKEND,
|
|
)
|
|
assert team is Team.BOARD
|
|
|
|
async def _not_board(_agent_id: UUID) -> bool:
|
|
return False
|
|
|
|
monkeypatch.setattr(service, "_assignee_is_board", _not_board)
|
|
team = await service._resolve_owning_team(
|
|
draft,
|
|
resolved_product_id=product_id,
|
|
resolved_assigned_to=uuid4(),
|
|
team_override=None,
|
|
default_lead=Team.BACKEND,
|
|
)
|
|
assert team is Team.MAIN_PM
|
|
|
|
|
|
def test_clean_list_extracts_dict_wrapped_items() -> None:
|
|
"""#159: _clean_list (via coerce_str_list) extracts text from the Claude
|
|
SDK's XML-ish dict wrappers (``<item>…</item>`` -> ``{"item": {"$text": …}}``)
|
|
instead of rendering ``str(dict)``. Pins the behavior so a regression to
|
|
``str(dict)`` in the rendered description is caught."""
|
|
out = _clean_list([{"item": {"$text": "build it"}}, "ship it", " ", ""])
|
|
assert out == ["build it", "ship it"]
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_create_task_from_draft_preserves_product_with_one_cell_map(
|
|
db_session: Any,
|
|
) -> None:
|
|
"""#57: a draft carrying a top-level product_id AND a 1-cell the_work map
|
|
keeps the product — the lone cell map is redundant, not a signal to drop the
|
|
product and force the cell's project_id."""
|
|
_project_id, ceo_id = await _seed_project_and_ceo(db_session)
|
|
product_id = uuid4()
|
|
db_session.add(
|
|
ProductTable(
|
|
id=product_id,
|
|
name="One-cell product",
|
|
slug=f"prod-{uuid4().hex[:8]}",
|
|
description="x",
|
|
created_by=ceo_id,
|
|
)
|
|
)
|
|
await db_session.flush()
|
|
service = get_prompter_service(db=db_session)
|
|
draft = {
|
|
"title": "Board-led single-cell product",
|
|
"acceptance_criteria": ["done"],
|
|
"product_id": str(product_id),
|
|
"the_work": [_work("backend", uuid4())],
|
|
}
|
|
task = await service.create_task_from_draft(draft, ceo_id)
|
|
assert task.product_id == product_id
|
|
assert task.project_id is None
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_create_task_from_draft_does_not_mutate_caller_draft(
|
|
db_session: Any,
|
|
) -> None:
|
|
"""#59: create_task_from_draft coerces + recomposes on a copy — the caller's
|
|
draft dict and its the_work unit dicts are left untouched (no in-place
|
|
rewrite of acceptance_criteria / items)."""
|
|
project_id, ceo_id = await _seed_project_and_ceo(db_session)
|
|
service = get_prompter_service(db=db_session)
|
|
original_items = [" trim me ", "keep"]
|
|
draft: dict[str, Any] = {
|
|
"title": "No-mutation check",
|
|
"acceptance_criteria": ["done"],
|
|
"project_id": str(project_id),
|
|
"the_work": [
|
|
{"team": "backend", "summary": "s", "items": list(original_items)}
|
|
],
|
|
}
|
|
await service.create_task_from_draft(draft, ceo_id)
|
|
# The caller's the_work unit items were NOT coerced in place...
|
|
assert draft["the_work"][0]["items"] == original_items
|
|
# ...and the top-level acceptance_criteria was NOT replaced.
|
|
assert draft["acceptance_criteria"] == ["done"]
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_create_task_from_draft_defaults_source_to_prompter(
|
|
db_session: Any,
|
|
) -> None:
|
|
project_id, ceo_id = await _seed_project_and_ceo(db_session)
|
|
service = get_prompter_service(db=db_session)
|
|
draft = {
|
|
"title": "Default-source draft",
|
|
"acceptance_criteria": ["done"],
|
|
"project_id": str(project_id),
|
|
}
|
|
task = await service.create_task_from_draft(draft, ceo_id)
|
|
assert task.source == "prompter"
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_create_task_from_draft_accepts_custom_source(
|
|
db_session: Any,
|
|
) -> None:
|
|
"""A non-intake caller (e.g. an approved roadmap item) stamps its own
|
|
source tag on the draft instead of the intake default."""
|
|
project_id, ceo_id = await _seed_project_and_ceo(db_session)
|
|
service = get_prompter_service(db=db_session)
|
|
draft = {
|
|
"title": "Roadmap-sourced draft",
|
|
"acceptance_criteria": ["done"],
|
|
"project_id": str(project_id),
|
|
"source": "roadmap",
|
|
}
|
|
task = await service.create_task_from_draft(draft, ceo_id)
|
|
assert task.source == "roadmap"
|
|
assert task.confirmed_by_human is True # the CEO approval IS the confirmation
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_create_task_from_draft_rejects_unwhitelisted_source(
|
|
db_session: Any,
|
|
) -> None:
|
|
"""An LLM-authored draft can't impersonate a privileged origin: a source
|
|
outside the whitelist falls back to 'prompter' (a 'release_manager' spoof
|
|
would otherwise wedge the release engine's one-open-proposal dedup)."""
|
|
project_id, ceo_id = await _seed_project_and_ceo(db_session)
|
|
service = get_prompter_service(db=db_session)
|
|
draft = {
|
|
"title": "Spoofed-source draft",
|
|
"acceptance_criteria": ["done"],
|
|
"project_id": str(project_id),
|
|
"source": "release_manager",
|
|
}
|
|
task = await service.create_task_from_draft(draft, ceo_id)
|
|
assert task.source == "prompter"
|
|
|
|
|
|
# =============================================================================
|
|
# Prompter memory v1 — history digest + compact search rows (pure, no DB)
|
|
# =============================================================================
|
|
|
|
|
|
def _task(title: str, **overrides: Any) -> TaskTable:
|
|
"""An unattached TaskTable instance — plain attribute assignment, no session.
|
|
|
|
Defaults to a completed backend task with no dates; pass ``completed_at`` /
|
|
``updated_at`` / ``created_at`` / ``status`` / ``team`` to override.
|
|
"""
|
|
fields: dict[str, Any] = {
|
|
"id": uuid4(),
|
|
"title": title,
|
|
"status": TaskStatus.COMPLETED,
|
|
"team": Team.BACKEND,
|
|
"completed_at": None,
|
|
"updated_at": None,
|
|
"created_at": None,
|
|
}
|
|
fields.update(overrides)
|
|
return TaskTable(**fields)
|
|
|
|
|
|
def test_task_activity_date_prefers_completed_at() -> None:
|
|
now = datetime.now(UTC)
|
|
task = _task(
|
|
"t",
|
|
completed_at=now,
|
|
updated_at=now - timedelta(days=1),
|
|
created_at=now - timedelta(days=2),
|
|
)
|
|
assert _task_activity_date(task) == now
|
|
|
|
|
|
def test_task_activity_date_falls_back_to_updated_at() -> None:
|
|
now = datetime.now(UTC)
|
|
task = _task(
|
|
"t", completed_at=None, updated_at=now, created_at=now - timedelta(days=1)
|
|
)
|
|
assert _task_activity_date(task) == now
|
|
|
|
|
|
def test_task_activity_date_falls_back_to_created_at() -> None:
|
|
now = datetime.now(UTC)
|
|
task = _task("t", completed_at=None, updated_at=None, created_at=now)
|
|
assert _task_activity_date(task) == now
|
|
|
|
|
|
def test_title_excerpt_leaves_short_titles_untouched() -> None:
|
|
assert _title_excerpt("Fix login bug") == "Fix login bug"
|
|
|
|
|
|
def test_title_excerpt_truncates_long_titles_with_ellipsis() -> None:
|
|
long_title = "A" * 100
|
|
excerpt = _title_excerpt(long_title)
|
|
assert len(excerpt) == _HISTORY_TITLE_EXCERPT_CAP
|
|
assert excerpt.endswith("…")
|
|
|
|
|
|
def test_build_history_digest_empty_is_blank() -> None:
|
|
assert build_history_digest([]) == ""
|
|
|
|
|
|
def test_build_history_digest_caps_at_limit_keeps_most_recent() -> None:
|
|
now = datetime.now(UTC)
|
|
# t0 oldest ... t19 newest.
|
|
ascending = [
|
|
_task(f"t{i}", created_at=now + timedelta(days=i), updated_at=None)
|
|
for i in range(20)
|
|
]
|
|
# Mimic the DB's most-recent-first ordering.
|
|
most_recent_first = list(reversed(ascending))
|
|
|
|
digest = build_history_digest(most_recent_first)
|
|
|
|
lines = digest.splitlines()
|
|
assert len(lines) == _HISTORY_DIGEST_PER_PROJECT_LIMIT
|
|
for i in range(5): # the 5 oldest are excluded
|
|
assert f"`{str(ascending[i].id)[:8]}`" not in digest
|
|
for i in range(5, 20): # the 15 most recent are present
|
|
assert f"`{str(ascending[i].id)[:8]}`" in digest
|
|
|
|
|
|
def test_build_history_digest_renders_oldest_first() -> None:
|
|
now = datetime.now(UTC)
|
|
a = _task("Task A", created_at=now - timedelta(days=2), updated_at=None)
|
|
b = _task("Task B", created_at=now - timedelta(days=1), updated_at=None)
|
|
c = _task("Task C", created_at=now, updated_at=None)
|
|
|
|
# DB order is most-recent-first: C, B, A.
|
|
digest = build_history_digest([c, b, a])
|
|
|
|
idx_a = digest.index("Task A")
|
|
idx_b = digest.index("Task B")
|
|
idx_c = digest.index("Task C")
|
|
assert idx_a < idx_b < idx_c
|
|
|
|
|
|
def test_compact_task_rows_shape() -> None:
|
|
now = datetime.now(UTC)
|
|
task = _task(
|
|
"Fix login bug",
|
|
status=TaskStatus.COMPLETED,
|
|
team=Team.BACKEND,
|
|
completed_at=now,
|
|
)
|
|
rows = compact_task_rows([task])
|
|
assert len(rows) == 1
|
|
row = rows[0]
|
|
assert set(row.keys()) == {"id", "title", "status", "team", "date"}
|
|
assert row["id"] == str(task.id)
|
|
assert row["title"] == "Fix login bug"
|
|
assert row["status"] == "completed"
|
|
assert row["team"] == "backend"
|
|
assert row["date"] == now.date().isoformat()
|
|
|
|
|
|
def test_compact_task_rows_preserves_none_team() -> None:
|
|
task = _task("No team", team=None, created_at=datetime.now(UTC))
|
|
rows = compact_task_rows([task])
|
|
assert rows[0]["team"] is None
|
|
|
|
|
|
# -----------------------------------------------------------------------------
|
|
# history_digest_layer — ambient-block assembly (project_history_digest stubbed)
|
|
# -----------------------------------------------------------------------------
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_history_digest_layer_empty_projects_returns_none() -> None:
|
|
assert await history_digest_layer(cast("AsyncSession", object()), []) is None
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_history_digest_layer_single_project_has_no_header(
|
|
monkeypatch: pytest.MonkeyPatch,
|
|
) -> None:
|
|
async def _fake(_session: Any, _project: Any, *, _limit: int = 15) -> str | None:
|
|
return "- `abc12345` Some task (completed, 2026-01-01)"
|
|
|
|
monkeypatch.setattr(prompter_module, "project_history_digest", _fake)
|
|
project = SimpleNamespace(slug="roboco", id=uuid4())
|
|
|
|
text = await history_digest_layer(cast("AsyncSession", object()), [project])
|
|
|
|
assert text is not None
|
|
assert text.startswith("## Task History\n\n### Recent tasks\n")
|
|
assert "### Recent tasks —" not in text
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_history_digest_layer_multi_project_headers_by_slug(
|
|
monkeypatch: pytest.MonkeyPatch,
|
|
) -> None:
|
|
projects = [
|
|
SimpleNamespace(slug="backend-svc", id=uuid4()),
|
|
SimpleNamespace(slug="frontend-app", id=uuid4()),
|
|
]
|
|
|
|
async def _fake(_session: Any, project: Any, *, _limit: int = 15) -> str | None:
|
|
return f"- `deadbeef` Task for {project.slug} (completed, 2026-01-01)"
|
|
|
|
monkeypatch.setattr(prompter_module, "project_history_digest", _fake)
|
|
|
|
text = await history_digest_layer(cast("AsyncSession", object()), projects)
|
|
|
|
assert text is not None
|
|
assert "### Recent tasks — `backend-svc`" in text
|
|
assert "### Recent tasks — `frontend-app`" in text
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_history_digest_layer_skips_projects_with_no_tasks(
|
|
monkeypatch: pytest.MonkeyPatch,
|
|
) -> None:
|
|
has_tasks = SimpleNamespace(slug="has-tasks", id=uuid4())
|
|
no_tasks = SimpleNamespace(slug="empty-proj", id=uuid4())
|
|
|
|
async def _fake(_session: Any, project: Any, *, _limit: int = 15) -> str | None:
|
|
return (
|
|
"- `deadbeef` A task (completed, 2026-01-01)"
|
|
if project is has_tasks
|
|
else None
|
|
)
|
|
|
|
monkeypatch.setattr(prompter_module, "project_history_digest", _fake)
|
|
|
|
text = await history_digest_layer(
|
|
cast("AsyncSession", object()), [has_tasks, no_tasks]
|
|
)
|
|
|
|
assert text is not None
|
|
assert "has-tasks" in text
|
|
assert "empty-proj" not in text
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_history_digest_layer_all_empty_returns_none(
|
|
monkeypatch: pytest.MonkeyPatch,
|
|
) -> None:
|
|
async def _fake(_session: Any, _project: Any, *, _limit: int = 15) -> str | None:
|
|
return None
|
|
|
|
monkeypatch.setattr(prompter_module, "project_history_digest", _fake)
|
|
projects = [
|
|
SimpleNamespace(slug="a", id=uuid4()),
|
|
SimpleNamespace(slug="b", id=uuid4()),
|
|
]
|
|
|
|
assert await history_digest_layer(cast("AsyncSession", object()), projects) is None
|