mirror of
https://github.com/rennf93/roboco.git
synced 2026-08-03 07:23:24 +02:00
* feat(batch): batch_id + collision descriptor columns
Sequenced batch intake ("Mega task") foundation: tasks.batch_id (indexed)
groups a batch of top-level tasks created together; intends_to_touch (text[]),
adds_migration and touches_shared (bool, NOT NULL default false) are the
per-task collision surface the SequencingService will read to wire dependency
waves. Mirrored on the Task model + TaskCreateRequest and wired through
TaskService.create. Migration 046 (real upgrade->downgrade->upgrade verified
vs a throwaway pgvector PG); a non-batch task declares no surface (defaults).
Task 1 of the 0.11.0 sequenced-batch-intake plan.
* feat(batch): flag + draft collision descriptors
Default-off ROBOCO_BATCH_INTAKE_ENABLED (config + FEATURE_FLAGS + panel card);
the propose_draft tool doc + the TS DraftProposal gain the per-task collision
surface intends_to_touch / adds_migration / touches_shared. The draft is a loose
dict so the descriptors ride it through the relay intact (test asserts the
forwarded payload); the analyzer (Task 3) reads them to wire dependency waves.
Task 2 of the 0.11.0 sequenced-batch-intake plan.
* feat(batch): deterministic collision-sequencing analyzer
SequencingService.analyze turns a batch's per-task collision surfaces into a
dependency DAG + execution waves — correctness in CODE, not agent judgment.
Rules in order: file overlap serializes (more-important first), migrations form
a serial chain (no concurrent Alembic heads), touches_shared runs last, cell
contention warns (never serializes); then dedupe, existence + cycle check, and
Kahn topological layering. Pure (no DB/services); SequencingError on a cycle or
out-of-range edge.
Golden test reproduces the CEO's hand-sequenced 4 waves of the 11-item
guard-core-app batch (the effort that deadlocked the Main PM): S6 alone last,
the R1/R3/R4 migration chain, R2/R3/S8 serialized on the shared threat service,
S1/S2/S7 in one parallel wave.
Task 3 of the 0.11.0 sequenced-batch-intake plan.
* chore(batch): brand the user-facing surfaces "MegaTask"
The user-facing name is MegaTask: the feature-flag label is "MegaTask intake",
the panel flag-card and the config description lead with MegaTask. Internal
names stay technical (batch_intake_enabled, batch_id, SequencingService).
* chore(batch): drop the feature flag — MegaTask is a core intake scope
MegaTask is additive and opt-in by its own nature (the Prompter proposes a
batch only when the CEO asks for several tasks; single-task intake is
unchanged), so there is no risk surface a flag protects — 'don't create a
MegaTask' is the off switch. Remove batch_intake_enabled from config, the
FEATURE_FLAGS registry, the panel flag card, and its tests. MegaTask will be
a third scope option in the Intake modal (single-cell / multi-project /
MegaTask), not a toggle.
* feat(batch): MegaTask identity predicate + orchestrator branchless recognition
The single source of truth for the umbrella's exemptions: pure
is_batch_umbrella / is_batch_root_subtask / is_branchless_coordination
(foundation/policy/batch.py) — an umbrella has a batch_id and is top-level; a
root-subtask shares the batch_id but is parented. The orchestrator's
_is_coordination_task now consults is_branchless_coordination, so a MegaTask
umbrella is recognized as doing no git of its own (git-exempt at spawn-readiness
/ stuck-detection) exactly like a product fan-out root. Non-batch behavior is
identical (the predicate reduces to the old no-project+product check; the
orchestrator coordination suite stays green), and the umbrella branch is inert
until the create path exists.
First slice of the MegaTask umbrella enforcement (branchless guard).
* feat(batch): branchless umbrella guard across the git-exemption sites
A MegaTask umbrella does no git of its own — every git-exemption site in
TaskService now consults the shared is_branchless_coordination predicate
instead of an inline product-only check, so the umbrella's exemptions
cannot drift between sites:
- the claimed->in_progress branch gate (GitContext.is_coordination) lets
an unbranched umbrella reach in_progress and delegate;
- _ensure_branch_for_task short-circuits an umbrella to "" instead of the
misconfigured raise (the claim path ignores the return, treating it as
branchless);
- CEO-reject routing sends a rejected umbrella to the Main PM in PENDING
(needs_revision is developer-claim-only and would deadlock it).
Covers both shapes via the predicate (product fan-out root OR umbrella);
a batch root-subtask keeps its own branch/PR. Adds orchestrator
recognition tests for the umbrella plus claim/branch/reject integration
tests.
* feat(batch): umbrella assembles no PR; completes branchless
submit_root now hard-rejects a MegaTask umbrella up front (a preflight
that also folds in the unknown-role refusal to stay within the
return-count budget): the umbrella spans many projects with no single
master, so each root-subtask opens and is reviewed on its own PR — the
umbrella never enters the in-path review gate. The Main PM completes it
directly once every root-subtask is terminal.
Umbrella completion needs no new code: it is branchless (no branch_name),
so _main_pm_complete_guard already accepts it from in_progress, checks
all_subtasks_terminal, and main_pm_complete walks it to awaiting_pm_review
and escalates to the CEO with no PR creation — exactly the product
fan-out root path. Adds the submit_root-reject and umbrella-completion
gateway tests; pins batch_id=None on the normal-root submit_root test
(a MagicMock auto-attr would otherwise read as an umbrella).
* feat(batch): MegaTask create path — umbrella + sequenced root-subtasks
PrompterService.confirm_live_batch turns N confirmed drafts into a real
MegaTask: it builds each draft's collision surface, runs the pure
SequencingService to get conflict-free waves, creates the branchless
umbrella (batch_id, no project/product), then one root-subtask per draft
(own project, parent=umbrella, sequence=wave index, descriptors), and
wires the analyzer's edges through add_dependency so the existing
dependency-gate runs the waves in order. The route picks the start path
like a single confirm: 'board' holds the root-subtasks in BACKLOG for the
batch review; 'main_pm' creates them PENDING so wave 0 dispatches at once.
create_task_from_draft gains a BatchPlacement (parent/batch/sequence/
team_override) and forwards the collision descriptors; the exactly-one-
target rule (here and the TaskService.create invariant) is relaxed for an
umbrella, which legitimately targets neither. New route
POST /live/{session}/confirm-batch + BatchConfirmRequest mirror the single
confirm. Adds the structural-invariant + board-hold + empty-batch tests.
* feat(batch): release MegaTask root-subtasks on CEO approval; board awareness
The board route holds a MegaTask's root-subtasks in BACKLOG so the work
waits for the batch review. approve_and_start (CEO gate #1, board->Main PM)
now releases them via _activate_batch_root_subtasks: each held child flips
BACKLOG -> PENDING + team=main_pm so the dependency-gate dispatches wave 0.
No-op for a non-umbrella; idempotent (children past BACKLOG untouched).
The Product Owner and Head of Marketing identity prompts gain a MegaTask
section so they review the whole batch + wave plan and adjust scope before
sign-off (they review drafts; the umbrella is their unit). Also extracts
the create() target invariant into _require_target_or_umbrella to keep the
method under the complexity gate after the umbrella exemption. Adds the
umbrella-approval activation test.
* feat(batch): multi-project intake scope for MegaTask
A MegaTask spans several possibly-unrelated repos, so the intake chat can
now be scoped to an explicit project list (not just one project or one
product). StartLiveRequest gains project_ids; /live/start threads it
through start/spawn_intake_session -> _spawn_intake_container ->
_clone_intake_scope. The multi-repo clone machinery already existed for
products; _intake_scope_slugs now also resolves an explicit project_ids
set (split into _slugs_for_project_ids / _slugs_for_product), cloning each
repo with the first as the primary cwd and the siblings readable. Scope
validation is now 'exactly one of project_slug / product_id / project_ids'
via the shared _require_one_intake_scope. Adds scope-resolution, spawn,
and route tests for the MegaTask path.
* feat(batch): propose_batch intake tool (MegaTask multi-draft hand-off)
The intake agent can now hand the panel a whole MegaTask in one tool call.
Both intake paths gain propose_batch alongside propose_draft:
- Claude (intake_driver): a propose_batch tool registered on the in-SDK
MCP server + allowlisted; the driver intercepts the ToolUseBlock and
emits ONE StreamChunk(kind="batch") carrying {drafts:[...], title}.
- grok (intake_server): a propose_batch tool that POSTs a "batch" relay
event via the shared _post_event helper (post_draft/post_batch).
A batch carries N drafts, each the propose_draft shape PLUS its own
project_id (a MegaTask spans unrelated repos) and collision surface so the
analyzer sequences the waves. The prompter prompt documents the MegaTask
scope + when to call propose_batch. Adds Claude-normalize and grok-relay
tests for the batch path.
* feat(batch): MegaTask intake panel — third scope, batch review, waves
The panel now drives a MegaTask end to end. The intake modal gains a
third scope, 'MegaTask', beside Single cell and Board-led: a multi-project
checklist (a MegaTask spans several possibly-unrelated repos), validated
to at least two. start() sends project_ids; use-prompter accumulates the
agent's single propose_batch hand-off as a 'batch' SSE event into a
BatchProposal and lands in a new batch_preview state.
A new BatchReviewCard lists every proposed task with its target project +
collision-surface badges (migration / shared) and offers one start path
for the whole batch — Board review & Start or Approve & Start — wired to
confirmBatch → POST /confirm-batch. The success card shows the sequenced
result: N tasks in M waves (+ any advisory notes). prompter.ts gains the
DraftScale 'megatask' + the BatchConfirm payload/result types; the SSE
client allows the 'batch' kind. Panel typecheck + lint + 113 tests green.
* docs(batch): MegaTask across changelog, CLAUDE.md, site, and RAG
The four documentation obligations for the MegaTask feature:
- CHANGELOG: an Unreleased entry covering the umbrella model, sequencing,
multi-project intake, propose_batch, and the create/approval path.
- CLAUDE.md: a MegaTask section (identity predicate, umbrella/root-subtask
hierarchy, sequencing rules, intake + create path, board activation).
- Published site: a user-facing company/megatask.md (scopes, waves, the
umbrella, the two start buttons) + nav entry; a pointer added to the
intake chapter of the Tour.
- RAG corpus: workflows/megatask.md so the Main PM (and any agent) can
retrieve the umbrella's branchless / no-PR / completion rules at runtime.
The runtime concurrent-migration guard is intentionally NOT added: the
analyzer already chains migration-adders into dependencies and the
dependency-gate serializes them, so a separate guard would be dead code.
* feat(batch): batch_id guardrail + wave preview + batch_id on TaskResponse
Guardrail (CEO): a batch_id is denied on any task that is not a well-formed
MegaTask member. is_valid_batch_shape permits batch_id only on an umbrella
(no parent → must target neither project nor product) or a root-subtask
(has a parent → exactly one target); TaskService.create enforces it AND
verifies a root-subtask's parent is the batch umbrella (same batch_id,
top-level). This closes a latent hole: is_batch_umbrella is true for a
batch_id + no-parent task even with a project, so a stray batch_id could
have spoofed the branchless branch-gate / no-PR exemption. (The public
task API never exposed batch_id for write; this guards the service layer.)
Wave preview: PrompterService.preview_batch + POST .../preview-batch
compute a MegaTask's waves from the proposed drafts WITHOUT creating
anything, so the panel can show the sequencing before confirm. Extracted
_sequence_drafts as the single source shared by preview and confirm, so
the previewed waves are exactly the ones wired.
TaskResponse now carries batch_id so the panel can badge the umbrella.
* feat(batch): MegaTask review — project editor, wave preview, persistence, badge
Closes the panel gaps in the MegaTask review experience:
- Per-task project editor: each proposed task gets an inline project
Select (updateBatchDraftProject), so a task the agent put in the wrong
or no repo can be fixed before launch — not only by re-chatting. Launch
stays blocked until every task has a project.
- Wave preview: on a batch proposal the panel fetches POST .../preview-batch
(no task created) and shows the conflict-free wave plan, so the human
reviews the sequencing before confirming.
- Refresh durability: the MegaTask review (batch + waves + projectIds) is
persisted, so a browser reload mid-review restores it like a single draft.
- MegaTask badge: TaskResponse exposes batch_id, the panel Task type
carries it, and the task table badges the umbrella row 'MegaTask'.
Panel typecheck + lint + 113 tests green.
* test(batch): stub task carries batch_id for task_to_response
task_to_response now serializes batch_id (TaskResponse field), so the
_stub_task SimpleNamespace fixture must provide it — without it the reader
hit AttributeError, failing the 8 task-schema serialization/enrichment
tests. Test-only; the real TaskTable carries the column (migration 046).
* fix(batch): close MegaTask audit gaps — completion crash, analyzer cycle, guardrails
An adversarial multi-agent audit of the feature surfaced 20 verified gaps;
this closes the backend ones.
HIGH:
- Umbrella completion crashed. escalate_to_ceo hard-required a pr_number,
which a branchless umbrella never has, so main_pm_complete dereferenced
None. Both pr_number gates now waive a MegaTask umbrella (escalate_to_ceo
+ the awaiting_pm_review->awaiting_ceo_approval lifecycle gate via a new
GitContext.is_umbrella), and main_pm_complete guards a None return. The
completion test had mocked escalate_to_ceo, hiding it — now a real
service test covers the waiver.
- The collision analyzer could fabricate a cycle (a touches_shared +
adds_migration draft overlapping another migration draft) and raise
SequencingError — a bare ValueError that escaped as an opaque 500. The
migration chain is now shared-last-aware (never contradicts rule 3), and
_sequence_drafts translates SequencingError to a clean 400.
MEDIUM:
- Collisions are now project-scoped: two repos can't collide on a
coincidental path or serialize independent migrations (DraftSurface
carries project_id; rules 1/2/3 respect it).
- The batch_id guardrail ran only at create. update() + the PATCH
null-clear path now re-assert is_valid_batch_shape, so a mutation can't
break a member's shape and spoof the branchless exemption.
- A draft missing title/acceptance_criteria now raises ValidationError
(was a bare KeyError -> 500).
- confirm_live_batch re-asserts every draft targets a scoped project and
the batch spans >=2 distinct projects (project_ids added to the request).
- Route-level tests for confirm-batch / preview-batch.
LOW: strict multi-repo clone (fail loud on any unresolvable project);
malformed/empty propose_batch surfaces an error chunk (Claude) / refuses
to POST (grok) instead of silently acking; dropped malformed drafts are
counted and surfaced; stale grok intake docstrings updated.
* fix(batch): MegaTask panel + doc audit gaps
Frontend half of the audit fixes:
- The confirm payload now carries project_ids (the schema requires it), and
the panel re-checks every task targets one of the scoped repos before
launching, naming the offending task.
- The Review-MegaTask project picker is filtered to the scoped repos and
the per-task validity (border + launch gate) keys off scoped membership,
so a task can only be (re)pointed at an in-scope project — also fixing the
case where the agent emitted a non-UUID / unknown project.
- Dropped malformed drafts are surfaced as a chat error so the human knows
the batch shrank instead of silently confirming fewer tasks.
- Doc wording: a wave releases on the previous wave's terminal state
(normally a merge; a cancellation releases it too), not strictly 'merged'.
* test(batch): lock the CEO's EXACT 4-wave hand-sequencing as the golden bar
The golden test asserted the constraints (S6 last, the migration chain, the
shared-threats serialization, S1/S2/S7 parallel) but not the full wave
partition. The bar for MegaTask is 'reproduce my exact waves or it's not
done', so assert the exact 4-wave partition the analyzer produces for the
guard-core-app batch:
wave 1: R1 R2 S1 S2 S3 S5 S7 · wave 2: R3 · wave 3: R4 S8 · wave 4: S6
Confirmed unchanged by the audit's analyzer fixes (no migration is shared;
single project).
* fix(batch): tolerate a stub task in assert_batch_shape_intact
The batch-shape re-validation read task.batch_id directly, but update()'s
partial-caller contract is exercised with a SimpleNamespace stub that has no
batch_id column → AttributeError. Use getattr(..., None) for batch_id and the
shape fields so the guard no-ops on any task lacking the column (a stub, or a
non-batch task) while still enforcing on a real batch member.
* fix(orchestrator): authenticate internal API self-calls with the system identity
The dispatcher httpx clients were built without an agent identity, so the
orchestrator's self-PATCHes to /api/tasks/{id} (auto-block, auto-resume,
auto-recover, SLA annotation) were rejected 401 "Missing X-Agent-ID" and
silently no-op'd. The auto-resume that lifts a PM's paused parent could never
write, so paused/blocked parents stayed wedged and stranded their dependents
(the fe-pm/be-pm respawn churn seen in prod).
Header propagation was inconsistent across the separate AsyncClient call-sites:
only the main dispatch client carried the system identity; the readiness and
sweep clients did not. Hoist the identity into a shared _SYSTEM_API_HEADERS
constant and apply it to every API-facing dispatcher client. The system role
holds TaskAction.ASSIGN, so it is authorized for the audited admin_set_status
path those write routes use. The external provider-recovery probe client is
intentionally left untouched.
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
633 lines
23 KiB
Python
633 lines
23 KiB
Python
"""Unit tests for PrompterService.
|
|
|
|
Covers the live-intake draft → task flow (``create_task_from_draft`` /
|
|
``confirm_live_draft`` + the enum/priority/team coercion) and the pure
|
|
draft/description helpers. DB-backed tests use an in-memory async session via
|
|
conftest fixtures.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
from typing import Any, cast
|
|
from uuid import UUID, uuid4
|
|
|
|
import pytest
|
|
from roboco.db.tables import (
|
|
AgentTable,
|
|
ProductTable,
|
|
ProjectTable,
|
|
TaskTable,
|
|
)
|
|
from roboco.models.base import (
|
|
AgentRole,
|
|
AgentStatus,
|
|
Complexity,
|
|
TaskNature,
|
|
TaskStatus,
|
|
TaskType,
|
|
Team,
|
|
)
|
|
from roboco.seeds.initial_data import AGENT_UUIDS
|
|
from roboco.services.base import ServiceError, ValidationError
|
|
from roboco.services.prompter import (
|
|
PrompterService,
|
|
compose_description,
|
|
derive_scale,
|
|
get_prompter_service,
|
|
parse_readiness,
|
|
)
|
|
|
|
# =============================================================================
|
|
# Pure function tests (no DB)
|
|
# =============================================================================
|
|
|
|
|
|
def test_parse_readiness_extracts_and_strips_tag() -> None:
|
|
content = (
|
|
"Here is my question about scope.\n\n"
|
|
'```roboco-meta\n{"covered": ["objective", "scope"], '
|
|
'"ready": true, "scale": "multi"}\n```'
|
|
)
|
|
clean, tag = parse_readiness(content)
|
|
assert clean == "Here is my question about scope."
|
|
assert tag is not None
|
|
assert tag.ready is True
|
|
assert tag.scale == "multi"
|
|
assert tag.covered == ["objective", "scope"]
|
|
# The control block must not leak into the user-visible text.
|
|
assert "roboco-meta" not in clean
|
|
|
|
|
|
def test_parse_readiness_absent_block_is_not_ready() -> None:
|
|
clean, tag = parse_readiness("Just a plain reply, no control block.")
|
|
assert clean == "Just a plain reply, no control block."
|
|
assert tag is None
|
|
|
|
|
|
def test_parse_readiness_malformed_json_is_graceful() -> None:
|
|
content = "Reply text.\n```roboco-meta\n{not valid json]\n```"
|
|
clean, tag = parse_readiness(content)
|
|
assert "roboco-meta" not in clean
|
|
assert clean == "Reply text."
|
|
assert tag is None
|
|
|
|
|
|
def test_parse_readiness_uses_last_block() -> None:
|
|
content = (
|
|
'```roboco-meta\n{"ready": false, "scale": "single"}\n```\n'
|
|
"Final answer.\n"
|
|
'```roboco-meta\n{"ready": true, "scale": "multi"}\n```'
|
|
)
|
|
clean, tag = parse_readiness(content)
|
|
assert tag is not None
|
|
assert tag.ready is True
|
|
assert tag.scale == "multi"
|
|
assert "roboco-meta" not in clean
|
|
|
|
|
|
def test_derive_scale_single_vs_multi() -> None:
|
|
assert derive_scale([{"team": "backend"}]) == "single"
|
|
assert derive_scale([{"team": "backend"}, {"team": "frontend"}]) == "multi"
|
|
# Non-cell teams (e.g. main_pm) do not count toward cell breadth.
|
|
assert derive_scale([{"team": "backend"}, {"team": "main_pm"}]) == "single"
|
|
assert derive_scale([]) == "single"
|
|
|
|
|
|
def test_compose_description_single_cell_markdown() -> None:
|
|
draft = {
|
|
"objective": "Let humans track token usage.",
|
|
"what_this_builds": ["A usage panel on the Metrics page"],
|
|
"the_work": [
|
|
{
|
|
"team": "frontend",
|
|
"summary": "Render the usage panel",
|
|
"items": ["Add the chart", "Wire the API"],
|
|
}
|
|
],
|
|
"notes": ["Reuse the existing Metrics layout"],
|
|
"acceptance_criteria": ["Panel shows totals", "Panel filters by range"],
|
|
}
|
|
md = compose_description(draft)
|
|
assert "## Objective" in md
|
|
assert "## What This Builds" in md
|
|
assert "## The Work" in md
|
|
assert "**Frontend** — Render the usage panel" in md
|
|
assert "## Notes" in md
|
|
assert "## Success Criteria" in md
|
|
assert "- Panel shows totals" in md
|
|
# Single-cell tasks get no board-led lead line.
|
|
assert "Board-led" not in md
|
|
|
|
|
|
def test_compose_description_multi_cell_has_board_led_lead() -> None:
|
|
draft = {
|
|
"objective": "Ship the Prompter.",
|
|
"the_work": [
|
|
{"team": "backend", "summary": "Chat endpoint", "items": []},
|
|
{"team": "frontend", "summary": "Chat UI", "items": []},
|
|
{"team": "ux_ui", "summary": "Interaction design", "items": []},
|
|
],
|
|
"acceptance_criteria": ["It works end to end"],
|
|
}
|
|
md = compose_description(draft)
|
|
assert "Board-led" in md
|
|
assert "**Backend**" in md
|
|
assert "**UX/UI**" in md
|
|
|
|
|
|
def test_compose_description_falls_back_to_provided_description() -> None:
|
|
# Sparse structured fields → fall back to a model-provided description.
|
|
draft = {"description": "A perfectly adequate fallback description here."}
|
|
md = compose_description(draft)
|
|
assert md == "A perfectly adequate fallback description here."
|
|
|
|
|
|
def test_lead_cell_team_prefers_the_work_cell() -> None:
|
|
draft = {"the_work": [{"team": "frontend"}], "team": "backend"}
|
|
assert PrompterService._lead_cell_team(draft, default=Team.BACKEND) is Team.FRONTEND
|
|
# Empty the_work falls back to the provided default.
|
|
assert PrompterService._lead_cell_team({}, default=Team.BACKEND) is Team.BACKEND
|
|
|
|
|
|
def test_lead_cell_team_skips_invalid_cell_names() -> None:
|
|
# An off-enum cell name is skipped, not raised on; falls through to a valid one.
|
|
draft = {"the_work": [{"team": "nonsense"}, {"team": "frontend"}]}
|
|
assert PrompterService._lead_cell_team(draft, default=Team.BACKEND) is Team.FRONTEND
|
|
|
|
|
|
def test_coerce_draft_enums_defaults_invalid_values() -> None:
|
|
# Regression: the LLM emits off-enum values (e.g. task_type="feature"). The
|
|
# confirm must coerce to defaults, never raise — a bad enum guess must not
|
|
# 400 the launch and force the agent to self-correct in-chat.
|
|
draft = {
|
|
"team": "backend",
|
|
"task_type": "feature", # not a valid TaskType
|
|
"nature": "bogus", # not a valid TaskNature
|
|
"estimated_complexity": "enormous", # not a valid Complexity
|
|
}
|
|
team, task_type, nature, complexity = PrompterService._coerce_draft_enums(draft)
|
|
assert team is Team.BACKEND
|
|
assert task_type is TaskType.CODE
|
|
assert nature is TaskNature.TECHNICAL
|
|
assert complexity is Complexity.MEDIUM
|
|
|
|
|
|
def test_coerce_priority_maps_words_clamps_and_defaults() -> None:
|
|
# Regression: priority is the one non-enum field the agent guesses, and it
|
|
# guesses a word ("high") as often as a number — int("high") used to 500.
|
|
# word/number -> expected priority int (0=urgent .. 3=low).
|
|
cases: dict[object, int] = {
|
|
"urgent": 0,
|
|
"high": 1,
|
|
"medium": 2,
|
|
"low": 3,
|
|
1: 1,
|
|
"3": 3,
|
|
99: 3, # clamped into range
|
|
"nonsense": 2, # unrecognized -> default medium
|
|
None: 2, # missing -> default medium
|
|
}
|
|
for value, expected in cases.items():
|
|
assert PrompterService._coerce_priority(value) == expected
|
|
|
|
|
|
def test_coerce_draft_enums_keeps_valid_and_derives_missing_team() -> None:
|
|
# Valid values pass through; a missing team is derived from the_work.
|
|
draft = {
|
|
"task_type": "documentation",
|
|
"nature": "technical",
|
|
"estimated_complexity": "medium",
|
|
"the_work": [{"team": "frontend"}],
|
|
}
|
|
team, task_type, nature, complexity = PrompterService._coerce_draft_enums(draft)
|
|
assert team is Team.FRONTEND
|
|
assert task_type is TaskType.DOCUMENTATION
|
|
assert nature is TaskNature.TECHNICAL
|
|
assert complexity is Complexity.MEDIUM
|
|
|
|
|
|
# =============================================================================
|
|
# Factory
|
|
# =============================================================================
|
|
|
|
|
|
def test_get_prompter_service_no_db() -> None:
|
|
service = get_prompter_service()
|
|
assert isinstance(service, PrompterService)
|
|
assert service._db is None
|
|
|
|
|
|
def test_get_prompter_service_raises_without_db_for_session_methods() -> None:
|
|
service = get_prompter_service()
|
|
with pytest.raises(ServiceError, match="DB session"):
|
|
_ = service._session
|
|
|
|
|
|
# =============================================================================
|
|
# DB-backed: assignee routing + confirm_live_draft
|
|
# =============================================================================
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_assignee_is_board_distinguishes_roles(db_session: Any) -> None:
|
|
"""Drives product team routing: a board reviewer keeps the root on the board.
|
|
|
|
A product confirmed via "Board review & Start" is assigned to a board
|
|
reviewer and must stay team=board so the CEO's Approve & Start gate appears;
|
|
one assigned to main-pm (or a cell dev) is not a board task.
|
|
"""
|
|
service = get_prompter_service(db=db_session)
|
|
|
|
def _agent(role: AgentRole) -> AgentTable:
|
|
return AgentTable(
|
|
id=uuid4(),
|
|
name="A",
|
|
slug=f"a-{uuid4().hex[:8]}",
|
|
role=role,
|
|
team=None,
|
|
status=AgentStatus.ACTIVE,
|
|
model_config={},
|
|
system_prompt="x",
|
|
capabilities=[],
|
|
permissions={},
|
|
metrics={},
|
|
)
|
|
|
|
po = _agent(AgentRole.PRODUCT_OWNER)
|
|
hom = _agent(AgentRole.HEAD_MARKETING)
|
|
dev = _agent(AgentRole.DEVELOPER)
|
|
db_session.add_all([po, hom, dev])
|
|
await db_session.flush()
|
|
|
|
assert await service._assignee_is_board(cast("UUID", po.id)) is True
|
|
assert await service._assignee_is_board(cast("UUID", hom.id)) is True
|
|
assert await service._assignee_is_board(cast("UUID", dev.id)) is False
|
|
# Unknown id is not a board agent — defensive, must not raise.
|
|
assert await service._assignee_is_board(uuid4()) is False
|
|
|
|
|
|
async def _seed_project_and_ceo(db_session: Any) -> tuple[UUID, UUID]:
|
|
"""Seed a system agent + project + CEO; return (project_id, ceo_id).
|
|
|
|
Returns plain ``UUID``s (not the ORM rows) so callers pass real uuids to the
|
|
service — no casting the ORM ``.id`` column type at the call site.
|
|
"""
|
|
system_id, project_id, ceo_id = uuid4(), uuid4(), uuid4()
|
|
system = AgentTable(
|
|
id=system_id,
|
|
name="System",
|
|
slug=f"system-{uuid4().hex[:8]}",
|
|
role=AgentRole.SYSTEM,
|
|
team=None,
|
|
status=AgentStatus.ACTIVE,
|
|
model_config={},
|
|
system_prompt="system",
|
|
capabilities=[],
|
|
permissions={},
|
|
metrics={},
|
|
)
|
|
db_session.add(system)
|
|
await db_session.flush()
|
|
project = ProjectTable(
|
|
id=project_id,
|
|
name="Intake Test Project",
|
|
slug=f"intake-{uuid4().hex[:8]}",
|
|
git_url="https://github.com/example/intake.git",
|
|
default_branch="main",
|
|
protected_branches=["main"],
|
|
assigned_cell=Team.BACKEND,
|
|
created_by=system_id,
|
|
is_active=True,
|
|
)
|
|
ceo = AgentTable(
|
|
id=ceo_id,
|
|
name="CEO",
|
|
slug=f"ceo-{uuid4().hex[:8]}",
|
|
role=AgentRole.CEO,
|
|
team=None,
|
|
status=AgentStatus.ACTIVE,
|
|
model_config={},
|
|
system_prompt="ceo",
|
|
capabilities=[],
|
|
permissions={},
|
|
metrics={},
|
|
)
|
|
db_session.add_all([project, ceo])
|
|
await db_session.flush()
|
|
# The "& Start" routes assign the draft to a fixed board/PM agent
|
|
# (product-owner for "Board review", main-pm for "Approve & Start"); those
|
|
# rows must exist for the assigned_to FK. merge() is idempotent, so this is
|
|
# safe whether or not another test already committed them on the shared DB.
|
|
for slug, role, team in (
|
|
("product-owner", AgentRole.PRODUCT_OWNER, None),
|
|
("main-pm", AgentRole.MAIN_PM, Team.MAIN_PM),
|
|
):
|
|
await db_session.merge(
|
|
AgentTable(
|
|
id=UUID(AGENT_UUIDS[slug]),
|
|
name=slug,
|
|
slug=slug,
|
|
role=role,
|
|
team=team,
|
|
status=AgentStatus.ACTIVE,
|
|
model_config={},
|
|
system_prompt=slug,
|
|
capabilities=[],
|
|
permissions={},
|
|
metrics={},
|
|
)
|
|
)
|
|
await db_session.flush()
|
|
return project_id, ceo_id
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_confirm_live_draft_board_route_assigns_po(db_session: Any) -> None:
|
|
""" "Board review & Start" (default route) → PENDING, assigned to the Product
|
|
Owner so the orchestrator fires the PO + HoM review."""
|
|
project_id, ceo_id = await _seed_project_and_ceo(db_session)
|
|
service = get_prompter_service(db=db_session)
|
|
|
|
draft = {
|
|
"title": "Add token metrics",
|
|
"objective": "See token usage at a glance.",
|
|
"acceptance_criteria": ["Dashboard shows total tokens"],
|
|
"team": "backend",
|
|
"the_work": [
|
|
{"team": "backend", "summary": "instrument", "items": ["count tokens"]}
|
|
],
|
|
}
|
|
task_id = await service.confirm_live_draft(draft, ceo_id, project_id=project_id)
|
|
|
|
row = await db_session.get(TaskTable, task_id)
|
|
assert row is not None
|
|
assert row.status == TaskStatus.PENDING # "& Start" — started now
|
|
assert row.assigned_to == UUID(AGENT_UUIDS["product-owner"]) # board review
|
|
assert row.source == "prompter"
|
|
assert row.confirmed_by_human is True
|
|
assert row.team == Team.BACKEND # lead cell from the_work
|
|
assert row.created_by == ceo_id
|
|
assert row.nature is not None and row.task_type is not None
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_confirm_live_draft_main_pm_route_assigns_main_pm(
|
|
db_session: Any,
|
|
) -> None:
|
|
""" "Approve & Start" (route="main_pm") → PENDING, assigned to the Main PM."""
|
|
project_id, ceo_id = await _seed_project_and_ceo(db_session)
|
|
service = get_prompter_service(db=db_session)
|
|
draft = {
|
|
"title": "Quick fix",
|
|
"acceptance_criteria": ["done"],
|
|
"team": "backend",
|
|
}
|
|
task_id = await service.confirm_live_draft(
|
|
draft, ceo_id, project_id=project_id, route="main_pm"
|
|
)
|
|
row = await db_session.get(TaskTable, task_id)
|
|
assert row.status == TaskStatus.PENDING
|
|
assert row.assigned_to == UUID(AGENT_UUIDS["main-pm"])
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_confirm_live_draft_product_routes_to_main_pm(db_session: Any) -> None:
|
|
"""A product-scoped draft via the "Approve & Start" path is a Main-PM root.
|
|
|
|
The board path (the ``route="board"`` default) keeps the root at
|
|
``team=board`` until the CEO approves; the Main-PM path is selected
|
|
explicitly with ``route="main_pm"``.
|
|
"""
|
|
_project_id, ceo_id = await _seed_project_and_ceo(db_session)
|
|
product_id = uuid4()
|
|
product = ProductTable(
|
|
id=product_id,
|
|
name="Intake Product",
|
|
slug=f"prod-{uuid4().hex[:8]}",
|
|
description="x",
|
|
created_by=ceo_id,
|
|
)
|
|
db_session.add(product)
|
|
await db_session.flush()
|
|
|
|
service = get_prompter_service(db=db_session)
|
|
draft = {
|
|
"title": "Board-led feature",
|
|
"acceptance_criteria": ["works end to end"],
|
|
"team": "backend",
|
|
}
|
|
task_id = await service.confirm_live_draft(
|
|
draft, ceo_id, product_id=product_id, route="main_pm"
|
|
)
|
|
row = await db_session.get(TaskTable, task_id)
|
|
assert row.team == Team.MAIN_PM
|
|
assert row.product_id == product_id
|
|
assert row.project_id is None
|
|
|
|
|
|
# =============================================================================
|
|
# MegaTask: confirm_live_batch (umbrella + sequenced root-subtasks)
|
|
# =============================================================================
|
|
|
|
|
|
async def _seed_second_project(db_session: Any, ceo_id: UUID) -> UUID:
|
|
"""Seed a second project so a MegaTask can span multiple repos."""
|
|
project_id = uuid4()
|
|
db_session.add(
|
|
ProjectTable(
|
|
id=project_id,
|
|
name="Intake Test Project 2",
|
|
slug=f"intake2-{uuid4().hex[:8]}",
|
|
git_url="https://github.com/example/intake2.git",
|
|
default_branch="main",
|
|
protected_branches=["main"],
|
|
assigned_cell=Team.FRONTEND,
|
|
created_by=ceo_id,
|
|
is_active=True,
|
|
)
|
|
)
|
|
await db_session.flush()
|
|
return project_id
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_confirm_live_batch_builds_umbrella_and_sequenced_subtasks(
|
|
db_session: Any,
|
|
) -> None:
|
|
"""A MegaTask creates one branchless umbrella + N root-subtasks across many
|
|
projects, with the collision-derived dependency edges wired so the
|
|
dependency-gate runs the waves in order."""
|
|
project1, ceo_id = await _seed_project_and_ceo(db_session)
|
|
project2 = await _seed_second_project(db_session, ceo_id)
|
|
service = get_prompter_service(db=db_session)
|
|
|
|
# A & B both add a migration → serial chain A→B (the migration rule orders
|
|
# them by priority then index). C is an independent frontend task in another
|
|
# project, so it runs in parallel with A in wave 0.
|
|
drafts: list[dict[str, Any]] = [
|
|
{
|
|
"title": "A: add table",
|
|
"acceptance_criteria": ["a"],
|
|
"team": "backend",
|
|
"project_id": str(project1),
|
|
"intends_to_touch": ["roboco/services/foo.py"],
|
|
"adds_migration": True,
|
|
},
|
|
{
|
|
"title": "B: extend table",
|
|
"acceptance_criteria": ["b"],
|
|
"team": "backend",
|
|
"project_id": str(project1),
|
|
"intends_to_touch": ["roboco/services/bar.py"],
|
|
"adds_migration": True,
|
|
},
|
|
{
|
|
"title": "C: frontend widget",
|
|
"acceptance_criteria": ["c"],
|
|
"team": "frontend",
|
|
"project_id": str(project2),
|
|
"intends_to_touch": ["panel/src/widget.tsx"],
|
|
},
|
|
]
|
|
result = await service.confirm_live_batch(
|
|
"Three things",
|
|
drafts,
|
|
ceo_id,
|
|
project_ids=[project1, project2],
|
|
route="main_pm",
|
|
)
|
|
|
|
# A (migration) and C (independent) run in wave 0; B chains after A.
|
|
assert result["waves"] == [[0, 2], [1]]
|
|
ids = result["root_subtask_ids"]
|
|
assert len(ids) == len(drafts)
|
|
|
|
umbrella_id = UUID(result["umbrella_task_id"])
|
|
umbrella = await db_session.get(TaskTable, umbrella_id)
|
|
assert umbrella.batch_id is not None
|
|
assert umbrella.parent_task_id is None
|
|
assert umbrella.project_id is None and umbrella.product_id is None
|
|
assert umbrella.team == Team.MAIN_PM
|
|
assert umbrella.status == TaskStatus.PENDING
|
|
assert umbrella.branch_name is None # branchless
|
|
|
|
a, b, c = [await db_session.get(TaskTable, UUID(sid)) for sid in ids]
|
|
for sub in (a, b, c):
|
|
assert sub.parent_task_id == umbrella_id
|
|
assert sub.batch_id == umbrella.batch_id
|
|
assert sub.team == Team.MAIN_PM
|
|
assert sub.status == TaskStatus.PENDING
|
|
assert a.project_id == project1
|
|
assert b.project_id == project1
|
|
assert c.project_id == project2
|
|
# sequence = wave index: A and C in wave 0, B in wave 1.
|
|
assert (a.sequence, b.sequence, c.sequence) == (0, 1, 0)
|
|
# Dependency wiring: B waits on A; C is independent.
|
|
assert UUID(ids[0]) in b.dependency_ids
|
|
assert c.dependency_ids == []
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_confirm_live_batch_board_route_holds_subtasks_in_backlog(
|
|
db_session: Any,
|
|
) -> None:
|
|
"""The "board" route sends the umbrella to the Product Owner for batch review
|
|
and holds the root-subtasks in BACKLOG until the umbrella is approved."""
|
|
project1, ceo_id = await _seed_project_and_ceo(db_session)
|
|
project2 = await _seed_second_project(db_session, ceo_id)
|
|
service = get_prompter_service(db=db_session)
|
|
drafts = [
|
|
{
|
|
"title": "One",
|
|
"acceptance_criteria": ["x"],
|
|
"team": "backend",
|
|
"project_id": str(project1),
|
|
},
|
|
{
|
|
"title": "Two",
|
|
"acceptance_criteria": ["y"],
|
|
"team": "frontend",
|
|
"project_id": str(project2),
|
|
},
|
|
]
|
|
result = await service.confirm_live_batch(
|
|
"Two repos", drafts, ceo_id, project_ids=[project1, project2], route="board"
|
|
)
|
|
|
|
umbrella = await db_session.get(TaskTable, UUID(result["umbrella_task_id"]))
|
|
assert umbrella.team == Team.BOARD
|
|
assert umbrella.assigned_to == UUID(AGENT_UUIDS["product-owner"])
|
|
assert umbrella.status == TaskStatus.PENDING
|
|
sub = await db_session.get(TaskTable, UUID(result["root_subtask_ids"][0]))
|
|
assert sub.status == TaskStatus.BACKLOG # held until batch review approves
|
|
assert sub.team == Team.BOARD
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_confirm_live_batch_rejects_empty(db_session: Any) -> None:
|
|
_project1, ceo_id = await _seed_project_and_ceo(db_session)
|
|
service = get_prompter_service(db=db_session)
|
|
with pytest.raises(ValidationError):
|
|
await service.confirm_live_batch(
|
|
"Empty", [], ceo_id, project_ids=[uuid4(), uuid4()]
|
|
)
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_confirm_live_batch_rejects_draft_outside_scope(db_session: Any) -> None:
|
|
"""A draft targeting a project NOT in the scoped project_ids is refused — the
|
|
intake agent only read the scoped repos."""
|
|
project1, ceo_id = await _seed_project_and_ceo(db_session)
|
|
project2 = await _seed_second_project(db_session, ceo_id)
|
|
service = get_prompter_service(db=db_session)
|
|
outside = uuid4() # never in scope
|
|
drafts = [
|
|
{"title": "A", "acceptance_criteria": ["a"], "project_id": str(project1)},
|
|
{"title": "B", "acceptance_criteria": ["b"], "project_id": str(outside)},
|
|
]
|
|
with pytest.raises(ValidationError, match="outside this MegaTask"):
|
|
await service.confirm_live_batch(
|
|
"Scoped", drafts, ceo_id, project_ids=[project1, project2], route="main_pm"
|
|
)
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_confirm_live_batch_rejects_single_project(db_session: Any) -> None:
|
|
"""A degenerate batch whose drafts all target one project is not a MegaTask."""
|
|
project1, ceo_id = await _seed_project_and_ceo(db_session)
|
|
project2 = await _seed_second_project(db_session, ceo_id)
|
|
service = get_prompter_service(db=db_session)
|
|
drafts = [
|
|
{"title": "A", "acceptance_criteria": ["a"], "project_id": str(project1)},
|
|
{"title": "B", "acceptance_criteria": ["b"], "project_id": str(project1)},
|
|
]
|
|
with pytest.raises(ValidationError, match="at least two distinct projects"):
|
|
await service.confirm_live_batch(
|
|
"One repo",
|
|
drafts,
|
|
ceo_id,
|
|
project_ids=[project1, project2],
|
|
route="main_pm",
|
|
)
|
|
|
|
|
|
def test_preview_batch_computes_waves_without_creating() -> None:
|
|
"""preview_batch is pure: it returns the same waves confirm would wire, with
|
|
no DB session and no task creation."""
|
|
service = get_prompter_service() # no db — pure compute
|
|
drafts: list[dict[str, Any]] = [
|
|
{"title": "A", "adds_migration": True, "intends_to_touch": ["a.py"]},
|
|
{"title": "B", "adds_migration": True, "intends_to_touch": ["b.py"]},
|
|
{"title": "C", "intends_to_touch": ["c.py"]},
|
|
]
|
|
result = service.preview_batch(drafts)
|
|
# A & B chain on the migration rule; C is independent → [[0, 2], [1]].
|
|
assert result["waves"] == [[0, 2], [1]]
|
|
assert isinstance(result["warnings"], list)
|
|
|
|
|
|
def test_preview_batch_rejects_empty() -> None:
|
|
service = get_prompter_service()
|
|
with pytest.raises(ValidationError):
|
|
service.preview_batch([])
|