Batch parity with the single-draft keep-alive redraft loop. A first
board-route confirm-batch parks the intake session against the umbrella
(instead of the unconditional reap), so the existing board-completion
injection reaches the still-live chat — now with a batch-aware brief
(compose_batch_redraft_message: live root-subtask snapshots + board
notes + a one-propose_batch re-proposal instruction). The re-confirm
carries BatchConfirmRequest.task_id and routes to the new
PrompterService.update_live_batch: in-place umbrella + root-subtask
update (positional patch of live children, cancel+recreate on scope
change, create/cancel on count change, dependency edges rewired to the
fresh wave plan) gated by the same _validate_batch_scope as create.
Readers use the CANCELLED-excluding get_live_subtasks view so
multi-round redrafts survive earlier cancels.
Cold path: re-interview now handles a branchless umbrella by recovering
its multi-repo scope from live children (distinct_projects_for_batch)
and returning project_ids — fixes the live 400 behind the task-detail
redraft button on umbrellas. Panel: confirmBatch board branch keeps the
chat open, threads batchRedraftTaskIdRef (persisted) into the
re-confirm, treats a redraft re-confirm as terminal on both routes, and
surfaces the server's real validation message on confirm failure.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
Within a MegaTask root-subtask, the per-cell tasks got sequence numbers but zero
dependency edges, so they ran fully in parallel (UX finished after backend
started, frontend self-blocked) — divergent branches, duplicated/wasted work.
The cross-cell wiring (_wire_ux_frontend_dependency: FE/BE cells depend on the UX
cell, bidirectional, propagated to dev subtasks via inherit_unmet_dependencies)
already exists, but it bails unless the parent has a product_id. A MegaTask
root-subtask has no product_id — it targets its cells via cell_projects — so the
wiring silently no-op'd for every MegaTask root (confirmed on the live video
root: product_id=None, three cells, all with empty dependency_ids).
Broaden the guard to fire on product_id OR is_batch_root_subtask(batch_id,
parent_task_id) (scalar fields; cell_projects is a lazy relationship). The same
tested wiring now holds MegaTask cells in order like a product fan-out.
Adds test_megatask_root_wires_cross_cell_ux_dependency.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
The Main PM claimed every MegaTask wave at once, ignoring the collision-ordered
dependencies. The sequencing data was correct (analyzer wired proper waves), but
enforcement was only half-wired: the unmet-dependency guard lives on the gateway
claim verbs (i_will_plan -> _run_claim_guards), while the orchestrator dispatches
coordination roots itself — _dispatch_pm_work fetches pending with no dependency
filter and _claim_task_for_agent system-claims via the raw POST /tasks/{id}/claim
route -> TaskService.claim, which had no dependency check. So the orchestrator
claimed every pending root-subtask for the Main PM regardless of wave.
Enforce the sequence at the claim chokepoint: _validate_claim_preconditions now
refuses to claim a PENDING task while any depends_on task is non-terminal
(extracted into _claim_blocked_by_dependencies for the complexity budget). This
guardrails every claim path — the gateway verbs (redundant) and the orchestrator
raw dispatch claim (the hole). Scoped to a PENDING start-of-work claim so a
mid-lifecycle QA/doc claim is unaffected; dependencies are monotonic so each wave
claims normally once the prior one completes.
Adds test_claim_pending_with_unmet_dependency_returns_none (blocked with an
unfinished dependency; claimable once it completes).
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* fix(prompter): batch confirm ignores a vestigial top-level project slug
The MegaTask confirm-batch 400'd with "Invalid project_id UUID: roboco-api".
The intake agent authors each draft's project as the repo slug it read, and the
panel only nulls the top-level project_id when the CEO toggles that draft's
picker — so an untouched draft carried the slug through to create_task_from_draft,
whose eager _resolve_uuid_field(project_id) raised on the non-UUID and rejected
the whole batch (guard-release from the prior fix surfaced it as a clean 400
instead of a wedged 500).
A batch root-subtask targets its repos via the the_work per-cell map (the panel
fills those with real project UUIDs); the top-level project_id/product_id is
vestigial for it. Strip it from a sub-draft when its cell map carries the real
target — a legacy no-the_work draft keeps its panel-filled top-level UUID, and
scope validation (which already runs off the cell map) is unchanged.
Adds a repro test with two cell-map drafts that also carry leftover top-level
slugs (roboco-api / roboco-panel); they now confirm instead of 400-ing.
* fix(tests): conventions PR integration test honors the #375 workspace-scope guard
#375 added a containment guard to open_conventions_pr (workspace_path must sit
under {workspaces_root}/{slug}); the unit test was updated but this integration
test still seeded a bare tmp_path/repo, so open_conventions_pr returned None and
test_open_conventions_pr_commits_locally_without_remote failed on master. Anchor
workspaces_root at the test dir and place the repo under the project's slug.
* refactor(prompter): extract batch sub-draft sanitize (xenon rank B)
The inline vestigial-target strip pushed _build_confirm_batch to cyclomatic
rank C (over the --max-absolute B gate). Move the assigned_to + top-level
project/product stripping into a pure _batch_subtask_draft helper; behavior is
unchanged, _build_confirm_batch drops back under the limit.
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* fix(prompter): keep a watched intake chat alive (idle-reap counted reading as idle)
Intake chats "dropped after a while" — the panel showed "Live connection lost".
The idle reaper retires an interactive session whose last_activity is older than
interactive_idle_reap_seconds (30m default), but last_activity was bumped only by
an agent event or a human turn. An open SSE stream — the human reading a proposed
draft / MegaTask spec without typing — bumped nothing, so a chat under active
review was reaped mid-read, closing the stream (the SSE transport error the panel
reports as "Live connection lost").
stream() now runs a keepalive task that refreshes last_activity every 60s while
the stream is connected, so an open, actively-watched chat counts as alive; when
the tab closes the generator ends, the keepalive is cancelled, and a genuinely
abandoned chat still reaps after the threshold. The keepalive runs beside an
un-cancelled queue.get() so no live token or the close sentinel can be dropped.
* fix(tests): conventions PR integration test honors the #375 workspace-scope guard
#375 added a containment guard to open_conventions_pr (workspace_path must sit
under {workspaces_root}/{slug}); the unit test was updated but this integration
test still seeded a bare tmp_path/repo, so open_conventions_pr returned None and
test_open_conventions_pr_commits_locally_without_remote failed on master. Anchor
workspaces_root at the test dir and place the repo under the project's slug.
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
The fleet-wide subagent ban was implemented as an allowlist omission, but Task
is a default-permitted Claude Code built-in — an allowlist auto-approves, it
does not restrict. Under permission_mode="dontAsk" (intake/secretary SDK) and
defaultMode="bypassPermissions" (fleet), Task ran regardless and can_use_tool
was never invoked for it, so every Claude-path agent could still spawn
subagents despite allows_subagent=False. Only the grok path blocked it.
Explicitly disallow the subagent tool at every Claude-path spawn point:
disallowed_tools=["Task"] on the intake and secretary SDK drivers, and "Task"
in the fleet settings.json base_deny (an explicit deny applies even under
bypassPermissions). This mirrors the grok path's --disallowed-tools Agent.
Pins the ban in test_cc_lockdown.py (fleet settings deny Task) and a new
test_sdk_driver_subagent_ban.py (intake + secretary options disallow Task).
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
The MegaTask review card was a static sibling of the scrollable chat list with
no height bound, so a tall batch overflowed the clipped container and stranded
the launch buttons off-screen with no scrollbar. It now owns the scroll area
(min-h-0 flex-1 overflow-y-auto), matching the pattern ChatMessages already uses.
confirm_live_batch's Redis idempotency guard (1h TTL) was acquired before the
build but only released via a success sidecar, so a build failure wedged the
session: every retry hit ServiceError("already in progress") -> HTTP 500 for an
hour. The build now releases the guard on any failure before the sidecar write,
so a retry re-attempts (and surfaces the real error) instead of being locked out.
Extracted _build_confirm_batch to keep the try/except thin.
Adds DB-backed tests for the panel the_work[].project_id shape, multi-cell
root-subtasks, dense same-repo collisions (both routes), and guard release.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
Fix the 4 real CodeQL path-injection alerts (open_conventions_pr trusted the
API-settable project.workspace_path with no containment) plus defense-in-depth
segment validation at the get_workspace_path chokepoint. Bump next 16.1.1->16.1.7
and transitive lockfile deps to clear 24 Dependabot alerts. Close the intake
subagent-ban gap: the Claude intake driver still carried the Task tool and the
prompter prompt told it to fan out research subagents, contradicting the
fleet-wide ban. The remaining 47 CodeQL + 29 Dependabot alerts are dismissed on
GitHub with per-alert justifications (guard patterns CodeQL can't model across
call hops; next 16.2.x blocked by the verified tab-hostage router regression).
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
CEO verdict on the blind 3-day timer: 'It should default to 1 day and
be more like... smart.' Now: interval defaults to 1 day; the cycle
skips (with logged reasons) while a spotlight draft is still awaiting
the CEO, and stretches to 3x the interval when nothing has shipped
(CHANGELOG sections via the read clone) since the last spotlight
activity — where activity is a materialized draft's seen_at or a
completed exploration's updated_at, deliberately excluding the stale-
cycle janitor's cancels. The HoM gains an explicit skip exit
(propose_feature_spotlight skip=true + reason: completes the
exploration, no draft, no seen-slug, still counts as activity), and
its spawn prompt now carries the seen ledger WITH dates, what shipped
since the last spotlight, and recently rejected drafts with the CEO's
reasons — fresh-but-unspotlighted first. Fail-open on changelog read
errors so a signal outage never starves the engine.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* [e0bc418f] add v019-release HyperFrames composition with vertical/square cuts, theme, props, and smoke tests
* [e0bc418f] fix(v019-release): remove local artifacts, replace em dashes, add captions.json and regression tests
* [e0bc418f] chore: delete .venv symlink and tighten .gitignore for symlink leakage
* [e0bc418f] chore(.gitignore): remove duplicate .pnpm-store line
* [e0bc418f] fix(v019-release): clean internal em dashes and add package-lock regression test
* [e0bc418f] chore(v019-release): remove committed vitest cache artifact from node_modules
* [e0bc418f] fix(.gitignore): restore .pnpm-store/ next to its GH001 comment and remove stray duplicate
* [e0bc418f] docs(motion): document v019-release composition and captions convention
* [e0bc418f] feat(release-recap): rebuild v0.19.0 video as panel-demo three-release recap
CEO rejected the text-card v019-release composition and rescoped this
to a recap of three releases (0.18.0-0.20.0) shipped in six days,
built in the panel-demo register per the CEO's brief. Brings
motion/kit/ and motion/compositions/panel-demo/ into this branch
(byte-identical to master, which merged them after this branch
diverged and cannot be rebased onto directly), deletes the obsolete
motion/compositions/v019-release/ text-card composition, and adds
motion/compositions/release-recap/: the panel frame, one intake typing
reveal, three release cards (v0.18.0/v0.19.0/v0.20.0) cycling through
a single shared kanban slot with their own progress-to-completed pill
flip, a cursor that glides in once and click-pulses at each flip, a
closing toast, and the roboco.tech outro, both at 1080x1920 and
1080x1080. Includes props.js, captions.json (X 136/280 chars, TikTok
365/2200 chars, no em dash), and a vitest smoke test (11/11 passing
across all three compositions in the package). Also refreshes
motion/README.md with the current design-bar/panel-demo-kit sections
and a release-recap example replacing the deleted v019-release one.
* [e0bc418f] fix(release-recap): polish pass — accumulate the three releases, fix facts, robust typewriter
Cards now stack and accumulate instead of cycling one absolute slot (the
three-release story never landed visually); duration tightened 20s to
15s; per-release copy corrected (0.18.0 fleet doctrine + design bar,
0.19.0 video engine + sandboxes, 0.20.0 root-owned coverage + the kit)
along with the same fixes in captions.json; one cursor with one glide
and one click landing at the pill edge instead of parking on pill text
for 17s; the per-char typing reveal (which partially froze) replaced by
a single width-steps mono typewriter sized to the supplied text.
* [e0bc418f] feat(release-recap): identity redesign — real logo, real-UI sidebar, explainers, stat beat
CEO review round 2: no logo, no similarity with the actual UI, explains
nothing. Rebuilt: branded cold open (the real logo on an app-icon tile
with a ring sweep, wordmark, tagline), a labeled sidebar mirroring the
real panel (nav labels + active state + search + Live dot + CEO chip),
a one-line explainer under each release title, a kinetic stat beat
(3 releases / 6 days / 1 human — full-frame interstitial on the square
cut where no side zone exists), and a logo lockup outro; 18s. The logo
PNG is vendored into motion/public (offline constraint, same pattern
as the fonts).
* [e0bc418f] fix(release-recap): darkmode logo hero — crisp 1024px mark replaces the pixelated 64px tile
The cold open now uses logos/roboco-logo-darkmode.png (vendored into
motion/public), built for dark backgrounds with the wordmark included:
full-size mark with a soft glow and the ring sweep, no white tile, no
separate typed wordmark. The small app-icon tiles in the sidebar and
outro keep the 64px icon where tiny sizes favor it.
---------
Co-authored-by: UX/UI Developer 1 <ux-dev-1@roboco.tech>
Co-authored-by: UX/UI Documenter <ux-doc@roboco.tech>
Co-authored-by: UX/UI Developer 2 <ux-dev-2@roboco.tech>
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
The last test asserting the prompter kept the Agent tool — it pinned
the allowance through the disallowed-tools string, which the #372 test
sweep (grepping the flag names) missed.
* [e0bc418f] add v019-release HyperFrames composition with vertical/square cuts, theme, props, and smoke tests
* [e0bc418f] fix(v019-release): remove local artifacts, replace em dashes, add captions.json and regression tests
* [e0bc418f] chore: delete .venv symlink and tighten .gitignore for symlink leakage
* [e0bc418f] chore(.gitignore): remove duplicate .pnpm-store line
* [e0bc418f] fix(v019-release): clean internal em dashes and add package-lock regression test
* [e0bc418f] chore(v019-release): remove committed vitest cache artifact from node_modules
* [e0bc418f] fix(.gitignore): restore .pnpm-store/ next to its GH001 comment and remove stray duplicate
* [e0bc418f] docs(motion): document v019-release composition and captions convention
* [e0bc418f] feat(release-recap): rebuild v0.19.0 video as panel-demo three-release recap
CEO rejected the text-card v019-release composition and rescoped this
to a recap of three releases (0.18.0-0.20.0) shipped in six days,
built in the panel-demo register per the CEO's brief. Brings
motion/kit/ and motion/compositions/panel-demo/ into this branch
(byte-identical to master, which merged them after this branch
diverged and cannot be rebased onto directly), deletes the obsolete
motion/compositions/v019-release/ text-card composition, and adds
motion/compositions/release-recap/: the panel frame, one intake typing
reveal, three release cards (v0.18.0/v0.19.0/v0.20.0) cycling through
a single shared kanban slot with their own progress-to-completed pill
flip, a cursor that glides in once and click-pulses at each flip, a
closing toast, and the roboco.tech outro, both at 1080x1920 and
1080x1080. Includes props.js, captions.json (X 136/280 chars, TikTok
365/2200 chars, no em dash), and a vitest smoke test (11/11 passing
across all three compositions in the package). Also refreshes
motion/README.md with the current design-bar/panel-demo-kit sections
and a release-recap example replacing the deleted v019-release one.
* [e0bc418f] fix(release-recap): polish pass — accumulate the three releases, fix facts, robust typewriter
Cards now stack and accumulate instead of cycling one absolute slot (the
three-release story never landed visually); duration tightened 20s to
15s; per-release copy corrected (0.18.0 fleet doctrine + design bar,
0.19.0 video engine + sandboxes, 0.20.0 root-owned coverage + the kit)
along with the same fixes in captions.json; one cursor with one glide
and one click landing at the pill edge instead of parking on pill text
for 17s; the per-char typing reveal (which partially froze) replaced by
a single width-steps mono typewriter sized to the supplied text.
* [e0bc418f] feat(release-recap): identity redesign — real logo, real-UI sidebar, explainers, stat beat
CEO review round 2: no logo, no similarity with the actual UI, explains
nothing. Rebuilt: branded cold open (the real logo on an app-icon tile
with a ring sweep, wordmark, tagline), a labeled sidebar mirroring the
real panel (nav labels + active state + search + Live dot + CEO chip),
a one-line explainer under each release title, a kinetic stat beat
(3 releases / 6 days / 1 human — full-frame interstitial on the square
cut where no side zone exists), and a logo lockup outro; 18s. The logo
PNG is vendored into motion/public (offline constraint, same pattern
as the fonts).
* [e0bc418f] fix(release-recap): darkmode logo hero — crisp 1024px mark replaces the pixelated 64px tile
The cold open now uses logos/roboco-logo-darkmode.png (vendored into
motion/public), built for dark backgrounds with the wordmark included:
full-size mark with a soft glow and the ring sweep, no white tile, no
separate typed wordmark. The small app-icon tiles in the sidebar and
outro keep the 64px icon where tiny sizes favor it.
---------
Co-authored-by: UX/UI Developer 1 <ux-dev-1@roboco.tech>
Co-authored-by: UX/UI Documenter <ux-doc@roboco.tech>
Co-authored-by: UX/UI Developer 2 <ux-dev-2@roboco.tech>
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
PLC0415 red on master's quality gate — the module already imports
ROLE_CONFIGS at top level; the function-level re-import from #372's
invariant test is deleted.
CEO directive 2026-07-09: every allows_subagent in role_config flips to
False (was True for cell_pm, main_pm, product_owner, head_marketing,
prompter, secretary); the grok path's drifted _SUBAGENT_ALLOWED_ROLES
allowlist empties to match. The spawn manifest already consumes the
flag, so Claude-path agents lose the Agent tool and grok-path agents
get it in disallowed-tools by construction. New invariant test iterates
every role config so a single role can't quietly regain it.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
Every URL-driven tabbed page (kanban/business/metrics/notifications/KB)
was stuck on the first tab in real Chrome 149 against a production
build: clicking a tab fetched the target RSC payload, then Next's
router MPA-reloaded the CURRENT url, snapping the tab back. Bisected
empirically with a scripted real-Chrome sweep: 16.1.1 green, 16.2.6
red, 16.2.10 (latest stable) red — the regression spans the 16.2 line
and reproduces with zero app data, no proxy, and a fresh profile. Dev
builds and Playwright's bundled Chromium mask it, which is how the
#354 bump (shipped in v0.20.0) passed review. Reverts the next half of
that bump; the axios half stays.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
Agent workspace clones commit as '<Display Name> <slug@roboco.tech>';
the emails map to no GitHub account, so CLA Assistant identifies those
committers by display name and reds any PR carrying agent-authored
commits with a signature no agent can post (live on PR #333). The
fleet is the org's own machinery, not an external contributor — the
roster names are allowlisted, wildcarded per team family.
The inline compound condition pushed _sync_branch_preflight_rejection
to cyclomatic rank C on the PR gate; the whole refusal predicate now
lives in _sync_base_refused.
A standalone task (video/CI-watch/dep-update — no parent) and a child
of a branchless coordination parent merge into the project default
branch by design, but the protected-base guard refused every
master/main base, hard-wedging their rebases into block/PM/respawn
churn (hit live on the v0.19.0 video task). The rebase force-pushes
only the task branch (with lease) and cannot write to the base, so the
guard now refuses master/main only when it is mis-resolved: a
branch-bearing parent exists or the parent row is missing/corrupt.
The '-'-prefixed injection guard stays unconditional.
* feat(video): pipeline visibility — strip, state-aware queue, render-error capture
Task 1 of the 2026-07-09 video-pipeline review. New CEO-gated GET
/video/pipeline lists every in-flight video item (authoring statuses,
rendering attempt n/max, terminal failures with the error — now stamped
onto the video_draft marker instead of dying as a log line).
source_task_id exposed on both video schemas. Panel: pipeline strip on
the Social page, state-aware queue empty copy, title/script on queue
rows, missing cuts disabled instead of a blank player, notifications
deep-link related_task_id. MAX_VIDEO_RENDER_ATTEMPTS moved to the
markers policy layer (single source of truth).
* feat(video): rich authoring briefs — changelog section, brand voice, kit pointer
Task 2 of the 2026-07-09 video-pipeline review. The release brief is
now a structured block (full CHANGELOG section capped at 4000 chars +
highlights) instead of one LLM-compressed sentence; brand_voice and a
motion/kit design-bar pointer are appended centrally in open_video_task
so release, spotlight, and on-demand paths all inherit them.
suggested_input_props seeded on the video_draft marker; third
acceptance criterion pins the design bar; propose_video docstring
points at the kit.
* fix(video): spotlight video drafts on CEO approval, renderer honors data-fps
Task 4 of the 2026-07-09 video-pipeline review. The companion-video
hook moves from propose_feature_spotlight (HoM authoring time) to
XPostService approve's posted-success branch for x_feature drafts,
mirroring the release-publish seam — a rejected spotlight no longer
burns a ux-dev cycle; wants_video/video_script ride the x_feature_ref
marker. Best-effort: a video-engine failure never breaks the post.
render.js reads data-fps from the composition HTML (clamped 24-60,
fallback 30) instead of hardcoding 30; parseFps covered by node --test.
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
Verb-description changes from the recent gateway PRs (declare_coverage,
sync_branch stash, reviewer unclaim) never regenerated the artifacts;
the master gate died at reflow-check before foundation-check could
flag it.
The kit READMEs from #365 were hard-wrapped and failed the master
quality gate's reflow-check (the gate runs on master pushes, not PRs).
uv.lock picks up the 0.20.0 project version from the release bump.
Reusable pk-* building blocks (motion/kit/) that recreate the control
panel's look for HyperFrames compositions: frame chrome, task card,
status pills with swap crossfade, chips, toast, typing reveal, cursor
sprite. Tokens lifted from the panel's dark theme and badge components.
Reference composition motion/compositions/panel-demo/ (12s, both cuts):
a task title types into intake, the card materializes, the cursor
clicks, the status pill flips to completed, a toast slides in.
Verified: motion vitest gate green (7/7); both cuts visually checked
in-browser at native resolution.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* [870467e6] Frontend: page-scoped refresh provider, hook, and navbar button (#347)
* [55376b8a] Create page-scoped refresh provider and context (#327)
* [55376b8a] feat(panel): add page-scoped refresh context and provider
* [55376b8a] docs(frontend): add page-refresh-provider component documentation
---------
Co-authored-by: Frontend Developer 1 <fe-dev-1@roboco.tech>
Co-authored-by: Frontend Documenter <fe-doc@roboco.tech>
* [a0c02d0f] Add public usePageRefresh hook (#332)
* [a0c02d0f] test(hooks): assert usePageRefresh is exported from hooks barrel
* [a0c02d0f] feat(hooks): add public usePageRefresh hook with provider and tests
* [a0c02d0f] fix(panel): move hook test wrappers to components and rename providers.tsx to unshadow barrel
* [a0c02d0f] docs(panel): document usePageRefresh hook and PageRefreshProvider API
---------
Co-authored-by: Frontend Developer 2 <fe-dev-2@roboco.tech>
Co-authored-by: Frontend Documenter <fe-doc@roboco.tech>
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* [5f28dd9b] Add navbar refresh button and remove inline dashboard refresh buttons (#336)
* [5f28dd9b] Align PageRefreshProvider with active hook API and remove inline dashboard refresh buttons
* [5f28dd9b] Remove unused scope-keyed PageRefreshProvider, context, and associated tests
* [5f28dd9b] Address QA revision: add header refresh tests, page-scoped label, remove dead provider code and .venv symlink, revert formatting-only changes
* [5f28dd9b] Remove remaining inline dashboard refresh buttons and committed .venv symlink
* [5f28dd9b] docs(frontend): update page-refresh provider docs and panel README for navbar refresh button
* [5f28dd9b] fix(panel): remove .venv symlink, ignore root .venv entries, and thin task-detail page data fetch into useTaskDetail hook
* [5f28dd9b] Extract GitBrowser data fetching into useGitBrowser hook and add tests; verify .venv cleanup and task-detail thin hook usage
* [5f28dd9b] fix(panel): remove root .venv symlink, restore .gitignore anchored rule, and revert lifecycle.json formatting noise
* Delete .venv
---------
Co-authored-by: Frontend Developer 2 <fe-dev-2@roboco.tech>
Co-authored-by: Frontend Documenter <fe-doc@roboco.tech>
---------
Co-authored-by: Frontend Developer 1 <fe-dev-1@roboco.tech>
Co-authored-by: Frontend Documenter <fe-doc@roboco.tech>
Co-authored-by: Frontend Developer 2 <fe-dev-2@roboco.tech>
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* [b8e1de1b] Fix navbar refresh button disabled state when registry is empty (#356) (#358)
* [b8e1de1b] fix(panel): derive navbar refresh disabled state from registry, not unused prop
PageRefreshProvider now computes `disabled` from whether any refresh
callback is currently registered (registry size > 0) instead of a
static, never-passed `disabled` prop that left the button permanently
enabled. header.tsx now destructures `disabled` from usePageRefresh()
and disables the button on `disabled || loading`. Updated the tests
that asserted the old always-enabled-by-default behavior and added a
new header test asserting the button is disabled with zero registered
callbacks.
* [b8e1de1b] docs(panel): document PageRefreshProvider disabled state derived from registry
Updated documentation to reflect the refactored PageRefreshProvider behavior: the `disabled` state is now derived from whether any refresh callbacks are currently registered (empty registry = disabled), rather than a static `disabled` prop. Clarified in both panel/README.md and the full component guide that the navbar refresh button disables when no callbacks are registered and when a refresh cycle is in progress. Updated API documentation to remove the now-removed `disabled` prop from PageRefreshProviderProps and updated code examples and test coverage descriptions to reflect the new callback-driven semantics.
---------
Co-authored-by: Frontend Developer 1 <fe-dev-1@roboco.tech>
Co-authored-by: Frontend Documenter <fe-doc@roboco.tech>
* test(panel): mock usePageRefresh in tests predating the provider
Merge-skew: the page-refresh feature makes CommandCenter and the agent
detail page call usePageRefresh; three tests merged from master render
them without the new provider. Mock the hook module, matching the
files' stub-everything style.
---------
Co-authored-by: Frontend Developer 1 <fe-dev-1@roboco.tech>
Co-authored-by: Frontend Documenter <fe-doc@roboco.tech>
Co-authored-by: Frontend Developer 2 <fe-dev-2@roboco.tech>
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
The coverage gates had no vocabulary for criteria only the root itself
can satisfy (the supersede PR from feature/main_pm/*, closing the
contributor's PR): once a Main PM declared coverage for the legitimate
cell criteria, the idle gate demanded a cell for the impossible ones
too, so they got pushed into a cell task and the cell PM (correctly)
escalated. declare_coverage now accepts the PM's own task: self-declared
criteria count as claimed for the idle gate and satisfied for the
roll-up (the roll-up actor is their owner by construction), surface as
claimed_by=root in the briefing, and both PM prompts say to never hand
a cell a criterion it cannot satisfy inside its own cell.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
CEO doctrine after an axios bump became a planning parent plus two
sequenced same-branch code children: single-concern work is ONE subtask
covering change, verification, and PR; same-branch sequenced siblings
are one task wearing two ids (the gateway serializes sibling code
subtasks anyway, so the split buys zero parallelism); the i_will_plan
sub_tasks list is a checklist, not a delegation quota.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
The lifespan awaited the reconcile, and the journal/learning backfill
made it expensive: up to 200 entries each needing an Ollama embedding
behind a busy Ollama held the API bind down for 30+ minutes on the NAS
(observed: ~6 embeds/min). uvicorn only binds after the lifespan
completes, so the whole stack 502'd while a best-effort maintenance
pass ran. The reconcile is now a background task scheduled at the end
of startup (crash-logged via done-callback, cancelled at shutdown); the
backfill still converges across boots under its per-boot cap.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* feat(api): x/video post history endpoints
Approved or rejected drafts vanished from both queues permanently --
the listers exclude terminal statuses and no history surface existed,
so a posted tweet or video was only findable in the raw task list.
GET /x/posts/history and GET /video/posts/history (CEO-gated, bounded)
return acted-on drafts newest-first with the posted platform ids and
reject reasons from the draft markers. Route tests assert by identity,
not emptiness: approve/reject commits the whole session, so prior
tests' rows legitimately persist in the shared test DB.
* feat(panel): Social page aggregating post queues and history
New dashboard page composing the X and video post queues with one
unified history section beneath them -- both platforms interleaved
newest-first, kind and outcome badges, posted X ids linking to the
live tweet, reject reasons shown. The command center's two full queue
cards become a compact pending-counts card linking to the page, so the
queues have one home instead of duplicated surfaces.
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
Bootstrap and the API lifespan both ran init_db in one process seconds
apart; the second call re-entered the alembic-in-thread machinery
(nested asyncio.run + NullPool engine + greenlet bridge in a reused
worker thread) for zero benefit and hung two consecutive NAS boots
there, blocking the API bind forever with zero SQL activity. init_db
now latches per database URL (drop_db resets it; a different DB always
runs fully), and the alembic worker is bounded at 300s -- a wedged
thread fails startup loudly with a pinpointed error so the container
restarts into a clean retry instead of hanging silently.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
The PR-gate turn cut (#295) already auto-submitted assembled tasks, but
its refusals fell back to a PM spawn silently -- in production the PM
turns the cut was meant to remove kept happening with no visible cause
(live case: an AC-coverage refusal). The flag is gone (the fallback is
the safety net), the umbrella/branchless exclusion uses the canonical
batch predicates, and every refusal reason now rides into the spawned
PM's prompt so the fallback starts informed.
A live task burned 5+ hours because every exit was locked. unclaim now
works from verifying and needs_revision (service guard + lifecycle edge);
the circuit breaker and the i_am_done push-failure remediate name the
working chain ending in unclaim(); sync_branch(stash=true) clears the
DIRTY_WORKSPACE dead-end (pop-conflict preserves the stash); blocking a
task QA already owns now says to idle instead of listing states; the
orchestrator auto-block logs real errors and skips states where blocking
is meaningless instead of force-blocking them.
declare_coverage (cell/main PM) retroactively stamps parent-AC refs on a
child that implements them -- closing the roll-up deadlock where the
declaring child was cancelled and its re-delegated replacement completed
the work uncredited. Cancelling a ref-declaring child now warns and
surfaces the orphaned criteria.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* feat(sandbox): on-demand request_sandbox verb replaces eager provisioning
Sandboxes were provisioned at every agent spawn for opted-in projects,
so every role paid the sidecar spin-up and a provisioning failure
refused the spawn. Provisioning now happens when an agent asks: the
request_sandbox do-verb (dev + QA) reaches the orchestrator through
ContentActionsDeps, ensure_sandbox provisions idempotently with an
in-memory per-agent cache (evicted at teardown and janitor sweep), and
creds return in the envelope payload including ready-to-export
ROBOCO_TEST_* values. Spawn now only injects a marker env naming the
available services plus a briefing line; sandbox failures can no longer
refuse a spawn. Teardown lifecycle unchanged.
* feat(sandbox): harden request_sandbox + Phase 3 wiring proof and docs
Hardening from adversarial review: ensure_sandbox now provisions the
project's full opted-in set on first request (a later superset can
never tear down a live sandbox mid-use), serializes per-agent behind an
asyncio lock (a client timeout-retry no longer races its own in-flight
provision), and verifies container liveness on every cache hit (a dead
sandbox evicts and re-provisions instead of serving dead creds). MCP
client budget 720->1080s for the full-set cold case. Phase 3: e2e smoke
wiring test (manifest grants + guard-chain envelopes over the real
API), sandbox-db/tools/map docs and CLAUDE.md rewritten for on-demand.
* feat(sandbox): release sandboxes when the agent's work ends
CEO directive: sidecars must not dangle once the agent is done. The six
work-ending verbs (i_am_done, unclaim, i_am_idle, pass_review,
fail_review, i_documented) now release the caller's sandbox best-effort
on their success path via release_sandbox (lock + teardown + cache
evict; a no-sandbox agent costs a dict lookup). Container removal and
the janitor remain the backstop; a re-request provisions fresh.
* test(sandbox): monkeypatch the release hook instead of method assignment
mypy method-assign rejected the direct AsyncMock assignments; the prior
static gate ran before this test file landed.
* test(sandbox): guard envelope evidence for mypy in verb tests
* chore(prompts): regenerate verb tables for request_sandbox
* chore: resolve merge with master (breadcrumbs + statement budget)
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* fix(rag): per-index chunk floors — journals and learnings were never indexed
The global 200-char garbage floor (sized for code/doc chunks) discarded
every templated journal note and most distilled org-memory lessons,
silently: ingest returned success with zero chunks, so agent journals
and learnings were never retrievable via RAG. IndexConfig now carries a
per-type min_chunk_length (journals 40, learnings 80, others unchanged).
* fix(mcp): git-readonly tools default project_slug from the container env
Agents 404ed /api/git/status with 'Project not found: roboco' — the
tools made the LLM supply the slug and six doc examples taught a slug
that matches no registered project. The tools now fall back to the
ROBOCO_PROJECT_SLUG the orchestrator already injects, and the stale
examples are corrected.
* feat(rag): startup backfill re-ingests zero-chunk journals and learnings
Before the per-index chunk-floor fix, ingest() returned success with
chunk_count=0 for undersized content: every historical journal entry and
distilled learning below the (then-global) 200-char floor was durably
recorded in journal_entries but silently never got a chunks_journals /
chunks_learnings row, and no exception meant the existing dead-letter
(rag_index_failures) never saw it either.
Extends the startup reconcile (roboco/api/app.py _reconcile_rag_indexes)
with a new pass: backfill_unindexed_journals (roboco/services/
rag_index_failures.py) queries journal_entries for rows missing from each
vector table and re-ingests them through the same live code paths
(_reindex_journal_entry / record_learning).
Journals and learnings are backfilled independently since a LEARNING entry
can clear the (lower) JOURNALS floor while still failing the (higher)
LEARNINGS floor — a learning's doc_source is a content hash, not the entry
id, so presence there is checked by hashing each candidate the same way
LearningsIndexPlugin.record_learning does and batch-querying chunks_learnings
for those exact sources.
Bounded to 200 rows per pass per boot (converges over restarts on a larger
backlog) and best-effort per row (one failure never aborts the pass). Rows
still under the current floor are excluded by a length filter in the SELECT
so they are never retried forever, and private entries are excluded from
the JOURNALS pass exactly like the live indexing path.
* test(rag): scope backfill assertions to their own rows
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* fix(api): default the event loop to asyncio + cancellation-safe commit
The recurring CI e2e segfault traced to uvloop: the harness's
uvicorn.run() auto-selected it while production's serve() path never
consulted Config.loop (stock asyncio, accidentally safe). Every launch
site now resolves ROBOCO_UVICORN_LOOP (default asyncio; uvloop opt-in),
and DbCommitMiddleware's commit-in-send can no longer be interrupted
mid-wire: on cancellation it gets a bounded grace to finish (committed
data survives the 504), else invalidate-and-reraise.
* feat(runtime): expected-stop breadcrumbs attribute container deaths
Two production exit-143s had no attributable source: every orchestrator
kill path now records a short reason breadcrumb, and the exit monitor
consumes it -- an expected stop logs its reason at info, a genuinely
unexpected one logs none_recorded plus docker-inspect diagnostics
(OOMKilled, timestamps) so the next mystery SIGTERM self-identifies.
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
The sidecar's render.js was written against an older producer API and
never rendered on this deploy: it called createRenderJob({inputPath,
outputPath, width, height, fps}) and executeRenderJob(job, onProgress),
but 0.7.36 moved the paths to executeRenderJob(job, projectDir,
outputPath, onProgress) and takes only render params in the job config.
outputPath arrived undefined and dirname(undefined) threw on every
render (every composition), HTTP 500.
- Pass projectDir (the composition dir) + outputPath as executeRenderJob
args; select the cut via config.entryFile.
- Drop width/height — 0.7.36 reads dimensions from the composition HTML
(data-width/data-height); add the required quality tier.
- Stage motion/public into the composition dir: the producer serves the
compiled entry at the file-server root, clamping the composition's
../../public/fonts to /public/fonts, so fonts must sit under projectDir.
Verified by rendering both v0.19.0 cuts against the running sidecar's
0.7.36 libraries: styled MP4s, no asset 404s.
* fix(auth): pass agent UUID to CLI-arg MCP servers (optimal/docs/search)
The container token is HMAC-signed over the agent UUID (#314), but the
optimal/docs/search MCP servers received the slug as their CLI arg and
sent X-Agent-ID=<slug>, so every research/RAG/docs call 401ed with
signature mismatch under enforced auth. Pass the already-computed
agent_uuid in the three args lists instead.
* fix(gateway): include remediate in gateway.rejected audit details
Conventions-gate rejections carry the offending file:line listing only
in the envelope's remediate field, which the audit row dropped -- ops
logs showed just the violation count with no way to see what blocked.
* fix(gateway): return envelope on do/commit git failure
A GitError from the commit verb propagated to the generic middleware
handler, so agents got a raw error blob with no remediate/next. Catch
it and return an error envelope; 'no changes added to commit' with an
explicit files list now names the mismatch and the omit-files fallback.
* fix(agent-sdk): absolute rejection cap breaks slow-drip verb loops
The verb circuit breaker only counted rejections inside a 60s sliding
window, so an agent retrying i_am_done every 3-4 minutes looped for 30+
minutes without tripping it. Add a session-scoped cumulative per-(verb,
task) cap at 3x the windowed limit that trips regardless of pacing.
* feat(a2a): CEO chime-in interjects into the viewed conversation
Previously reply_as_ceo re-homed the message into a canonical CEO<->target
conversation with no panel surface, so a chime-in reported success but was
invisible and only opportunistically delivered. interject_as_ceo now inserts
the message into the conversation being viewed (from_agent=ceo, directed via
an @target content prefix), bumps that conversation's counters with the
unread ping keyed to the addressed participant, and both participants see it
in transcript and read_a2a.
* feat(panel): manual spawn carries task + message, surfaces refusals
The agent detail page spawned with no request body (task/message impossible),
the spawn button could double-fire (2.5ms double-POST seen live), and refusal
reasons never reached the UI: readiness refusals were generic 500s and the
already-running no-op looked like success. Detail page now uses
SpawnAgentDialog, a synchronous ref guard blocks re-entry, AgentReadinessError
maps to 409 with its reason shown, already_running is signalled and toasted,
and a task_id builds a task-aware prompt instructing the claim (task_id alone
never did), with the CEO's message appended as a note.
* test(panel): align a2a page test with the interjection footer copy
The chime-in rebuild changed the composer footer; the page-level test
asserting the old copy was outside the rebuild's scoped vitest run.
* fix(api): commit the request DB session before the response is sent
FastAPI unwinds yield-dependencies after the response bytes go out, so
get_db's post-yield commit raced the client's next request -- a verb
could return ok while its claim/status write was still uncommitted (the
e2e ok-without-effect flake family), and a failed commit was silently
lost behind an already-sent 200. DbCommitMiddleware (innermost, pure
ASGI) commits the session stashed by get_db_committed before forwarding
http.response.start; commit failure now surfaces as a 5xx. get_db is
untouched for its direct non-request callers.
* fix(db): invalidate, not rollback, the session on request cancellation
With the commit moved into the send path, the flow-verb timeout can
cancel mid-commit; rolling back then issues another command over an
asyncpg connection stranded mid-wire-protocol, and the poisoned
connection segfaults uvloop/asyncpg when a later checkout recycles it
(3/3 identical CI faulthandler dumps). On CancelledError discard the
connection via session.invalidate() -- SQLAlchemy's documented handling
for a timeout during commit -- and keep rollback for plain exceptions.
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
ReleaseExecutor.publish_release shelled out to 'gh release create', but no
Dockerfile installs the gh CLI (verified missing in the live orchestrator
container), so an armed release manager died at publish AFTER the release
commit was pushed. Publish now POSTs /repos/{owner}/{repo}/releases with the
project's decrypted token — same auth/httpx pattern as PR creation, same
fail-closed semantics (non-201 -> structured publish_failed, CEO retries;
the 300s deadline is the httpx client timeout). Subprocess publish-timeout
test replaced with REST-path tests (201/non-201/transport-error/no-token).
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
- agent-dev-fe / agent-qa-fe: remove playwright chromium + its system libs (~770MB each; verified unused — panel tests are vitest, e2e harness is scripted Python; pnpm kept)
- agent-grok: drop the redundant chown -R that duplicated the 149MB CLI tree into a second layer (install already runs as agent)
- agent-base + orchestrator runners: split the single /app COPY into .venv-first / source-last layers so a source-only deploy re-layers ~13MB instead of ~380MB per image
- agent-base: split the 813MB apt+node+claude-code RUN so a CLI bump no longer re-downloads the OS/node layer
- .dockerignore: exclude gitignored docs/internal from the orchestrator's docs COPY; pin uv helper image to 0.11
- verified: all four images rebuilt + runtime-probed (claude/git/jq/node/uv/pnpm/grok, import roboco, docs/alembic/agents present); .venv layer proven CACHED across a source-only change
- mongo:8-alpine → mongo:8 (tag never existed; a mongo-opted project could spawn no agents) + a Docker Hub tag-existence e2e guard for every sandbox engine
- flow-verb timeouts at both walls: shared SLOW_VERBS policy (i_am_done / submit_up / submit_root / open_pr / i_will_work_on get the 900s server budget); the MCP client now outlasts the server budget (+10s headroom, orchestrator-injected env) so agents receive the middleware's clean 504 envelope instead of dying at the old flat 30s client timeout
- cancellation safety: the quality gate kills+reaps its child on CancelledError; create_pr records the PR via a shield-with-wait-out helper so the write can neither be skipped nor race get_db's rollback
- video engine: renderer sidecar isolated on a render-only network, 2g/2cpu caps, 570s render watchdog with exit-on-hang, 512MB tar decompression cap, CEO notification on terminal render failure, reject under the approve mutex (fail-closed on Redis-down)
- dead python-jose dependency removed (drops ecdsa and its unfixable Minerva advisory PYSEC-2026-1325); panel --font-mono now a real monospace stack
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* feat(tasks): task-content guardrails — structured plans + constraints split
Bound task PLANNING content the way journals/notes already are, fixing the
poor task quality flagged 2026-07-07 (degenerate roots, over-decomposed
leaves, descriptions bloated by an auto-attached conventions dump).
Phase A — plan/AC guardrails (no migration):
- _pm_sub_tasks_gate: cap sub_tasks at 7; per-subtask ceilings (title <=200,
description <=600) enforced at both the Pydantic boundary and the gate.
Dropped the min-2-roots and no-subtasks-on-code rules: both contradict the
2026-05-08 rule (test_cell_pm_can_plan_code_typed_parent_via_i_will_plan)
and break legitimate single-cell roots. Long comment in the gate explains.
- IWillPlanRequest: plan <=2000, approach <=800 (floor 150 kept), typed
SubTaskCreate/RiskCreate/OpenQuestionCreate replacing loose list[dict].
- DelegateRequest + task_completeness: acceptance_criteria capped at 7 items,
each <=200 chars. New FieldRule.MAX_LENGTH_LIST + _post_rule_reject helper
(extracted to keep the gate under xenon B).
- Routes dump typed models to dicts for the existing rich_plan shaper.
Phase B — conventions split (migration 068):
- New nullable tasks.constraints Text column; _attach_baseline_constraints
now writes the ## Constraints block there instead of appending to
description, so description is the human-authored instruction only. The
conventions still reach the agent independently at spawn via the ambient
block, so agent correctness is unaffected.
- TaskResponse / Task model / panel Task type carry constraints; panel shows
a read-only Constraints card. Field is optional on the TS type (backend
returns null for flag-off / pre-migration rows).
Tests: 5 new gate unit tests, 7 schema tests, 3 AC policy tests, 3 e2e smoke
scenarios; 4 baseline-constraints integration tests updated. ruff/mypy/xenon
clean; 10026 unit+foundation+e2e green; panel typecheck clean.
Refs: plan breezy-imagining-kahn
* test(tasks): use typed SubTaskCreate instead of dict literals in plan tests
make quality runs mypy over tests/ (1079 files), not just roboco/ — the
four sites passing dict literals to the now-typed sub_tasks: list[SubTaskCreate]
field failed mypy. Construct SubTaskCreate directly; the typed model raising
ValidationError IS the boundary the rejection tests assert.
* fix(deps): drop unused python-jose — clears PYSEC-2026-1325 (ecdsa, no fix)
CI's pip-audit went red on a freshly-published advisory PYSEC-2026-1325
against ecdsa 0.19.2 (no fix published — 0.19.2 is the latest). ecdsa is a
transitive dep of python-jose, which is a DIRECT dep of roboco but is NOT
imported anywhere in roboco/ or tests/ (grep-verified). The actual JWT path
uses PyJWT (import jwt) + fastapi_users.jwt, not python-jose.
So python-jose is a dead dependency. Removing it (deletion over an
--ignore-vuln waiver) drops ecdsa + rsa + pyasn1 + their type stubs from the
lockfile, eliminating the CVE at the source. deptry roboco/ stays clean
(no missing-dep), mypy clean, auth + schema tests pass.
Master CI was green 9h before this PR's run, so the advisory published in
that window would red any run including master — this fix unblocks both.
* chore(prompts): regenerate verb tables for typed plan sub_tasks
Phase A's IWillPlanRequest schema change (sub_tasks/risks/open_questions from
loose list[dict] to typed SubTaskCreate/RiskCreate/OpenQuestionCreate) made
the auto-generated verb tables stale. Regenerated via
scripts/regenerate_verb_tables.py — the diff is purely the signature
reflection (list[str|str] -> list[SubTaskCreate], etc.). Required by the
foundation-check gate (Makefile:559).
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* feat(sandbox): pluggable per-engine registry (postgres/redis/mongo)
Replaces the hardcoded postgres+redis branches in the provisioner and the
env emitter with a registry of SandboxEngine specs (image, run args,
readiness probe, connection, ROBOCO_TEST_* env) in a pure low module
(roboco/models/sandbox.py). VALID_SANDBOX_SERVICES is derived from the
registry — single source of truth — and the provisioner + orchestrator
iterate it, so adding an engine is one class + one registry line, not
another branch. Adds a mongo:8-alpine engine (ROBOCO_TEST_MONGO_*) as the
third service alongside postgres/redis.
Also fixes the cold-pull loop that stranded v0.19.0 board agents with
empty error strings: docker run pulled inline under a 20s deadline, so a
NAS cold pull was killed, cancelled, and re-pulled from scratch forever.
_ensure_image now inspects + pulls (300s) before run; provisioning errors
log type+message so a bare TimeoutError no longer shows as "".
Panel edit-project dialog: postgres/redis toggles -> a Set<string>
multi-select driven by a SANDBOX_SERVICES catalog, so new engines appear
in the UI by adding to the catalog.
Tests: engine parity (allowlist==registry, unique slugs/images, no None
leak in env, SandboxInfo aggregates every engine), mongo provision + env
injection, plus the existing postgres/redis provision/env/spawn/janitor
suite updated to the registry shape. 821 unit / 5 skip green; ruff + mypy
(360 files) clean.
* docs(sandbox): reflect pluggable engine registry + mongo across docs
CHANGELOG (0.19.0): Added entry for the pluggable sandbox engine registry
(postgres/redis/mongo) + Fixed entry for the cold-pull loop/empty-error
strand that boarded v0.19.0 board agents.
docs/map (9 files): sandbox subsystem blurbs, SandboxProvisioner rows,
_maybe_provision_sandbox/_append_sandbox_env rows, feature-flag rows, the
migration-057 row + v0.17.0 delta, and the models.md VALID_SANDBOX_SERVICES
note — all retitled to DB/Redis/Mongo via the engine registry
(roboco/models/sandbox.py), with the one-class-one-line extension story and
the _ensure_image cold-pull fix. Production-network (roboco_data) lines left
as postgres+redis — mongo is sandbox-only, not a prod service.
docs/rag (3 files): sandbox-db.md rewritten around the registry (engine list,
generic _provision_engine, image pre-pull, ROBOCO_TEST_DB_*/REDIS_*/MONGO_*
incl. MONGO_AUTH_DB=admin, single emit_env); config-reference sandbox flag
row + subsection retitled; db-network-isolation framing broadened to
postgres/redis/mongo. preconditions-and-rejections left untouched (its hit
was an unrelated gateway see-also link).
* test(e2e): harden umbrella close terminal reads with bounded wait-for-state
The MegaTask umbrella close test flaked once on CI (ceo-approve returned
200 but the re-fetch saw awaiting_pm_review) then passed on re-run. The
production path is deterministic: complete -> main_pm_complete ->
submit_pm_review -> escalate_to_ceo -> ceo_approve -> commit, all on one
session, all awaited; the fire-and-forget completion hooks are isolated
(own session, best-effort, never touch task.status or the request session).
20 local runs could not reproduce it.
The one real surface is the read pattern: the e2e stack commits on the
uvicorn thread's loop and reads via a separate loop (run_db -> asyncio.run
with a fresh engine), so a terminal single point-read can race a
still-draining completion hook on a contended runner. Replace the two
terminal point-reads with a bounded wait_for_status poll. Strictly better
than a one-shot read: absorbs the transient, and a genuine state bug still
surfaces via the timeout branch asserting against the last-read state.
* fix(gateway): bound hung flow-verbs with a server-side timeout
A gateway intent-verb whose request transaction held the SELECT ... FOR
UPDATE lock on the task row never committed: uvicorn does not cancel the
endpoint coroutine on client disconnect and get_db only rolled back on
Exception (not a hang/cancellation), so the row lock was held indefinitely
and every later task-row write on that task wedged (2026-07-07
kimi-k2.7-code:cloud agent on task 79d686f0). Reads (evidence) and journal
writes (note) stayed fast — the symptom that pointed at a task-row lock.
Fix: pure-ASGI FlowVerbTimeoutMiddleware wraps each /api/v1/flow/* request
in asyncio.timeout(flow_verb_timeout_seconds, default 120s). On expiry the
inner app is cancelled; CancelledError now propagates through get_db (which
catches it alongside Exception and rolls back), releasing the FOR UPDATE
lock, and a retryable 504 gateway_timeout envelope is returned. Pure ASGI
(not BaseHTTPMiddleware) so cancellation reaches the route coroutine +
get_db dependency directly, with no spawned-task gap. Registered innermost
so correlation + logging still wrap the 504.
E2E: two fault-injection scenarios in tests/e2e_smoke/test_flow_verb_timeout.py.
A hang is injected inside the verb's own transaction (claim acquires the
FOR UPDATE lock, then set_plan sleeps past the timeout; only the first
set_plan call runs — a retry short-circuits as idempotent re-entry).
- ARMED (server timeout 1s): verb-1 returns a bounded 504 gateway_timeout,
verb-2 re-acquires the row and reaches the post-claim gate (tracing_gap)
— proving the lock was released by verb-1's cancellation.
- DISARMED (server timeout 1000s, MCP client timeout 3s): verb-1 holds the
lock past the client's HTTP timeout — the empirical reproduction of the
wedge on the same branch, by turning the fix off.
Full e2e suite green (32 passed).
* feat(sandbox): pluggable per-engine registry (postgres/redis/mongo) (#324) (#325)
* feat(sandbox): pluggable per-engine registry (postgres/redis/mongo)
Replaces the hardcoded postgres+redis branches in the provisioner and the
env emitter with a registry of SandboxEngine specs (image, run args,
readiness probe, connection, ROBOCO_TEST_* env) in a pure low module
(roboco/models/sandbox.py). VALID_SANDBOX_SERVICES is derived from the
registry — single source of truth — and the provisioner + orchestrator
iterate it, so adding an engine is one class + one registry line, not
another branch. Adds a mongo:8-alpine engine (ROBOCO_TEST_MONGO_*) as the
third service alongside postgres/redis.
Also fixes the cold-pull loop that stranded v0.19.0 board agents with
empty error strings: docker run pulled inline under a 20s deadline, so a
NAS cold pull was killed, cancelled, and re-pulled from scratch forever.
_ensure_image now inspects + pulls (300s) before run; provisioning errors
log type+message so a bare TimeoutError no longer shows as "".
Panel edit-project dialog: postgres/redis toggles -> a Set<string>
multi-select driven by a SANDBOX_SERVICES catalog, so new engines appear
in the UI by adding to the catalog.
Tests: engine parity (allowlist==registry, unique slugs/images, no None
leak in env, SandboxInfo aggregates every engine), mongo provision + env
injection, plus the existing postgres/redis provision/env/spawn/janitor
suite updated to the registry shape. 821 unit / 5 skip green; ruff + mypy
(360 files) clean.
* docs(sandbox): reflect pluggable engine registry + mongo across docs
CHANGELOG (0.19.0): Added entry for the pluggable sandbox engine registry
(postgres/redis/mongo) + Fixed entry for the cold-pull loop/empty-error
strand that boarded v0.19.0 board agents.
docs/map (9 files): sandbox subsystem blurbs, SandboxProvisioner rows,
_maybe_provision_sandbox/_append_sandbox_env rows, feature-flag rows, the
migration-057 row + v0.17.0 delta, and the models.md VALID_SANDBOX_SERVICES
note — all retitled to DB/Redis/Mongo via the engine registry
(roboco/models/sandbox.py), with the one-class-one-line extension story and
the _ensure_image cold-pull fix. Production-network (roboco_data) lines left
as postgres+redis — mongo is sandbox-only, not a prod service.
docs/rag (3 files): sandbox-db.md rewritten around the registry (engine list,
generic _provision_engine, image pre-pull, ROBOCO_TEST_DB_*/REDIS_*/MONGO_*
incl. MONGO_AUTH_DB=admin, single emit_env); config-reference sandbox flag
row + subsection retitled; db-network-isolation framing broadened to
postgres/redis/mongo. preconditions-and-rejections left untouched (its hit
was an unrelated gateway see-also link).
* test(e2e): harden umbrella close terminal reads with bounded wait-for-state
The MegaTask umbrella close test flaked once on CI (ceo-approve returned
200 but the re-fetch saw awaiting_pm_review) then passed on re-run. The
production path is deterministic: complete -> main_pm_complete ->
submit_pm_review -> escalate_to_ceo -> ceo_approve -> commit, all on one
session, all awaited; the fire-and-forget completion hooks are isolated
(own session, best-effort, never touch task.status or the request session).
20 local runs could not reproduce it.
The one real surface is the read pattern: the e2e stack commits on the
uvicorn thread's loop and reads via a separate loop (run_db -> asyncio.run
with a fresh engine), so a terminal single point-read can race a
still-draining completion hook on a contended runner. Replace the two
terminal point-reads with a bounded wait_for_status poll. Strictly better
than a one-shot read: absorbs the transient, and a genuine state bug still
surfaces via the timeout branch asserting against the last-read state.
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>