* feat(x): redraft loop on CEO reject — feedback re-enters the draft flow
A rejected X draft's reason used to die with the cancel. reject() with a
non-blank reason now schedules a redraft after its commit
(defer_after_commit; fresh session; never blocks or fails the HTTP
response): XEngine.redraft_from_rejection re-drafts the same source kind
via the local model with the reason and rejected body folded in as
revision guidance, originating ONE fresh held draft — mirroring the
video pipeline's reauthor_from_rejection. Local-model failure or empty
output originates nothing (no degraded copies); markers carry forward
whole so a redrafted reply/spotlight stays fully functional downstream;
bodies ride the same 280 clamp; the open-posts cap holds.
Hardened per adversarial review: reject() is now idempotent on an
already-CANCELLED target at both check sites (mirroring approve's
already_rejected guard — a replayed reject schedules nothing), and the
dedup check+originate runs under a non-blocking identity-keyed Redis
lock (SET NX + compare-and-del, matching the approve/reject mutex
style) so racing rejects can't stack duplicate drafts — lock held or
Redis down skips the redraft, which is always safe. Tests pin the
fresh-session contract by session identity, the replay no-op, the
lock-skip, and clean up their own committed rows.
* fix(tests): runtime UUID import + typed task-id coercion in x cleanup helper
CI's quality gate runs mypy over tests/ (the local pass covered only
roboco/): the _delete_tasks calls handed ORM-typed ids where uuid.UUID
was expected. Coercing at the call sites then exposed that UUID was
imported under TYPE_CHECKING only — a runtime NameError. Import moved
to runtime; both call sites coerce explicitly.
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
Root cause first: no viewport export existed anywhere, so viewport-fit
was never 'cover' and every env(safe-area-inset-*) resolved to 0 on
notched iPhones — content under the status bar, dock without real
home-indicator clearance. The export lives on the (tg) group layout
(server component), NOT app-wide: the dashboard shell has no safe-area
padding and must not inherit cover.
Also: min-w-0 on four truncating flex children that overflowed their
justify-between rows (chat names/previews, fleet task titles);
object-contain on the approvals video (letterbox instead of distort on
short phones); break-words on the changelog pre / task description /
quoted mention; touch targets bumped to >=36px (sheet close, segmented
controls, ack button, chips, back button, bell, jump-to-latest,
cut-toggle); overflow-x-hidden backstop on the (tg) main scroller; fleet
avatar strip sliced to 3 with a +N badge instead of silent clipping.
Verified: pnpm typecheck clean, lint 0 errors, panel suite 870/870.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
fastapi-guard peels a fixed trusted_proxy_depth=1 from X-Forwarded-For
(the rightmost entry, which nginx itself recorded). That is correct for
every chain except host-proxied tailnet traffic (Tailscale Serve →
nginx), which arrives as [tailnet-client, loopback-or-bridge-gateway] —
depth-1 resolves it to a whitelisted hop IP, leaving WAF/ban/rate-limit
inert for the whole /tg surface (the documented ceiling).
ClientIpResolutionMiddleware (pure ASGI, wraps SecurityMiddleware so it
runs first) stamps guard_core's request.state.client_ip cache — its
supported pre-resolution seam — for EXACTLY that shape: peel known local
hops (loopback + docker bridge pool) from the right, stamp only when at
least one hop was peeled AND the candidate is in the tailnet CGNAT range
(100.64.0.0/10). Every other shape abstains, so direct LAN clients,
agent containers relaying through nginx (even with forged public-IP
prefixes), and all-hops operator traffic resolve byte-for-byte as
before. Documented residual: a same-bridge container forging a CGNAT
prefix only DE-privileges itself (loses its whitelist exemption). XFF is
read first-occurrence to match Starlette's own header semantics, and a
wiring test pins the middleware mount ORDER, not just presence.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
The pg_dump sidecar wrote its dumps to the same disk it protects — one
disk failure lost both. Setting ROBOCO_BACKUP_MIRROR_DIR in .env to a
path on a different disk (external/remote mount) arms a mirror step after
every successful dump: tmp+rename copy, mirror pruned to the same
BACKUP_KEEP, unwritable mirror logs-and-skips without blocking the
primary. Unset, the script never attempts a copy — no fake off-disk
copies on the same disk. Docs gain the mirror setup and a quarterly
restore drill (throwaway pgvector container, pg_restore, row-count
sanity check).
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* feat(lifecycle): inherit advanced upstream base on work re-claims
A re-claim reused a branch cut at an earlier claim, so upstream work
merged since (UX/UI landing on the root after the cell branch was cut)
never reached BE/FE branches — divergence and avoidable conflicts.
_finalize_claim now merges the advanced base into the pre-existing branch
via the dependency-lineage merge: already-ancestor is a no-op, a conflict
aborts at the cut point and leaves a transition note steering the agent
to sync_branch, a clean merge logs an audit trail; never fails the claim.
Double-gated: by role (developer/cell_pm/main_pm — QA/documenter/gate
claims review the branch as pushed and never move it) AND by pre-claim
status (pending/needs_revision only — a PM's i_will_plan re-claim of its
own awaiting_pm_review task must not move a branch that already passed
QA + the PR gate). Fresh cuts already branch from the live remote base.
Cell-PM prompt now orders reading the upstream design docs before
planning.
* fix(lifecycle): extract base-inheritance gate predicate for xenon budget
The four-condition inline gate pushed _finalize_claim to cyclomatic
rank C; the quality gate caps blocks at B. The decision moves to a pure
module-level predicate, byte-for-byte the same logic.
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
The escalation/approval dispatchers spawn a notification's recipient every
cooldown window for as long as it stays pending. These spawns carry no
task_id, so the PM respawn breaker never sees them — a single wedged
alert/escalation whose recipient never resolves it respawns that recipient
forever. Observed live: fe-pm's unacked alerts kept main-pm/fe-pm spawning
every ~2-3 min for 6+ hours.
Two guards, both gating the spawn after the existing cooldown:
- A hard per-(agent, notification) attempt cap (notification_spawn_max_attempts,
default 5): once a notification has respawned its target that many times
without being acknowledged, stop and log once. The count is id-scoped and
survives map pruning (re-stamp), so a fresh escalation is unaffected.
- A live-work check before spawning: skip when the notification has expired,
is stale past notification_spawn_max_age_seconds (default 6h — wedged or
reloaded from before a restart), or its related task is already terminal.
Fail-open — a failed task fetch or unparseable field never suppresses a
real escalation.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* [943d8c4d] Frontend data freshness and approval-queue reliability audit (#631)
* [233a8b0f] WebSocket reconnect message-loss audit and fix (#625)
* [233a8b0f] fix(panel): add REST catch-up to useNotificationStream on WS reconnect
connection.ts has no message buffering/replay, so a notification published
while the CEO bell's socket was down (disconnected/reconnecting) was lost
forever instead of merely delayed. Add a reconnect-triggered GET
/notifications?unread_only=true catch-up folded into the existing
notification_id dedup so a notification delivered both via catch-up and
live WS is never double-counted, and make clearMessages drop the held
catch-up batch too. use-a2a-live.ts and use-rate-limit-websocket.ts were
audited and already have working reconnect-triggered REST fallbacks
(verified via a2a/page.tsx, rate-limit-banner.tsx, usage-overview-panel.tsx
and their existing F083 tests) so no fix was needed there.
* [233a8b0f] docs(panel): add comprehensive WebSocket hooks reference and reconnect architecture guide
Add panel/docs/frontend/hooks.md with full API reference for useWebSocket, useNotificationStream (with new REST catch-up behavior), useAgentStream, useA2ALiveStream, and useConnectionStatus. Include examples, best practices, and testing guidance.
Add panel/docs/architecture/websocket-reconnect.md documenting the message-loss mitigation pattern: Strategy 1 (REST catch-up for events, used by useNotificationStream) and Strategy 2 (REST invalidation for state, used by A2A/rate-limit consumers), plus the dedup logic ensuring no notification is double-counted on reconnect.
---------
Co-authored-by: Frontend Developer 1 <fe-dev-1@roboco.tech>
Co-authored-by: Frontend Documenter <fe-doc@roboco.tech>
* [d5315683] fix(frontend): add distinct toast feedback for silently-swallowed x-post and release-proposal statuses, plus regression tests for all 4 approval queues (#626)
Co-authored-by: Frontend Developer 2 <fe-dev-2@roboco.tech>
* [cd953838] Data-hook null-guard audit and API client 429 retry-by-method fix (#630)
* [cd953838] fix(panel): gate 429 retry by HTTP method, add hook null-guard regression tests
* [cd953838] chore(conventions): waive test-fixture wrapper in hooks null-guard test
* [cd953838] docs(frontend): document API rate-limit retry behavior and null-guard audit results
Added `docs/frontend/api-rate-limiting.md` to document the 429 retry strategy: GET/PUT auto-retry, POST/PATCH/DELETE require X-Idempotency-Key header. Updated `docs/frontend/hooks.md` to confirm the data-hook null-guard audit found all hooks already have correct `enabled` guards and include a regression test suite for the board-review poll on/off behavior and enabled-guard assertions.
---------
Co-authored-by: Frontend Developer 1 <fe-dev-1@roboco.tech>
Co-authored-by: Frontend Documenter <fe-doc@roboco.tech>
---------
Co-authored-by: Frontend Developer 1 <fe-dev-1@roboco.tech>
Co-authored-by: Frontend Documenter <fe-doc@roboco.tech>
Co-authored-by: Frontend Developer 2 <fe-dev-2@roboco.tech>
* [4534c71a] Backend concurrency, state-machine, and engine audit (#634)
* [41de844a] fix(lifecycle): sync CLAIM_RULES with runtime + clear stale claimant on PM hand-off (#627)
Two confirmed state-machine gaps found while auditing lifecycle.py,
task_lifecycle.py, the _ESCALATABLE_TO_BLOCKED bypass, and every
_REVIEW_QUEUE_STATES entry point:
- lifecycle.py's CLAIM_RULES/claim-ActionSpec/StatusTransition table
did not grant CELL_PM/MAIN_PM re-claim of AWAITING_PM_REVIEW even
though task.py's runtime _ROLE_CLAIM_STATUSES already granted it
and claimed the spec agreed -- the two tables had silently drifted,
breaking i_will_plan re-claim on an awaiting_pm_review task.
- docs_complete's _maybe_advance_to_pm_review pre-assigns a specific
owning PM via assigned_to but left claimed_by/active_claimant_id
pointing at the outgoing documenter, unlike every sibling transition
into a review-queue state. A stale active_claimant_id makes
content_actions.py's _active_claim_violation wrongly reject the
newly-assigned PM's own content writes before it formally claims.
Reassign claimed_by + active_claimant_id to the owning PM alongside
assigned_to.
Adds a regression test asserting the documenter's stale claim does not
survive the docs_complete -> awaiting_pm_review hand-off.
Co-authored-by: Backend Developer 2 <be-dev-2@roboco.tech>
* [0c46f666] Engine dedup race + sequencing.py edge-case audit (#628)
* [0c46f666] fix(sequencing): dedup race audit + collision-edge fallback bug
Audited the list-open-then-originate dedup pattern across six engines:
RoadmapEngine, XEngine.run_cycle, DepUpdateEngine, and CIWatchEngine each
run inside exactly one sequential orchestrator-loop asyncio task (no other
call site invokes run_cycle), so they cannot race with themselves; their
in-cycle dedup sets/keys are correctly built before any commit. SelfHealEngine
is the same shape. VideoEngine.open_video_task is genuinely different: it is
reachable from the release-publish hook, the feature-spotlight hook, and the
on-demand POST /video/request route, so two overlapping calls for the same
occasion can both pass the "no open task yet" check before either commits.
Fixed by wrapping the check+insert in a short-lived Redis mutex (reusing
HeartbeatMutex) keyed by occasion, mirroring XPostService's existing
lock pattern, with a regression test proving only one of two concurrent
calls creates a task.
Verified ReleaseExecutor's half-landed retry path (release_commit_sha):
apply_version_bumps and write_changelog_entry both run as uncommitted
working-tree edits before commit_and_push's single `git add -A` + commit,
so a bumped-version-without-changelog state can never reach origin (and
therefore can never be observed by a fresh retry clone) - confirmed correct
with a real-git-repo regression test, no fix needed.
Fixed sequencing.py's dev_task_collision_edges: the `if edges: return edges`
short-circuit dropped the same-assignee-lane fallback entirely whenever ANY
surfaced sibling pair produced a collision edge, even for a completely
unrelated same-assignee pair with no declared surface. Now the fallback
always runs, skipping only pairs the analyzer already ordered (so the two
mechanisms can never disagree on direction for the same pair).
Verified sequencing.py rule 3 (all-shared batch generates no edges): correct
by inspection (_shared_last_edges skips every pair when both are shared) and
confirmed with a regression test - no fix needed.
* [0c46f666] docs(reference): concurrency audit summary - engine races, fixes, verified patterns
---------
Co-authored-by: Backend Developer 2 <be-dev-2@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>
* [8f7f167a] Redis mutex pre-lock write audit (#629)
* [8f7f167a] Redis mutex pre-lock write audit: add cross-session regression test for XPostService.approve
Audited x_post_service.py, video_post_service.py, release_proposal.py, and
heartbeat_mutex.py for the pre-lock DB-write anti-pattern (a session write
that happens before the SET NX / HeartbeatMutex acquire returns a token,
letting a losing racer's stale write clobber a winner's committed state).
XPostService.approve, VideoPostService.approve, and
ReleaseProposalService.approve/reject already implement the correct
validate-pure-pre-lock, apply-under-lock pattern (the XPostService fix
already shipped per CHANGELOG.md: "X edited_body write deferred into the
single-flight lock (M5)"). HeartbeatMutex holds no AsyncSession at all, so
the anti-pattern is structurally inapplicable there.
Adds a genuine cross-session concurrency regression test to
test_x_post_service.py (a real second DB connection, not an in-process
mock) mirroring VideoPostService's existing cross-session test, proving a
concurrently-committed post survives and the CEO's edited body never lands
on the just-posted row.
* [8f7f167a] Remove redundant inline comments flagged by QA in cross-session regression test
Both comments restated what the surrounding docstrings already say
explicitly, per QA findings F-dbadd8f0 (line 294) and F-27ac051e (line
631) — no behavior change, tests re-verified green against a sandbox
Postgres.
* [8f7f167a] Remove inline trailing comments flagged by QA (correct file this time)
QA findings F-e6f3e6a6 and F-24189858 cited tests/unit/services/
test_x_post_service.py:294 and :631 across 5 revision rounds, but that
file never contained the flagged comment text — a repo-wide grep for
the exact quoted strings shows both comments actually live in the
mirrored tests/unit/services/test_video_post_service.py file, in its
own cross-session concurrency regression tests (the caption-edit and
tiktok-skip tests). Removed both there:
- "# externally visible to the "concurrent" session below" on the
db_session.commit() call
- "# never attempted without credentials" on the tiktok_poster.calls
assertion
Both restated what the surrounding docstrings/test names already say;
no behavior change. Verified with the full make quality gate against a
sandbox Postgres/Redis: 13,717 passed, 94.41% coverage, clean except
one pre-existing unrelated failure in tests/unit/api/test_cloud_auth.py
::test_login_route_parses_oauth2_form_not_query_params, which connects
to the app's default localhost:5432 Postgres (not the db_session
sandbox fixture) and is unreachable in this sandboxed environment —
structurally unrelated to the auth subsystem this task never touches.
* [8f7f167a] Redis mutex pre-lock write audit (round 7): add cross-session regression tests for reject() lock protection
Round-7 QA findings F-7eb9fbcb, F-06f39a2e, and F-4d56e49b claim
XPostService.reject(), ReleaseProposalService.reject(), and
release_executor._await_proc() lack lock protection / a CancelledError
handler — but their cited line ranges (255-267, 429-454, 241-257)
describe a pre-fix, shorter version of these functions that predates
commit fb293a787d, already on this branch. At current HEAD:
- x_post_service.py reject() (lines 275-299) acquires _LOCK_PREFIX,
re-reads under the lock, applies markers.set_x_reject_reason() +
CANCELLED only inside the critical section, releases in finally.
- release_proposal.py reject() (lines 460-486) does the identical
dance with _RELEASE_LOCK_PREFIX.
- release_executor.py _await_proc() (lines 257-265) already has an
`except asyncio.CancelledError` block that kills + reaps the child
and re-raises, mirroring the TimeoutError handler, with an existing
dedicated regression test
(test_await_proc_kills_child_on_outer_cancellation).
The one genuine gap: neither reject() path had a cross-session
(real second DB connection, not an in-process mock) regression test
proving the in-lock re-read catches a concurrent approve/publish that
completes mid-lock-wait — only approve() had one. Added
test_reject_concurrent_approve_completes_during_lock_wait to both
test_x_post_service.py and test_release_proposal_status_guards.py,
mirroring the existing approve() cross-session test: a second engine
commits COMPLETED between reject's pre-lock read and lock acquisition,
and the test asserts the CANCELLED write / reject-reason marker never
lands on the just-completed row.
No production code changed — verified via 103 targeted tests green
against a sandbox Postgres/Redis, plus `make -o sync gate` clean.
* [8f7f167a] Regenerate stale lifecycle artifacts (restore auditor waive_finding)
foundation-check was the only failing gate: the committed lifecycle artifacts
were missing the auditor's waive_finding verb that the lifecycle source
defines, so make quality regenerated them and failed on the diff — nothing to
do with the mutex fix (which passes ruff/mypy/tests/coverage/bandit clean).
make lifecycle restores the drift; this is what the 8 revision rounds kept
missing.
---------
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* [d615f2e3] fix(tests): sync stale CLAIM_RULES pinning assertions with lifecycle.py (#635)
test_claim_rules_match_pre_gateway_table still asserted the pre-audit
two-member frozenset for CELL_PM/MAIN_PM claim rules. CLAIM_RULES in
lifecycle.py already grants both roles claim rights on
Status.AWAITING_PM_REVIEW (added by the state-machine exhaustiveness
audit) so a PM can re-claim its own review-queue task after a respawn.
Updated both assertions to include AWAITING_PM_REVIEW, matching the
actual dict. Grepped the repo for sibling stale copies of the old
literal; found none beyond this test.
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
* [3c4e7a35] fix(quality-gate): reflow CONCURRENCY_AUDIT.md and stub occasion lock in video tests (#636)
Root cause: PR #634's CI failed at the markdown-reflow check (make quality
Makefile:285) on CONCURRENCY_AUDIT.md — a hard-wrapped audit doc left over
from the merged "Engine dedup race + sequencing.py edge-case audit" unit
(PR #628). Fixed with `make reflow-docs` (the exact remedy the CI output
itself named).
Running the full local `make quality` (with a sandbox Postgres/Redis to
get past DB-gated skips) surfaced a second real regression from that same
PR #628 unit: it added a Redis-backed HeartbeatMutex occasion lock to
VideoEngine.open_video_task, but two pre-existing test files
(tests/unit/runtime/test_video_render_loop.py and
tests/integration/test_video_routes.py) call open_video_task without
stubbing that lock, so they failed closed against the suite's
deliberately-unreachable test Redis (_no_live_redis). Fixed by applying
the same lock-stub pattern tests/unit/services/test_video_engine.py
already uses for its own occasion-lock tests: an autouse HeartbeatMutex
stand-in fixture in test_video_render_loop.py, and wrapping the two
route-level video-request tests in test_video_routes.py with the file's
existing _LOCKED patch pair (already used by every other lock-dependent
test in that file).
The one remaining local failure,
test_cloud_auth.py::test_login_route_parses_oauth2_form_not_query_params,
is a pre-existing environment gap unrelated to this branch: it needs a
real Postgres reachable at localhost:5432 (which .github/workflows/ci.yml
provides as a service container) but this dev sandbox has no such binding
— confirmed unrelated to any of the four merged audit units.
make quality now passes clean: 13729 passed, 0 regressions, 94.49% coverage.
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
* [ddc8121f] regenerate lifecycle artifacts for awaiting_pm_review claim rules and waive_finding intent (#637)
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
---------
Co-authored-by: Backend Developer 2 <be-dev-2@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* fix(sequencing): drop lane-fallback edges that would cycle against analyzer edges
The dev-task collision fallback unioned the analyzer's authoritative edges
with same-assignee lane edges, deduping only the direct pair. A lane chain
through an unsurfaced middle sibling could still contradict an analyzer edge
transitively (the shared-last migration order inverts plain priority order),
closing a 3-cycle that made add_dependency raise ConflictError and wedged
every later delegate to that parent. Fallback edges are now accepted only
when they can't close a cycle against the edges already kept; a regression
test reproduces the exact scenario.
Also strip pre-merge cruft: remove the root CONCURRENCY_AUDIT.md working
report, delete the near-duplicate websocket-reconnect.md doc, fix the stale
a2a/page.tsx doc citation, correct the api-rate-limiting doc to state
idempotency-key retry is unimplemented, and fix two lifecycle.py comments
that referenced a guard function which never existed.
---------
Co-authored-by: Frontend Developer 1 <fe-dev-1@roboco.tech>
Co-authored-by: Frontend Documenter <fe-doc@roboco.tech>
Co-authored-by: Frontend Developer 2 <fe-dev-2@roboco.tech>
Co-authored-by: Backend Developer 2 <be-dev-2@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* fix(panel): A2A transcript/list poll as fallback when the socket drops
The desktop A2A view refreshed ONLY on /ws/system a2a.message frames —
refetchInterval was off (only the /tg mini app polled). So when the socket
flaps (the NAS stack flaps often), the open transcript froze: agent replies
never landed and the thread stuck on the last frame received, even though the
messages persisted server-side. useA2AMessages and useA2AConversations now
take a 10s REST poll gated on the live-stream connection — polls only while
disconnected, never when the WS is healthy, so it's a true fallback with no
wasted requests. Backend was fine (get_messages_admin returns the full
transcript incl. CEO interjects; the frame carries the right conversation_id).
* fix(panel): make the A2A poll an unconditional backstop, not disconnect-gated
Live-checked the NAS: /ws/system connections stay open (10 opens / 0 closes in
30m) and events publish — the socket is NOT flapping, so a disconnect-gated
poll wouldn't fire. The real freeze mode is a silent-dead / half-open WS that
still reports readyState OPEN (no client keepalive ping detects it), where
isConnected stays true. So poll unconditionally: 20s while the socket claims
up, 8s once known-down. Guarantees liveness regardless of why a frame didn't
land.
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
#621 only let a NEW project bind to a GitHub App installation (the create
dialog's Select repo picker). An already-imported project on a PAT had no way
to re-route to the App. The Edit Project dialog now carries a GitHub App
section: when App creds are configured, it shows the current binding (App
installation vs PAT), reuses the same SelectRepoPicker to bind, and an Unbind
button to revert to PAT (sends explicit null). Hidden/disabled for non-GitHub
providers. Once bound, git ops (commits, PR reviews) are attributed to the
App bot instead of the operator's account.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
A project that checks in generated artifacts (RoboCo's lifecycle renders,
verb tables) drifts whenever their source changes. The agent pre-submit gate
(make gate) omits foundation-check, so drift is invisible at the desk and only
fails on CI's drift gate — a failure with no link back to the task, which made
one live task thrash 8 revision rounds.
New per-project codegen_command (migration 078): run in the task's worktree
right before push, and any drift committed into the same push, so CI never
sees stale artifacts. Fail-open — a broken/timeout codegen command logs and
lets the push proceed (CI's drift gate is the safety net); a null command
(every project without checked-in codegen) is a pure no-op. Hooked at both
push_branch (open_pr's first push, the PR head CI grades) and push_task_branch
(later re-pushes). RoboCo sets codegen_command='make codegen' (a new Makefile
target — the write counterpart to foundation-check's read) via the panel.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
The Auditor and PR reviewers get the real DM button (this branch). Secretary
and Intake aren't A2A-DMable — they run their conversation over a live-session
bridge — so they now carry the same chat icon but it navigates to their own
screen instead: Intake -> /prompter, Secretary -> /business?tab=secretary.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* feat(a2a): CEO can DM the Auditor and PR reviewers
A mid-flight PR reviewer or Auditor that's stuck was unreachable — the CEO
had no way to DM them. Both roles now carry dm/read_a2a, so the CEO can open
a 1:1 and they can reply in-thread through the existing CEO-reply path.
Scoped deliberately: the Auditor stays a silent observer to its peers — it
gains no peer-initiation surface (can_a2a_direct routes it through
_check_auditor_a2a, which refuses every initiation target; it can only reply
inside a CEO-opened DM). PR reviewers keep their owning-PM scope. Intake and
Secretary stay excluded — they have their own dedicated chat pages.
NO_COMMS_ROLES drops to {prompter, secretary}; the panel's EXCLUDE_NON_DM_ROLES
matches. KB/docs updated so the 'auditor/pr_reviewer have no dm' claim isn't
left stale.
* test(a2a): smoke guard checks _NO_COMMS_ROLES, not a hardcoded 'auditor'
The dm() runtime guard no longer names the auditor (it now carries dm to
reply to the CEO); it refuses the canonical _NO_COMMS_ROLES set. Assert on
that set so the smoke test tracks the guard, not a stale role name.
* chore(foundation): regenerate verb tables for auditor/pr_reviewer dm+read_a2a
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
formatBucket guessed hourly vs daily from the timestamp string
(bucket.endsWith('T00:00:00.000Z')) and then rendered LOCAL getHours(), so
every daily midnight-UTC bucket rendered as the viewer's local hour — '02:00'
at UTC+2 — across the 7d/30d/90d windows. Granularity is now derived from the
data's own bucket spacing (bucketGranularity: min gap >2h = daily) and daily
buckets render as a short UTC date. The 24h/hourly view is unchanged. Added
minTickGap so the 90d axis doesn't crowd.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* feat(github-app): App credentials, installation tokens, and a Select repo picker
RoboCo was 100% PAT-based. A singleton Fernet-encrypted github_app_credentials
row (migration 077, telegram-credentials pattern) now stores the App id +
private key; github_app_auth mints RS256 app JWTs and caches installation
tokens until 5 minutes before expiry. Projects can bind an installation
(projects.github_installation_id): get_decrypted_token returns a minted
installation token for bound projects and falls back to the stored PAT on
any minting failure, so all ten token consumers work unchanged.
CEO-gated routes expose credentials CRUD plus installation/repo listing, and
the New Project dialog gains a Select repo picker (disabled with a HelpTip
until the App is configured) that fills the git URL and binds the
installation; manual URL + PAT stays the default path.
* test(panel): mock the GitHub App credentials card in the settings page test
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
#615 merged on a false green: the CI paths filter excludes motion/**, so a
motion-only commit fired no quality run and the merge landed two latent
breakages on slave.
- _get_ceo_agent used scalar_one_or_none on role==CEO, which raises
MultipleResultsFound once a second CEO-role row exists. It now pins to the
earliest-created CEO, mirroring the sibling _get_auditor_agent. This is what
made test_brand_voice_nudge_fires_once fail under the full-suite ordering.
- _process_mentions tipped to xenon rank C when the skip-drafting branch was
added; extracted the cap check and the skip filter into two small helpers.
- ci.yml push paths now include motion/** so a motion-only commit can't
false-green the quality gate again.
Regression test: two CEO rows no longer break the lookup.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* feat(x): real voice guide, slop ban, and caption craft for drafted content
Release posts and mention replies were drafted from a one-sentence voice
stub while the reasoning-backed Head-of-Marketing voice guide only reached
the off-by-default spotlight path. The drafting prompts now carry the full
voice rules, a banned AI-slop list, three style exemplars, hook-first
structure, and an under-240-char budget so the 280 clamp never truncates
mid-sentence.
An empty brand_voice now nudges the CEO exactly once (durable
system_settings marker) instead of silently shipping baseline voice forever.
A failed reply draft skips origination instead of shipping 'Thanks for the
mention!'. Video dev prompts and motion/README gain per-platform caption
templates (X: hook + specifics + outro; TikTok: hook + short lines + few
niche hashtags).
* docs(motion): reflow the new Captions section to satisfy the prose gate
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* fix(notifications): task titles and agent slugs replace raw UUIDs
Notification producers interpolated raw task/agent UUIDs into subjects and
bodies ('Task 68e1e4db-... unblocked', 'handed back to 00000000-...-0004').
A tiny notification_text helper (task_display: title-first with a #id8
fallback; agent_display: identity-map slug first, DB lookup fallback) now
feeds every producer: all 13 NotificationService methods, the 7
delivery-service bodies whose subjects were already title-based, the
substitute-PM ad-hoc insert, and the orchestrator/choreographer callers,
which thread the task row's title one call deeper. Fixes the literal
'cell_pm' role string sent as an agent slug in the merge-conflict
notification. Tool-call examples like unblock('<uuid>') keep the raw id on
purpose — agents need it.
* test(notifications): board-review subject assertion matches the humanized format
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* fix(git): branch list classifies remote refs correctly and prunes stale ones
The branches route detected remote-tracking refs via a 'remotes/' prefix
that --format=%(refname:short) never emits, so every origin/* ref rendered
under LOCAL and origin/HEAD surfaced as a fake branch. Listing now uses the
full %(refname) and classifies on refs/heads/ vs refs/remotes/.
Cleanup's remote deletion worked, but no code path ever pruned the viewing
clone's remote-tracking refs, so deleted branches persisted in the UI
forever. The branches route now runs a best-effort 'git remote prune origin'
before listing remote refs, and the manual Fetch fetches with --prune.
* refactor(git): extract branch-line classifier to satisfy the complexity gate
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
The header chip and the Settings User Info card rendered a literal 'Renzo'.
The name now lives in the system_settings store under ceo_name (validated:
trimmed, non-empty, max 60 chars) with the same client-served default the
transcript-retention card uses, editable inline from the User Info card.
Agent prompts already refer to 'the CEO' generically, so no prompt rewiring;
the two agent-facing RAG docs drop the name too. License/CLA copyright is
untouched.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
New Project claimed GitLab/Gitea were 'planned' while both providers are
fully shipped, showed a hardcoded GitHub badge, and never sent git_provider
at all — a non-GitHub project could not be created without an immediate
edit. It now carries the same forge Select the edit dialog has; the edit
dialog's own 'GitLab support is planned' tooltip is corrected too.
Also: the git actions panel's hardcoded 'main' (wrong PR-eligibility and
target label for master-default and env-ladder projects) is replaced by the
project's resolved head branch; acceptance criteria become editable in the
edit-task dialog; three feature flags get their missing descriptions; and
three forms swap raw-UUID text inputs for the existing Task/Agent selectors.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
The per-agent override grid rendered a hand-maintained AGENT_GROUPS literal
that had drifted: ux-dev-2 and all four PR reviewers were absent, so their
model overrides could not be viewed or edited at all (the backend accepts
any slug). The grid now derives its sections from useAgentDefinitions()
with the same team helpers the Fleet page uses, org-ranked ordering, a
loading skeleton, and group HelpTips. The static offline-fallback maps
(agent-utils, use-agents, mock-data) get the missing agents too.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
Every chart hand-rolled its own k-only formatter (22596k-style ticks, left-
clipped y labels), used recharts' default white tooltip (invisible header on
the dark theme), and two charts pinned a numeric XAxis interval that collapses
short series to a single tick. The agent Token Activity chart had no axes at
all and blanked its tooltip date on purpose.
One shared formatTokens/formatBucket (lib/format.ts) and one shared themed
tooltip style (components/charts/chart-tooltip.tsx) now feed all 10 charts;
axis widths/margins sized to the labels; preserveStartEnd tick intervals.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* feat(tg): Mini App V6 — premium overhaul (design system, Chat parity, Metrics drilldown, CEO verbs)
Design system: native type with tabular-numeral heroes (mono demoted to
the wordmark), borderless elevated cards, floating dock, Telegram
window-chrome painting via the theme bridge; Inbox moves behind a header
bell with humanized notifications (UUIDs resolve to task names).
Chat: honest Mine/Fleet split — participant-scoped CEO threads with real
unread counts and mark-read, watched fleet threads with reply-as-CEO on
task-linked conversations (watch-only otherwise), markdown transcripts,
live pulse flashes, and a pinned Secretary live chat on the panel's SSE
session runtime.
Metrics: new tab with period-segmented spend hero, by-agent/team/model
breakdowns, delivery + efficiency health, and a per-agent drilldown over
usage time-series (agent_slug) + member scorecard.
Board: tg-native grouped pipeline replacing the MobileTaskBoard wrapper;
task sheet gains the CEO decide verbs (approve / request changes /
unblock).
Security: /api/dashboard router now require_panel_token-gated at router
level (mirrors /api/usage), closing unauthenticated metrics exposure.
* fix(tg): restore Share Tech Mono brand voice, Phosphor icon set, borderless avatars
The mono returns as the numeral/brand voice (.tg-display — heroes, stat
values, wordmark) while labels stay native sentence case. The hand-drawn
duotone glyphs and lucide feature icons are replaced by Phosphor (MIT):
duotone at rest via an IconContext at the shell, filled weight on the
dock's active tab; row glyph maps (board statuses, inbox kinds, approval
kinds, quick actions) all move over. Team avatar tiles drop their borders
— tint-only squircles.
* fix(tg): fleet avatar strip breathes — spaced tiles instead of overlap
* polish(tg): taste-skill audit pass — em-dash purge, one icon family, separator rationing
Applied the design-taste audit against the cockpit: every em-dash in
visible UI copy rewritten (periods/commas/colons), the remaining lucide
chrome (carets, arrows, send, close, spinners) moved to Phosphor so the
tg tree ships one icon family (send is the native paper-plane, carets
bold), the hand-rolled chevron SVG deleted, and metadata lines rationed
to a single middle-dot separator.
* polish(tg): pipeline chip strip scrolls without a visible scrollbar
* fix(tests): metrics observability fixture uses a relative timestamp
The hardcoded _T0 (2026-06-20) aged out of the service's 30-day window
exactly 30 days later, detonating the suite on every branch. Two days
back from now() stays inside every window (30d metrics, 7d scorecards)
permanently.
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* [68e1e4db] feat(video): author release-0.26.0 marketing clip in the panel-demo kit register
* [68e1e4db] fix(video): extend pk-frame clip window to full length to kill black-gap render bug
* [68e1e4db] docs(video): document release-0.26.0 composition structure in motion/README
Added a comprehensive "Release-specific example: release-0.26.0" section documenting the v0.26.0 marketing video composition: a 40-second panel-demo kit clip showcasing four key features (orchestrator security hardening, three-forge support, Telegram Mini App V4, active guard mode) with choreographed panel elements, stats overlay, and dual-format captions. Includes preview/test instructions, props.js shape, captions schema with verified character counts, and smoke-test invariants matching the pattern established by release-0.25.0.
---------
Co-authored-by: UX/UI Developer 1 <ux-dev-1@roboco.tech>
Co-authored-by: UX/UI Documenter <ux-doc@roboco.tech>
A source=video authoring task reaches awaiting_ceo_approval with no MP4
yet — rendering only happens after it completes — so the CEO had nothing
to review. Two CEO-gated routes now serve the request_render preview
frames: GET /video/preview-frames/{task_id} lists them per orientation
(parsed from the self-describing .previews/{task8}/{orientation}/ filenames
rather than the render_preview marker, which only holds the last call's
single orientation), and .../{orientation}/{filename} streams a frame's
PNG behind the existing path-confinement guard. The task-detail Overview
gains a Video preview card — a 9:16/1:1 toggle + prev/next/scrubber frame
stepper with composition id, duration, and a dirty badge — shown for a
video task with preview frames or awaiting CEO approval.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
draft_release_post was fed highlights=list(report.change_summary) — raw
per-commit subjects — so the announcement model parroted the top commit
('RoboCo API v0.26.0 is out: docs: curate the full Unreleased body #601').
New pure changelog_highlights() extracts the bold feature leads from the
curated release entry (report.drafted_changelog), stripping PR refs and
trailing periods; approve() prefers those and falls back to change_summary
only when the changelog yields nothing. The video captions were already
good because the authoring dev read the changelog — same source now.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
With the guard active on the NAS, a documenter's journal-entry POST body
tripped a WAF signature and the guard banned its docker-bridge IP
(172.18.0.7) — after which EVERY gateway verb from that agent (dm,
i_am_idle, claim_review) was blocked by ip_security, wedging the agent
into a respawn loop. Confirmed live: roboco:guard:banned_ips:172.18.0.7
in redis with passive=False. The guard's threat-ban targets the external
attack surface arriving via nginx; internal HMAC-authenticated agents
reach the orchestrator DIRECTLY on the docker bridge and must not be
subject to it. build_security_config now sets whitelist to the RFC1918 +
loopback ranges. External traffic keeps its real client IP (XFF, trusted-
proxy depth 1 — un-spoofable into a private range), so the WAF still
fires on genuine attackers; the middleware tests model that with a
public TEST-NET-3 IP.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
The release engine's bell notification for a just-originated proposal
inserted through a fresh session while the proposal task sat uncommitted
in the engine's own transaction — the related_task_id FK rejected the
row and the ping was silently lost (caught live in the postgres log; the
DB-free Telegram DM still went out). send_ack_notification now accepts
db_session, forwarded to _create_notification so the insert joins the
caller's transaction, and the release engine passes its session. The
other five callers pass no task_id or reference committed tasks.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
The release proposal's drafted changelog carried only the five entries
PRs had remembered to write, and its own gaps list flagged ~53 missing
items — the entire feature story (Mini App V4+V5, the forge program,
Telegram V2/V3, Agents hub, Workstation, video craft, the panel-perf and
hermetic-test-suite work). The Unreleased body now tells the whole
release: Security amended for #599's calibration (the stale
placeholder-trips-the-guard warning was inverted by the fix) + adm-zip
and registry-auth entries; Added carries the thirteen feature programs;
Changed covers perf/agnosticism/test-hermeticity/docs; Fixed absorbs the
remaining baskets. The readiness drafter prefers this curated body, so
the re-originated proposal ships it verbatim.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
The app-level agent-roster sync fires /api/agents on every surface; on
/tg it races the initData sign-in, takes the cloud-auth 401, and the
interceptor bounced the Telegram webview to the password /login page it
cannot complete — hijacking the cockpit into the dashboard. The Mini App
owns its auth UX (initData sign-in + its own wall), so the redirect now
exempts /tg via an exported, tested path predicate.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
The 2026-07-19 outage root cause: enforce_https keyed off
environment==production, but nginx is the single entry point — the app
only ever sees proxy-HTTP, and the production NAS terminates no TLS at
all — so the moment the guard went active, https_enforcement blocked
the entire request stream. It is now hardcoded off at the guard-config
level (TLS and http->https redirects belong to nginx, not the app).
Same calibration pass for the two prose validators the flip armed: the
secret-exfil key pattern requires a real b64-shaped value so the
documented placeholder lines can't block, and the injection override
pattern requires the second-person 'your' so neutral engineering prose
about the guard subsystem passes.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* perf(panel): kanban virtualization + row memoization + scorecards batch endpoint
The audit's remaining phases: the kanban card lists render through
@tanstack/react-virtual windows (columns are the dnd drop targets, cards
only drag — no sortable conflict) with memoized columns/cards and a
stabilized handleAction; the task table's desktop row and mobile card are
extracted and memoized (the table itself already client-paginates to 100).
Backend: the per-member scorecard N+1 (~20 requests x 3 queries per poll)
collapses into GET /dashboard/metrics/members backed by
get_all_member_scorecards with grouped rollup/overlay SQL shared with the
single-agent path.
* test(metrics): shared-DB-safe scorecard tests — unique seeds, delta assertions
The two new batch-scorecard tests assumed a private DB: fixed ceo/system
slugs collided with other tests' seeds (ix_agents_slug) and a global
exactly-one CEO lookup + exact roster count broke in the one-process
suite. Unique slugs, subset/disjoint assertions, dead count constant
dropped.
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
An operator .env with deploy-flavored values skewed local runs three
ways: environment=production flipped the GHSA-4f7g fail-closed auth gate
(header-trust tests 401), an armed cloud_auth 401'd every agent request,
an armed guard rate-limited the suite, and a missing encryption key
failed every crypto path. The autouse fixture pins environment,
encryption_key (per-process Fernet), cloud_auth_enabled, and
guard_enabled to schema defaults; suites exercising those postures arm
them per-test. Full suite locally: 13641 passed, 0 failed — was 174
failures under an armed .env.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
Slice updates cover the forge package + provider parity, Telegram cockpit
V4/V5 surfaces, panel perf (virtualized kanban, scorecards batch route),
close_task_pr_best_effort + dependents guard, NO_COMMS_ROLES + CEO A2A
refusal at conversation creation, notification ack-TTL, docs-site config,
pr_labels base-branch signature, taste-skill prompt layers, prompter
history exclusions, live forge e2e suites, and migrations 075/076.
_complete_map.md is regenerated from the slices — the committed concat
had drifted ~8KB behind its own sources.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
Seventeen-page sweep of the agent-facing KB against shipped behavior:
required covers_parent_criteria and per-AC criteria_verified reach the
QA/PM/task-tools pages (the QA docs also named non-callable pass_review/
fail_review — the MCP tools are pass/fail); collision_context lands in the
QA/gate/planning evidence docs; the possibilities matrix gets its own
architecture page + config entry; the auditor page gains its missing
waive_finding and playbook-curation verbs; git-pr-types.md is rewritten
off the long-dead is_root_pr model; PR/workspace/git-error pages stop
assuming GitHub (forge-agnostic + env-ladder semantics).
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* feat(tg): Mini App V5 — brand typography, icon depth, motion, detail sheets
Share Tech Mono (the vendored motion-brand face) becomes the cockpit's
display voice via next/font/local scoped to #tg-shell; icon tiles, circle
actions, and avatars get gradient/ring depth; a dependency-free motion
vocabulary lands (spend count-up, tab rise-in, staggered sections,
sparkline draw-in, sheet slide-up with native BackButton dismiss); the
Board tab gains a tap-through task sheet (ACs, open findings, PR link),
Today's fleet opens a full-roster sheet, and Board/Inbox join the
/tg?demo=1 fixtures.
* feat(tg): custom RoboCo icon set + operations ring
The cockpit stops using stock lucide on its hero surfaces: a hand-drawn
duotone icon set (speedometer, seal, bell, kanban, brand-cursor bubble,
rocket, double-check, broom, robot head) covers the tab bar and the Today
ring. The ring itself stops duplicating the tab bar and becomes real
operations: Ship deep-focuses the release proposal in Approvals, Ack all
bulk-acknowledges pending notifications, Sweep runs the stale-branch
cleanup across every git-configured project behind a confirm sheet, and
Fleet opens the roster.
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* fix(security): guard goes active; CEO A2A respects no-comms roles; ack notifications expire
ROBOCO_GUARD_PASSIVE_MODE defaults to false in both compose files — the
deferred post-calibration flip; fail_secure stays off and the env override
remains the rollback. can_a2a_direct no longer short-circuits the CEO past
the no-comms set (auditor/pr_reviewer/prompter/secretary), now canonical
in foundation.policy.communications.NO_COMMS_ROLES and shared with the
content-actions gate; the A2A service refuses at conversation creation
instead of silently suppressing the wake. Ack-required notifications get
expires_at stamped from ROBOCO_NOTIFICATION_ACK_TTL_HOURS (default 48,
0 disables), so the re-escalation sweeper's expires_at query matches rows
for the first time.
* refactor(notification): extract _ack_and_expiry — xenon rank back under B
The expires_at stamping pushed _create_notification_with_session to
rank C; the requires_ack + expiry derivation moves into a helper with
the same semantics and comments.
* test(conftest): dispose the global DB engine after every test
Production code reaching get_db_context()/get_engine() lazily creates the
process-global engine bound to the current event loop; with per-test
function-scoped loops, any later test touching the global path inherits a
dead-loop engine and dies with 'Future attached to a different loop' —
the order-dependent class that has been wandering the suite (cloud_auth
login, metrics, tasks-routes, full-lifecycle) whenever collection order
shifts. An autouse fixture now close_db()s after every test, keeping the
global path loop-local; no-op when untouched.
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
Thread the deployer's product name through the X reply + feature-spotlight
prompts (B6 leftover; release/video paths shipped in #570); make the
docs-site repo/URL config (ROBOCO_DOCS_SITE_*, defaults unchanged) instead
of a roboco-website hardcode (B8); de-assert our repo from the Main PM
prompt (B10); derive PR labels from the real target branch instead of
literal to-master/to-slave; drop the stale headcount from base.md; and
make the bash-guard's Makefile check require an actual quality/gate/lint/
test target before denying raw package-manager commands (no more
false-remediation loop on Go/Rust Makefiles).
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
Task cancellation left the task's PR open on the forge forever: cancel()
now best-effort-closes the recorded PR for the task and its cascaded
descendants (close_task_pr_best_effort resolves owner/repo off git_url —
no clone needed; never raises into the cancel). The bulk stale-branch
sweep gains a dependents guard: a branch still recorded by a non-terminal
task, or serving as a live child's resolve_parent_branch base, is excluded
from the candidate window — mirroring the existing env-ladder-rung skip.
Scoped to the sweep, not delete_task_branch, so the BFS cascade-cancel
can't falsely block a parent's branch on its own about-to-cancel child.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
The history-digest pipeline (PR #297) injected past-task data but nothing
told the intake agent what to do with it: prompter.md now carries an
explicit don't-re-propose section (cite duplicates by short id, reference
precedent in notes, let history inform depends_on sequencing), and
list_recent_for_project excludes cancelled tasks so dead work can't pose
as precedent.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
The deferred half of the Leonxlnx/taste-skill (MIT) adoption: an
industrial-brutalist / minimalist-editorial / premium-agency aesthetic
vocabulary keyed onto the existing design-bar dials for both FE and UX/UI
team prompts, and a ux_ui-only image-direction section (composition,
palette discipline, anti-slop imagery, mockup conventions) with a pointer
from frontend.md. Layer tests pin presence and team scoping.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
Four Phase-0 tests still asserted gitlab.com URLs / git_provider='gitlab'
get rejected, which the GitLab forge provider (#581) made false. They fail
deterministically in isolation and only pass CI when one-process test
ordering masks them — the latent red behind today's flaky quality gates.
Flipped to assert acceptance (auto-detect + explicit), mirroring the GHE
escape-hatch test; unknown-host/unknown-provider rejection tests stay.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* feat(tg): premium Today — spend hero, trend, quick actions, fleet avatars
The cockpit home stops being flat cards and becomes a real app surface:
- Spend HERO: the day's cost at 40px with a signed delta-vs-yesterday
chip and a live 7-day amber area sparkline (hand-rolled inline SVG, no
charting lib in the Mini App bundle).
- Quick-action ring: circular Approve (amber + needs-you badge) / Board /
Inbox / Chat, the wallet-style primary-verb row.
- Needs-you as a rich amber gradient banner (top items + draft chips)
instead of a plain section.
- Fleet as live avatar tokens (stable per-name hue, pulse dot) over the
working list.
- "Shipped this week" day-bars (today emphasized) + week total.
Backend: /telegram/today gains spend.series (7-day cost) + delta_pct and
a velocity series (per-day completed tasks) — two cheap grouped-by-day
queries, same DB-only ethos, degrading to zeros on error.
* feat(tg): color-code approval rows by kind
TgRowIcon gains a tone prop; the approvals list tints each tile per kind
(amber Release / sky X post / violet Video / emerald Roadmap) so a mixed
queue reads as color-coded instead of a monochrome column.
* feat(tg): sender/peer avatars on Inbox + Chat
Inbox notifications and Chat conversation rows adopt the fleet-avatar
language: a per-name-hued initials token leads each card, unread inbox
items carry a subtle primary tint, and both cards move to the rounded-2xl
surface — so every tab now shares one visual system. Board keeps the
shared MobileTaskBoard (already grouped/pill-styled, and reused outside
the cockpit).
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
GitLab's create_org_repo (services/forge/gitlab.py) replaces the Phase-3
synthetic 501 with a real implementation: resolves org (a group's full
path, subgroups included) to a numeric namespace id via GET
/groups/{path}, falling back to the token's own namespace on a 404
(personal-namespace projects); POSTs /projects with the
name/path/description/visibility/initialize_with_readme payload
(visibility mapped private->"private"/"internal"); reshapes the 201
onto the GitHub fields callers read (full_name/clone_url/html_url) and
GitLab's duplicate-path 400 "has already been taken" onto GitHub's 422
shape, text preserved. Gitea's create_org_repo was already real but
untested — added transport-level coverage.
GitHubProvisioningService (services/github_provisioning.py) is now
provider-aware: ROBOCO_PROVISIONING_PROVIDER (github default / gitlab /
gitea) and ROBOCO_PROVISIONING_HOST (self-hosted instance, required for
gitlab/gitea or the service stays disabled exactly like a missing
token/org) pick the target forge; the class/factory names stay
GitHub-flavored for backward compatibility (pitch.py and existing
imports untouched). A shared _is_already_exists() helper recognizes
GitHub's "already exists" (422), Gitea's (409/422 "already exists"),
and GitLab's reshaped "has already been taken" (422). The existing-repo
re-fetch now builds a provider-aware RepoRef (GitLab packs org/name
into the owner field; GitHub/Gitea keep the owner,repo pair). Default
behavior (no new env set) is byte-for-byte the Phase-1 GitHub path,
pinned by a regression test.
Gates: ruff format/check clean, mypy roboco/+tests/ clean (1235 files),
xenon A/A/B clean, targeted suite (forge + provisioning + pitch) 79/79
green.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
Caught live on the NAS: every cloud-auth login failed with
422 {"loc": ["query", "db"]} regardless of credentials. Root cause:
auth/manager.py used postponed annotations with a TYPE_CHECKING-only
AsyncSession import, so FastAPI could not resolve get_user_db's
Annotated[AsyncSession, Depends(get_db)] at runtime and silently
demoted `db` to a required query parameter.
The module now evaluates annotations eagerly (no future-annotations,
runtime imports) so an unresolvable annotation is loud instead of a
silent contract change. Regression test mounts the REAL login router —
no dependency overrides, which is exactly why the existing suite never
caught this — and asserts wrong credentials yield 400
LOGIN_BAD_CREDENTIALS, never a 422.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* test(forge): live-GitLab contract suite — verified against gitlab.com
Mirror of the Gitea live suite for GitLabProvider (forge Phase 3):
env-gated (ROBOCO_GITLAB_E2E_URL/_TOKEN + ROBOCO_E2E_SMOKE=1),
self-seeding — creates a throwaway private project under the token's
namespace, pushes real commits, and drives the provider end to end: MR
open → duplicate 409→422 reshape → native source/target filter →
GitHub-shape adaptation (iid→number, head/base, merged) → per-file diff
reassembly → note review → commit-status→check_runs reshape → squash
merge with a mergeability-settle retry (GitLab computes merge status
asynchronously) → branch delete → release (shaped html_url) → the
oauth2 Basic-auth git-CLI claim. Project deleted afterwards.
Green against gitlab.com on first full run — no adapter fixes needed
(the settle-retry is the one live-behavior accommodation). Requires a
token with Project: Create (classic `api` scope, or fine-grained with
project create + API read/write).
* fix(tests): split None-guarded assert so its message can't dereference None
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* feat(forge): Phase 2 — Gitea provider, per-call routing, host registry
Gitea support lands behind the Phase-1 seam:
- GiteaProvider (services/forge/gitea.py): Gitea v1 transport addressed
by instance host (api base from the project's git_url). Where Gitea's
wire contract diverges from GitHub's, the provider adapts responses
back into the shapes GitService already classifies (ShapedResponse):
`token` auth scheme, duplicate-PR 409→422 with the "already exists"
text GitService keys on, commit statuses reshaped into check_runs /
workflow_runs envelopes, APPROVE→APPROVED review mapping, Do-keyed
POST merge, merge-method repo keys, label-color '#' prefix,
client-side head/base PR filtering. Deliberate postures per the spec:
zero-workflows fail-open (statuses-free repo → no_ci_configured) and
merge_branch as a shaped 501 (env-sync cascade lands on missing_ref;
the shared local-git fallback is Phase-2.1).
- ForgeRouter (services/forge/router.py): GitService._forge now routes
per call from RepoRef.host — every existing call site unchanged in
shape. RepoRef gains an optional host; _parse_git_url returns the
host-stamped ref and it is threaded through GitService/release
executor instead of being rebuilt from strings (helpers re-signatured
to take RepoRef).
- Host registry (services/forge/registry.py): in-memory host→provider
map, self-healing — ProjectService.get/get_by_slug re-register on
every read; provider_for resolves gitea projects by git_url host.
- Registration validation now accepts git_provider="gitea"; GitLab
remains recognized-but-rejected. Panel: the read-only Forge badge
becomes a real picker (Auto-detect / GitHub-GHE / Gitea / GitLab
disabled).
Plain git (clone/fetch/push) needs no changes — the Basic-auth
extraheader works on Gitea unchanged. Gates: mypy 392 files, xenon A,
full unit suite 6356 green, integration suite 2257 green.
* feat(forge): live-Gitea contract suite + scheme support + slash-safe refs
Hardening from running the provider against a real dockerized Gitea
1.22.6 (the spec's Phase-2 contract suite, now committed as the
env-gated tests/e2e_smoke/test_gitea_live.py — self-seeding: creates its
own repo, pushes real commits, and drives PR open → duplicate reshape →
list/filter → diff → review → labels → commit-status CI reshapes →
squash merge → branch delete → release, plus a live verification of the
x-access-token Basic-auth git-CLI claim).
Two real findings fixed:
- Branch refs weren't URL-encoded — every RoboCo branch carries slashes
(feature/backend/...), and Gitea's router 404s on the extra path
segments. list_ci_runs + delete_branch_ref now quote the ref
(regression-pinned in the unit suite).
- The API base hardcoded https; a LAN instance serving plain http is a
real deployment shape. GiteaProvider gains a scheme (recorded per host
by the registry from the project's git_url).
ShapedResponse moves to forge/shaping.py (shared by the upcoming GitLab
transport, which needs its text override for diff reassembly).
* feat(forge): Phase 3 GitLab provider + Phase 2.1 local-merge fallback
GitLabProvider (services/forge/gitlab.py): GitLab v4 transport addressed
by host+scheme, subgroup-safe (the MR project path packs into
RepoRef.owner, URL-encoded per call). Adapters translate MR semantics
into the GitHub shapes GitService classifies: iid→number,
source/target_branch→head/base with a merged bool, per-file diffs
reassembled into unified-diff text (ShapedResponse text override,
3-page cap), approve-vs-note review routing (GitLab has no
request-changes verb), pipelines/statuses reshaped into
workflow_runs/check_runs, merge-method repo-key mapping, duplicate-MR
409→422. Reviewer mirroring is skipped (needs numeric ids RoboCo
doesn't store); provisioning stays Phase 4. gitlab.com now auto-detects
at registration like github.com; self-hosted GitLab sets the provider
explicitly (panel picker enabled).
Phase 2.1: neither Gitea nor GitLab has GitHub's server-side merges
API — their merge_branch returns a shaped 501 and
GitService.sync_env_branch now runs the shared local-git fallback
(_local_merge_branch: throwaway clone → ancestor check → merge → push;
a conflict aborts with the remote untouched; same status vocabulary as
the merges-API path).
Also aligns the whole tree with the full gate's tests/-scoped mypy
(provider-test responder typing, e2e_smoke's stale owner/repo shapes).
Gates: mypy 1229 files clean, xenon A, unit suite 6393 green, forge
suites 85 green, panel typecheck/lint clean.
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
The full-suite CI run executes every test tree in one process against one
database, so other suites' committed rows are visible to the Today-brief
queries and the "empty company" / absolute-count assertions failed
(assert 4 == 0). The suite now snapshots a pre-seed baseline and asserts
deltas, seeds at priority 0 so its rows stay inside the brief's item
caps, and uses a unique agent slug (the fixed be-dev-1 could collide
with leaked rows). Verified by co-running with the known-polluting
integration suites in one process.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
CI runs `mypy roboco/ tests/`; the bridge/cockpit suites were only gated
against `mypy roboco/` locally. Real credentials dataclass instead of
SimpleNamespace, AsyncMock casts where mocks sit behind typed fields, a
return annotation on the fake stream, a None-guard on the consumer task,
and a fresh registry lookup where mypy's literal narrowing read an
assert as always-false.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* feat(tg): P0 — dev mock bridge + Telegram-native foundations
Mini App V4 phase 0. The (tg) shell gains the groundwork every later
phase builds on:
- Dev mock bridge: outside Telegram, a development build falls back to a
no-op WebApp object and skips the webapp-auth POST (the regular panel
session cookie authorizes API calls), so the cockpit is workable in a
plain browser. Production keeps the "Open from Telegram" wall.
- Telegram theme adoption: themeParams map onto the shadcn CSS variables
scoped to #tg-shell (desktop dashboard untouched), colorScheme drives
the dark class, themeChanged re-applies live. Non-hex values are
dropped at the trust boundary.
- Viewport/swipe correctness: shell height rides Telegram's own
--tg-viewport-stable-height (100dvh fallback), vertical swipe-to-close
disabled so list scrolling can't dismiss the app.
- Native chrome bindings: TgWebAppProvider context plus useMainButton /
useBackButton declarative hooks and a null-safe haptics helper —
consumers never touch window.Telegram directly.
* feat(tg): P1 — Today home tab + one-round-trip /telegram/today brief
Mini App V4 phase 1: the cockpit now opens on a "Today" brief answering
"does anything need me?" in one glance.
Backend: GET /api/telegram/today (CEO-gated, rate-limited) returns the
whole brief in one round trip via the new TgCockpitService — needs-you
items (awaiting-CEO + blocked tasks capped for a phone screen, held-draft
counts across release/X/video/roadmap queues), a fleet snapshot with
per-agent current-task titles, today's spend from the day rollup
(degrading to zeros on a usage hiccup, mirroring the CEO overview), and
ship state. Deliberately DB-only: no live GitHub calls, no readiness
snapshot (that path clones), no orchestrator singleton — the CI
red/green proxy is the set of open ci_watch fix tasks.
Panel: TgTodayTab is the new default tab (Gauge icon) — needs-you rows
and draft chips deep-link into the tab that acts on them (with a haptic
tap), fleet/spend/ship render as dense cards, 45s refetch until the P3
WebSocket wiring lands.
* fix(tg): dev mock engages when the CDN bridge loads outside Telegram
Live browser smoke caught it: a bare tab still loads telegram-web-app.js,
so window.Telegram.WebApp EXISTS outside Telegram — just with empty
initData. The dev fallback keyed on a null bridge only, so a dev browser
went down the real-auth path and posted empty initData instead of
mounting the mock. The fallback now treats bridge-with-no-initData the
same as no bridge (a real Telegram launch always carries initData);
production behavior is unchanged.
* feat(tg): P2 — native approvals card stack
Mini App V4 phase 2: the Approvals tab stops stacking the four desktop
queue cards and becomes a phone-native flow — one normalized list across
release proposal / X drafts / video drafts / proposed roadmap items, and
a full-context detail per item:
- Release: version/bump/gate badges, changelog draft, gaps, migration
notes, in-flight + failed-execute banners; approve runs the fail-closed
executor, reject requires a substantive change request (10 chars).
- X: editable body with the live 280 counter, replied-to mention quoted;
approve sends the edited body only when actually edited.
- Video: cut-toggled player (blob-fetched through the authed client — a
bare <video src> would 401), per-platform caption edits with 280/2200
counters; approve sends only checked-in edits.
- Roadmap: the PO's full pitch (description, rationale, ACs); approve
materializes into the backlog per item.
The detail's primary action rides Telegram's native MainButton and back
navigation rides the BackButton, with visible fallbacks outside Telegram
(dev mock, old clients). Haptics fire on outcomes. An acted-on item
vanishes from the refetched queue, popping back to the list by
construction. A failed queue source is surfaced ("list may be
incomplete" / "couldn't load") instead of masquerading as an empty
queue — caught live in the browser smoke.
* feat(tg): dev demo mode — /tg?demo=1 renders canned cockpit data
Development-only: with the flag param present, the Today brief and the
four approval queues resolve typed fixtures (dynamically imported, so
production bundles never carry them) instead of hitting the backend —
the cockpit is fully browsable with zero stack running. Mutations still
go to the real API and fail loudly; it's a showroom, not a simulator.
* feat(tg): P3 — cockpit rides /ws/system live
Chat adopts the desktop A2A invalidate-on-frame idiom over the shared
ref-counted /ws/system socket: every a2a.message frame refreshes the
conversation list and the affected thread, missed-frame gaps are healed
by a reconnect refetch, and the 10s thread poll turns off entirely while
the socket is up (it remains the fallback). The Today brief refreshes on
each USAGE_SNAPSHOT push so the spend line tracks the sweeper live, with
the 45s poll as the socket-down fallback. No new sockets, no backend
changes — the WS gate already accepts the cloud-auth session cookie.
* feat(tg): P4 — deterministic bot command tier + self-syncing menu
Mini App V4 phase 4 (deterministic half): three new bot commands beside
/status /queue /task —
- /agents: who's mid-task right now, from the same TgCockpitService
fleet snapshot the Today brief renders (now public `fleet()`).
- /usage: today's spend from the day rollup.
- /blocked: awaiting-you + blocked tasks, deep-linked into the panel,
capped per section, titles HTML-escaped.
BOT_COMMANDS is the single registry driving /help AND a once-per-process
Bot API setMyCommands sync on the first poll cycle (new client method,
best-effort), so the Telegram command menu can never drift from what the
code implements. The interactive tier (/secretary, /newtask riding a
live Intake interview in-thread) is specced but not in this commit.
* feat(tg): direction-C styling pass — Telegram palette, RoboCo voice
The cockpit stops wearing default-shadcn and gets its own visual
language on top of the P0 themeParams bridge (colors stay CSS-variable
driven, so inside Telegram everything still adopts the user's theme):
- Shared primitives (components/tg/ui.tsx): TgSection grouped cards with
tracked micro-label headers, TgRow list rows (44px targets, press
feedback, 1/2-line clamp), TgRowIcon glyph tiles, TgStat tabular-nums
figures. Every tab composes the same three, so density and rhythm are
identical across the surface.
- Shell renders a centered 430px column (sm:border-x) — the phone UI no
longer stretches across a desktop dev browser.
- Tab bar: tighter type, active stroke-weight shift, backdrop blur.
- Today: needs-you count badge, divided task rows with inline blocked
marker, fleet as mono-named rows, spend/ship as stat tiles.
- Approvals rows as icon-tile cards; detail header gains the kind glyph.
- Inbox/Chat rows aligned to the same card language.
* feat(tg): P5 — /secretary and /newtask live-chat bridges
The bot's interactive tier: both commands bridge the CEO's Telegram chat
into the same in-process runtimes the panel drives — the persistent
Secretary container and the scoped Intake interview.
There is no synchronous send→reply seam (replies land on the session's
single-consumer relay queue), so each bridged session runs one long-lived
consumer task (roboco/services/telegram_bridge.py) that drains
PrompterLiveRegistry.stream and pushes one Telegram message per completed
turn. While a session is live, plain chat text IS the conversation;
/end closes it.
/newtask resolves the intake scope (single project auto-picked, multiple
offered as a tap-to-pick keyboard holding the initial text), and the
interview happens in-thread. A draft proposal renders as a card with
Send-to-Board / Discard buttons: confirm routes through the normal
board-review path (PrompterService.confirm_live_draft, route=board) and
PARKS the session — board feedback later streams straight back into the
same thread, closing the redraft loop from the phone. MegaTask batches
still confirm in the panel only.
The consumer's open stream arms the registry's 60s keepalive, so the
bridge runs its own idle TTL (same setting, parked sessions exempt).
State is per-process in-memory by design (the _PENDING_REPLIES posture);
intake/secretary containers are process-wide singletons, so a bridged
session preempts a live panel session of the same kind by construction.
* feat(tg): cockpit skin — RoboCo dark deck with a constant amber accent
The cockpit no longer inherits the dashboard's white default outside
Telegram: #tg-shell carries its own standing skin (deep slate surfaces,
amber primary) so the Mini App looks like RoboCo everywhere. Inside
Telegram the themeParams bridge now overrides SURFACE tokens only —
background/card/text/hint/border repaint to the user's Telegram theme
while --primary/--ring stay RoboCo amber: Telegram's surfaces, RoboCo's
voice. Demo fixtures also rewritten to neutral content (they previously
depicted unbuilt forge work and already-shipped roadmap items as live).
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* [467e263d] fix(ci): restore green Python quality gate on slave (mypy + xenon)
Two independent bugs broke run 29653535468's 'Python quality gate' job,
not the pydantic-settings pin (already correctly 2.14.2 on this branch).
- git.py: _delete_remote_branch_best_effort's success path fell off the
function with no return, failing mypy's missing-return-statement check
and silently returning None instead of the documented True.
- company_goals.py: CompanyGoalsService.upsert's six repetitive
`if key in data` branches pushed the module's average cyclomatic
complexity past xenon's --max-modules A gate; refactored into a loop
over the mutable-field tuple (behaviourally identical).
* [467e263d] docs(changelog): restore green Python quality gate on slave (CI run 29653535468)
Documented two independent code-level bugs fixed in commit 5ed93ca7:
1. git.py: missing return True statement in _delete_remote_branch_best_effort success path (mypy failure)
2. company_goals.py: cyclomatic complexity refactoring for xenon gate pass
Both bugs identified and fixed by dev team; verified via full CI reproduction locally.
---------
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>
* feat(forge): Phase 0 — git_provider column + registration-time forge validation
Pointing a project at a GitLab/Gitea git_url used to fail silently, several
steps deep, at first PR. New pure policy module (foundation/policy/forge.py)
detects the provider from the git_url host and validates at the
ProjectService create/update chokepoint: github auto-detects and
auto-stamps, explicit git_provider=github is the GitHub Enterprise escape
hatch, gitlab/gitea are recognized-but-not-yet-supported, unknown hosts get
a loud rejection with guidance. An update changing git_url does NOT inherit
a stored auto-stamped provider (restating the override is required), so a
host swap can't smuggle the escape hatch past validation. Migration 075
adds the nullable projects.git_provider column; the panel project dialogs
show the detected forge. Phase 0 of the forge-providers spec.
* feat(forge): Phase 1 — GitProvider seam, GitHub transport extracted
roboco/services/forge/: base.py holds the pure contracts (RepoRef +
GitProvider ABC, stdlib-only — a later GitLabProvider is implemented by
reading this file alone), github.py the httpx transport (20 endpoint
methods behind one shared request-plumbing helper set), registry.py the
wiring (git_provider column -> provider, failing loud on gitlab/gitea).
GitService keeps its exact public surface and all response
classification; its 26 inline REST call sites route through a lazy
_forge property (several suites build GitService via __new__, so an
__init__-set attribute breaks them). github_provisioning and
release_executor ride the same provider. Zero behavior change — the
pre-existing suites pass unmodified; per-project provider resolution
lands with the second provider.
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
Item B+C of the video/X per-project targeting spec, plus the
company_goals.company_name field they depend on (migration 075).
- CompanyGoalsService.resolve_product_name is the single fallback chain
(project name -> charter company_name -> RoboCo); XEngine and
VideoEngine both call it and their prompt builders are pure functions
taking product_name — release posts/videos stop hardcoding RoboCo.
- The X and video queue responses carry project_slug/project_name via one
shared unloaded-guard helper (api/schemas/project_fields.py); both
panel queues render a shared ProjectBadge so multi-project drafts are
tellable apart.
- Business -> Goals editor gains the company-name input.
- Fixes a pre-existing test-isolation leak: the company-goals routes test
commits the charter singleton into the session-scoped test DB and
polluted later suites; it now deletes the row on teardown.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* feat(forge): Phase 0 — git_provider column + registration-time forge validation
Pointing a project at a GitLab/Gitea git_url used to fail silently, several
steps deep, at first PR. New pure policy module (foundation/policy/forge.py)
detects the provider from the git_url host and validates at the
ProjectService create/update chokepoint: github auto-detects and
auto-stamps, explicit git_provider=github is the GitHub Enterprise escape
hatch, gitlab/gitea are recognized-but-not-yet-supported, unknown hosts get
a loud rejection with guidance. An update changing git_url does NOT inherit
a stored auto-stamped provider (restating the override is required), so a
host swap can't smuggle the escape hatch past validation. Migration 075
adds the nullable projects.git_provider column; the panel project dialogs
show the detected forge. Phase 0 of the forge-providers spec.
* fix(panel): mock-mode forge detection extracts the real host
CodeQL js/incomplete-url-substring-sanitization: the substring check
matched github.com anywhere in the URL. Extract the hostname (URL parse
or scp-form regex, mirroring forge.py) and require an exact match.
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* [e530aa5e] Diagnose and fix roboco-api CI failure (run 29629255153) (#561) (#562)
* [e530aa5e] fix(tests): narrow None before indexing validate_init_data() result in telegram_initdata self-check
CI run 29629255153 failed on mypy, not the historical pydantic-settings
issue (uv.lock already pins 2.14.2). The __main__ self-check block in
test_telegram_initdata.py indexed the dict[str, object] | None return
of validate_init_data() without narrowing away None first.
* [e530aa5e] docs(qa): document CI fix for mypy type narrowing in telegram_initdata test
Explains the root cause (mypy type error in __main__ block), the solution (None narrowing before indexing), and the safe pattern for future test self-checks that call functions returning optional types.
---------
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>
* [0884b737] Diagnose and fix Python quality gate + e2e lifecycle smoke CI failures on PR #563 (#564) (#565)
* [0884b737] fix(tests): isolate ROBOCO_SDK_URL for scripted e2e-smoke agents
tests/e2e_smoke/harness.py already isolates ROBOCO_AGENT_TOKEN from the
host environment (the #503/#504 fix) but left ROBOCO_SDK_URL leaking
through. flow_server/do_server both default it to
http://localhost:9000 and forward every rejection there for the
per-verb circuit breaker; inside a real spawned agent container that
port is a live SDK loopback, so the breaker records genuine attempts
for the ephemeral test-agent IDs and trips circuit_open mid-test
(test_sandbox_on_demand.py::test_request_sandbox_guard_chain_over_real_api,
which deliberately causes 3 rejections in a row). Point it at a
guaranteed-refused loopback address so every environment gets the same
fail-open bypass a bare CI runner already gets by having nothing
listening on 9000 at all.
* [0884b737] docs(changelog): document e2e-smoke harness ROBOCO_SDK_URL isolation fix
Document the fix that isolates ROBOCO_SDK_URL in the ScriptedAgent harness to prevent the per-verb circuit breaker from leaking state into ephemeral test-agent identities when the e2e-smoke suite runs inside a live agent container. This ensures the suite passes consistently regardless of whether it runs on bare CI or inside a spawned agent.
---------
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>
* [3b9a1771] Diagnose and fix ALL make quality + e2e-smoke stage failures on PR #563; confirm real CI green (round 3) (#566) (#567)
* [3b9a1771] fix(e2e-smoke): match real embedding dimension when seeding fake journal chunk
test_c3_deleted_journal_unindexed inserted a 4-dim placeholder vector
into chunks_journals, but the e2e stack's app lifespan eagerly creates
that table with the real settings.embedding_dimensions (1024) before
the test runs, so the insert failed with "expected 1024 dimensions,
not 4". Derive _SMOKE_DIM from settings.embedding_dimensions instead
of a hardcoded constant so the seeded vector always matches the
table's actual column width.
* [3b9a1771] docs(qa): document e2e-smoke embedding dimension fix in round 3 CI diagnosis
Recorded the root cause, solution, and pattern for the final e2e-smoke test failure found in comprehensive sandbox testing: the test seeded a 4-dim placeholder vector but the app's eager lifespan init created chunks_journals with the real 1024-dim embedding column. Updated _SMOKE_DIM to derive from settings.embedding_dimensions instead of a hardcoded constant.
---------
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>
---------
Co-authored-by: Backend Developer 1 <be-dev-1@roboco.tech>
Co-authored-by: Backend Documenter <be-doc@roboco.tech>
* feat(panel): Journals joins the Agents hub as its third tab
* docs(map): Journals tab on the Agents hub; registry retargets
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* feat(panel): Agents hub — Fleet and Conversations tabs, /a2a redirect, DM quick-action
* fix(panel): validate deep-linked DM targets against roster and exclusions; re-arm the dm latch
* docs(map): panel entries for this wave
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* feat(panel): customizable Quick Actions on Overview — registry, defaults, persisted picker
* fix(panel): legacy bar actions join the quick-action defaults; empty state; persist-contract test
* docs(map): panel entries for this wave
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* feat(panel): Workstation card grids with persisted view toggle and sorting
* fix(panel): stable multiplier sort with cell-label fallback after adversarial review
* docs(map): panel entries for this wave
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* feat(motion): demo-register visual design bar, vendored craft references, catalog vocabulary index
* fix(motion): correct design-bar composition claims after adversarial fact-check
* docs(map,rag): design-bar, vendored references, and catalog index on the video-engine surfaces
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* feat(panel): Workstation page — Products and Projects as tabs behind one sidebar entry
* docs(map): workstation page, extracted views, and the local-filter scroll trade-off
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* feat(panel): session-start, 7d overview spend, and 30d business spend charts
* feat(panel): surface work sessions as a Git page tab (route was orphaned)
* feat(git): reap spent task branches and render previews at lifecycle chokepoints; guarded stale-branch sweep
* feat(panel): stale-branch cleanup button on the Git page
* fix(git,panel): cursor-resumable sweep, force-delete spent refs, local filter state
* docs(map,rag): branch/preview reaping, cleanup sweep, git-tab work sessions, wave-2 charts
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
A draft targeting a platform with no credentials could never complete:
the X leg posted and committed its id, the TikTok leg failed as
'transient', and the card sat pending in the CEO queue with retry
semantics that structurally cannot succeed. Unconfigured platforms are
now an explicit skip: the draft COMPLETES when every configured
platform has posted (detail names what was skipped), an approve whose
targets are all unconfigured refuses loudly instead of silently
completing, and genuine post failures keep their partial/retry
semantics. Re-approving a parked card clears it without re-posting
(the already-posted guard is covered by a dedicated test).
Verified: 32/32 test_video_post_service (3 new: skip-and-complete,
all-unconfigured refusal, re-approve recovery), ruff/mypy/xenon clean.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
Playwright, the taste-skill design bar, and hyperframes were all
implemented — and each stopped one hop short of the hands doing video
work: playwright reached only fe-qa/ux-qa, the design bar's web-UI dials
actively pointed a video task at 'dense product UI -> motion 2-3', and
hyperframes' agent-facing doctrine never reached any agent. Three wires:
- vendor the official HyperFrames agent skills (hyperframes-core,
-keyframes, -creative) under motion/skills/ at a pinned upstream
commit (Apache-2.0, attribution headers; prose reflowed to house
style, re-vendor note in each header); README and the dev video
prompt block point at them
- register the playwright MCP for a ux-dev spawned onto a source=video
task (_is_video_authoring_spawn: fail-closed role/team/task-source
probe) so the composition author can watch their HTML live in a real
browser between renders — gating-only, agent-ux already bakes the
browser; QA gating unchanged, be-qa/ordinary ux-dev still excluded
- design bar video-mode override in the ux_ui team prompt: video tasks
are films, the web dials do not apply — use the cinematography bar
and the vendored doctrine instead
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
The release cuts read as static screen recordings: one 6s cursor glide
that then vanished, a locked-off camera, metronomic flat card fades. And
two renderer defects were silently wrecking every cut:
- @hyperframes/producer floated (^0.7.36, no lockfile — image builds get
whatever is latest): 0.7.60 fails EVERY render with "Cannot access 'rt'
before initialization". Pinned 0.7.36 exact and committed
package-lock.json so sidecar builds are reproducible.
- the producer's per-clip visibility scheduler runs on a clock that lags
~50% behind the encoded timeline on a 40s cut — the final frame showed
the authored ~18s state, so the tail scenes (bell card, toast, outro)
were silently missing from the MP4. This, not authoring, is why
rendered cuts kept losing their late scenes. Fix: clip windows are for
structural layers only (hero, panel frame, status); every beat rides
base-hidden styles + delayed CSS animations, which run on the true
clock. Verified frame-by-frame: all four cards, stats, toast, and
outro now land exactly on schedule in both cuts.
Craft, made reusable in the kit instead of per-composition heroics:
- kit.js choreographCursor: data-waypoints="t x y [click]; ..." generates
a multi-leg eased path with fade in/out, an idle-hand sway, and click
rings + glyph press dips — the cursor behaves like a hand, never pops
in, freezes, or blinks out
- kit.js choreographCamera: data-shots="t x y scale; ..." — push-ins
toward each beat's focal point, pull-backs for reveals, settle to end
- release-0.25.0 both cuts re-choreographed: the cursor is the CEO's
hand (settle on the intake while it types, ONE submit click, witness
each card completing, acknowledge the toast, exit off-frame); the
camera lives on every beat; springy card entrances replace flat fades
- motion/README.md gains 'Cinematography & rhythm' (shot-list-first, no
locked-off camera, verify motion with frame PAIRS) + the clip-window
rule; kit/README.md documents both engines; the dev video prompt block
carries the craft bar
Verified end to end through the real sidecar: both cuts render green,
32-frame strips read visually — cursor travels and clicks on schedule,
camera moves, every scene present. motion pnpm test 15/15.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* feat(gateway): delegation detail-fidelity — details survive hand-off, both directions
Details thinned out at every delegation hop: a PM child task mapped to no
parent criterion was legal (coverage only surfaced at submit_up, after the
whole wave ran — a 12-subtask docs tree grew through 8 review rounds that
way, one child titled 'docs page and route wrapper' shipping only the
page), and QA could pass work on a gestalt read (a 4-scene video brief
shipped 3 scenes past every gate because the features existed only in
prose). Three chokepoint gates:
- delegate (down): every child must declare covers_parent_criteria
resolving against the parent's real acceptance criteria — no mapping or
an unresolvable ref rejects naming every offending child and the valid
criteria; the success envelope carries parent_ac_coverage
{covered, uncovered} so a wave-planning PM sees remaining gaps in the
same turn. Full coverage stays enforced at submit_up (waves stay legal).
- pass_review (up): mandatory criteria_verified — one {criterion,
evidence} entry per task AC, matched by the findings ledger's
id-or-exact-text matcher, evidence soup-checked and capped; rejects
naming the unverified criteria; entries render deterministically into
qa_notes as '[AC] <criterion> — verified: <evidence>' lines. The old
count-only ac_verdicts gate is superseded (arg kept for back-compat).
- video briefs (structured detail at origination): an enumerable feature
list (release highlights, or input_props.highlights carried onto a
reject re-author) becomes its own scene acceptance criterion, bounded to
the AC caps; a re-author without highlights carries the
feedback-addressed criterion instead.
Extracted findings.py's criterion matcher into shared unmatched_criteria /
uncovered_acceptance_criteria instead of duplicating it; criteria_verified
joins the WAF free-text exclusion set like findings/issues.
* fix(gateway): break the block/unblock wedge — four hardening fixes from the live PM loop
A cell task looped fe-pm/main-pm block/unblock for hours (10 cycles, 43
spawns): a transient GitHub API error resolving CI became an unwaivable
blocker finding whose own fix text said no code change was required, the
submit freshness guard then demanded a commit no finding called for,
escalate_up auto-blocked, and main-pm's correct recovery plan 422'd on
the approach length cap, degrading it to a bare unblock. Four fixes:
- pr_pass CI-unresolvable refusal is now explicitly transient-worded:
retry pr_pass shortly, do NOT pr_fail over a CI-status lookup error —
a platform blip is not a code finding
- submit freshness guard grants ONE unchanged-head resubmission per
head sha when the findings ledger has zero open rows (all addressed
without code changes) — stamped via the resubmit_unchanged_head
marker so the same head can never loop a second time
- unblock carries a flip breaker: block_flip_count marker, and at the
third flip a one-shot CEO notification flags the task as structurally
wedged (unblock itself still succeeds — the breaker signals, it does
not wedge recovery)
- i_will_plan's approach cap truncates at 800 chars instead of
rejecting — an over-detailed plan must never cost the PM its turn
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* fix(release): CI wait polls the prod rung; escape the header tooltip apostrophe
get_latest_ci_conclusion defaults to the ladder's head rung, so
wait_for_ci searched slave for a release commit that lives on master
and timed out after 40 minutes with the run already green. The wait
now passes the prod branch explicitly. Also fixes the
react/no-unescaped-entities error that turned master's Panel CI red.
* fix(panel,video): dead dialog triggers behind tooltips; dotted composition ids render
HelpTip nested inside a Dialog/AlertDialog trigger puts the trigger's
click handler on the Tooltip root, which renders no DOM — the agents
Spawn item and the KB Reindex-All / Delete-index confirms were dead.
Tooltips now wrap the triggers. The video renderer accepts interior
single dots in composition ids (release-0.25.0) with '..' still
unrepresentable, and propose_video refuses an unrenderable id at
authoring time.
* fix(dispatch): restart-safe PM review turns
A leaf task in awaiting_pm_review had no periodic pickup: the closure
dispatcher bailed on childless tasks and skipped PR-bearing review
tasks as already-promoted, assuming the submit-time PM session was
still alive — an assumption every restart breaks. Proven live on the
docs-sync leaf after the 0.25.0 redeploy, which also dependency-blocked
its sibling dev task. Childless awaiting_pm_review tasks now flow to
the PM's review turn, and the merge turn respawns its PM when none is
active.
* feat(video): verify the rendered artifact, not the source
The 14s release-0.25.0 cut shipped with only one of four scenes visibly
registering: the dev authored DOM, the smoke asserted DOM, QA read code —
nobody consumed the rendered MP4 before the CEO did. Close that loop, and
the reject loop behind it:
- sidecar frames mode: POST /render with frames=1..32 renders the cut,
ffprobes the REAL duration, extracts midpoint-sampled keyframe PNGs
(timestamps in filenames), streams a tar.gz back with X-Video-Duration
- request_render do-verb (developer/QA, request_sandbox's shape): renders
the caller's ACTUAL composition — dev's own worktree (head_sha/dirty
provenance), QA a read-only git-archive export of the assembled branch —
extracts frames to the container-shared .previews/ path, stamps the
render_preview marker, returns the paths as envelope evidence
- gate: i_am_done on a source=video task refuses without a stamped
render_preview (Requirement.RENDER_VERIFIED; canonical source string
moved to foundation as markers.VIDEO_TASK_SOURCE; mirrored in the
possibilities-matrix fast path so it cannot bypass the check)
- QA claim_review evidence carries video_context (composition id, the
dev's preview, a re-render instruction) so review checks output
- dev spawn prompt block + a 4th authoring AC order Read-every-frame
verification before submitting
- reject -> re-author: a CEO reject with a reason opens a fresh authoring
task carrying the verbatim feedback + a revise-in-place pointer at the
existing composition (best-effort, never fails the reject) — rejection
feedback no longer dies on the cancelled draft
E2E: rendered the committed release-0.25.0 composition through the new
frames mode locally — the returned keyframes show exactly the reported
failure (blank frame at 5.8s, only 'Env ladder' by 12.8s), the check the
fleet was missing.
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* fix(release): CI wait polls the prod rung; escape the header tooltip apostrophe
get_latest_ci_conclusion defaults to the ladder's head rung, so
wait_for_ci searched slave for a release commit that lives on master
and timed out after 40 minutes with the run already green. The wait
now passes the prod branch explicitly. Also fixes the
react/no-unescaped-entities error that turned master's Panel CI red.
* fix(panel,video): dead dialog triggers behind tooltips; dotted composition ids render
HelpTip nested inside a Dialog/AlertDialog trigger puts the trigger's
click handler on the Tooltip root, which renders no DOM — the agents
Spawn item and the KB Reindex-All / Delete-index confirms were dead.
Tooltips now wrap the triggers. The video renderer accepts interior
single dots in composition ids (release-0.25.0) with '..' still
unrepresentable, and propose_video refuses an unrenderable id at
authoring time.
* fix(dispatch): restart-safe PM review turns
A leaf task in awaiting_pm_review had no periodic pickup: the closure
dispatcher bailed on childless tasks and skipped PR-bearing review
tasks as already-promoted, assuming the submit-time PM session was
still alive — an assumption every restart breaks. Proven live on the
docs-sync leaf after the 0.25.0 redeploy, which also dependency-blocked
its sibling dev task. Childless awaiting_pm_review tasks now flow to
the PM's review turn, and the merge turn respawns its PM when none is
active.
* [1dae04a7] Revise release-0.25.0 composition to 40s scene-based pacing with four feature cards
* [1dae04a7] docs(motion): update release-0.25.0 README section for 40s four-card revision
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
Co-authored-by: UX/UI Developer 1 <ux-dev-1@roboco.tech>
Co-authored-by: UX/UI Documenter <ux-doc@roboco.tech>
* fix(release): CI wait polls the prod rung; escape the header tooltip apostrophe
get_latest_ci_conclusion defaults to the ladder's head rung, so
wait_for_ci searched slave for a release commit that lives on master
and timed out after 40 minutes with the run already green. The wait
now passes the prod branch explicitly. Also fixes the
react/no-unescaped-entities error that turned master's Panel CI red.
* [2f806123] feat(motion): add release-0.25.0 composition extending panel-demo register
* [2f806123] fix(scope): revert out-of-scope backend and panel changes from video branch
* [2f806123] fix(scope): revert out-of-scope backend and panel changes from video branch
* [2f806123] fix(scope): restore out-of-scope files from current origin/master after stale-master revert
* [2f806123] fix(scope): restore motion/pnpm-workspace.yaml from origin/master
* [2f806123] docs(motion): add release-0.25.0 composition example to README
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
Co-authored-by: UX/UI Developer 1 <ux-dev-1@roboco.tech>
Co-authored-by: UX/UI Documenter <ux-doc@roboco.tech>
Pure derive_pr_labels (foundation/policy/pr_labels.py) maps a PR's shape
to a stable org-structure label set: to master/to slave (is_root_pr
discriminator), root, MegaTask, and the owning layer (main-pm /
cell/{team} / subtask/{team}). Mirrors batch.py: object|None inputs,
enum-or-string normalization, no DB/I/O. Full slave-targeting semantics
(base_branch vs default_branch) land with the slave/master wiring (W-H);
YAGNI now.
GitService._apply_pr_labels posts the result to the GitHub labels API
best-effort (create-before-add, swallow 422/409, never raises) so a label
failure can never block PR creation. Wired at all three PR-opening sites:
create_pr (gateway path), create_pull_request (REST/task path), and
_push_and_open_conventions_pr (static chore label). Existing PR tests
mock _apply_pr_labels so they never hit the real labels API.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
Reject cancelled the proposal's status but never moved it out of the
held-proposal set, so the one-open-proposal dedup blocked the release
manager from ever re-assessing — a rejected proposal deadlocked the cycle.
reject() now sets CANCELLED (mirroring video_post_service), which
list_open_release_proposals already excludes, so a fresh proposal can
originate next cycle.
A failed ~40min background execute (gate red, CI red, or an unexpected
crash) left the proposal silently PENDING with no signal to the CEO.
_run_approve_background now writes a release_execute_outcome marker
(status + detail) on every terminal outcome, and an 'error' marker on an
unhandled exception. GET /proposal surfaces execute_status / execute_detail
/ execute_in_flight (derived from the in-memory _INFLIGHT_APPROVES registry)
so the panel can show a running badge, a failure block with the reason, and
a Retry-approve label instead of a silent wait.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* feat(gateway): reviewer/PM collision map (W5)
The collision surface (intends_to_touch / adds_migration / touches_shared)
is authored at delegate time, consumed once by SequencingService to wire
dependency edges, then never shown to a reviewer again. This surfaces it:
- Pure builder (services/gateway/choreographer/collision.py): for a task
under review, the surfaced siblings (same parent) that would collide —
file-overlap globs or a shared migration chain (both adds_migration) —
with the overlapping globs and a declared-vs-actual drift check. No
DB/IO; callers fetch siblings (one indexed get_subtasks query, mig 069)
+ actual files (git). Caps: 10 siblings, 5 globs.
- Evidence envelopes: collision_context block injected into QA
claim_review, PR-gate claim_gate_review (both carry real touched files
so drift is populated), and the PM i_will_plan briefing (no actual
files at plan time, drift omitted). Best-effort — a failure omits the
block, never breaks the verb/briefing. Empty block omitted (zero token
cost via _EVIDENCE_OMIT_WHEN_EMPTY).
- Panel: GET /api/tasks/{id}/collision-map (declared surface + sibling
overlap; no drift — the panel route resolves no workspace) + a Collision
tab on the task detail (8th tab). Mock-mode returns an empty map.
- docs/map added to the RAG auto-index dirs so the collision-map concept
is fleet-retrievable; skipped gracefully if the dir is absent.
19 new tests (15 unit on the pure builder + 4 integration on the route).
Gate green: ruff/mypy/xenon (module rank A)/pytest 13000/coverage 94.81%,
panel typecheck/lint/516 tests.
* [w6-telegram] Add Telegram notifications bridge (V1)
CEO-facing Telegram DM bridge, flag-gated off by default
(ROBOCO_TELEGRAM_ENABLED). Mirrors the X-credentials / X-client pattern:
- TelegramCredentialsTable (migration 073) — singleton Fernet-encrypted
bot_token + chat_id, all-or-nothing set/clear; API never returns plaintext.
- TelegramClient ABC / NullTelegramClient (no-op, configured->False, never
raises) / LiveTelegramClient (httpx POST sendMessage) / build_telegram_client
factory (Null when creds unset).
- /telegram/credentials CEO-only routes (write-only, guard-decorated).
- Best-effort _notify_telegram fan-out from the two CEO-notify producers
(notify_ceo_of_escalation, notify_ceo_of_completion) — guarded by the flag,
never raises into the producer, carries a panel deep-link when
panel_base_url is set.
- panel credentials card (2 fields) nested in the Telegram feature-flag row.
- panel_base_url + telegram_timeout_seconds config fields.
V1 scope only: credentials + flag + panel card + client + one-line fan-out.
Out of scope (V2): inbound commands, a TelegramEngine background loop, a
dedup ledger, a bus subscription.
* [w6-telegram] fix: slave mypy/xenon regression (product tests + helper extract)
Pre-existing on slave from prior session's merges — no PR's CI caught them
(squash merges don't re-CI the result; each branch was based on older slave).
- test_product: _product helper returned MagicMock -> list invariant error;
cast to ProductTable, move import under TYPE_CHECKING.
- test_usage: svc.session.execute (AsyncSession) has no call_args_list;
cast to MagicMock at the two call sites.
- product.progress_for_products: xenon rank C -> extract module-level
_project_to_products_map helper (repo pattern: helper-extract).
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* [env-bran] EnvSyncEngine: orchestrator-side prod→dev cascade (default-off)
- EnvSyncEngine mirrors CiWatchEngine: cascade ladder_pairs top-down via
GitHub merges API; clean→auto-push lower rung, conflict→one sync PR +
tracked MAIN_PM task + stop. Never pushes prod (lower rung is never prod
by construction).
- GitService.sync_env_branch (merges API) + open_sync_pr (idempotent) +
_env_merge_status/_post_sync_pr helpers (constants for 201/204/409).
- TaskService.ENV_SYNC_SOURCE + list_open_env_sync_tasks (per-repo dedup).
- config env_sync_enabled/_interval_seconds(1800)/_max_open_tasks(3)/_max_per_cycle(1).
- Orchestrator 4-touch registration + _load_env_sync_set (ladder+token opt-in).
- Feature-flags card + settings FEATURE_FLAGS entry for ROBOCO_ENV_SYNC_ENABLED.
* [env-bran] Panel: environment ladder editor + types + validation
- EnvironmentRung type + environments on Project/ProjectCreate/ProjectUpdate.
- EnvironmentLadderEditor (plain useState, add/remove/up-down reorder, head/
prod labels) reused by create + edit project dialogs.
- validateLadder (non-empty name+branch, no duplicate branches) shared,
toast.error on submit; empty editor => null => inherits default_branch shim.
- default_branch input kept with override-hint; API client passthrough.
- 6 unit tests for validateLadder.
* [env-bran] Tests + gate green: env ladder, EnvSyncEngine, promotion chain
- tests/unit/models/test_env_branches.py: shim, head/prod, ladder_pairs,
promotion_chain, normalize (20 tests)
- tests/integration/services/test_env_sync_engine.py: cascade clean/conflict/
missing_ref/tokenless/degenerate/caps/dedup/disabled (9 tests, DB)
- tests/integration/test_migration_env_branches.py: 073 defaults null + round-trip
- tests/unit/services/test_release_executor*.py: add env_chain=[] to
_ReleaseContext constructions (promotion_chain field is now required)
- tests/unit/runtime/test_orchestrator_shutdown_drain.py: register _env_sync_task
in the stop()-drain fixture (new named background loop)
- roboco/services/git.py: revert _project_head_branch rename back to
_project_default_branch (modify-in-place per plan); the rename in the
consumers commit broke ~15 unit-test mocks that bind the original name
- roboco/services/env_sync_engine.py + models/env_branches.py: ruff format
- roboco/api/schemas/project.py: trailing-newline format
Backend gate green (13013 passed / 439 skipped), mypy clean, ruff clean.
Panel gate green (typecheck/lint/522 tests).
* [env-bran] fix: add env_chain to _ReleaseContext in e2e smoke (CI red)
The release-executor promotion_chain change made _ReleaseContext.env_chain
required. I fixed the three unit/release test files but missed the
construction in tests/e2e_smoke/test_background_engines.py:98 — my local
gate ran 'mypy roboco/' (excludes tests/) and I skipped 'make e2e-smoke',
so CI's mypy-on-tests + the e2e runtime job caught it instead of me.
Verified locally with the CI-equivalent gates:
uv run mypy roboco/ tests/ -> 1170 files, clean
ROBOCO_E2E_SMOKE=1 uv run pytest tests/e2e_smoke -> 50 passed, 1 skipped
* [env-bran] fix: extract _ensure_prod_fetched to clear xenon rank C (CI red)
_production_assess grew past xenon --max-absolute B (rank C) when the
env-branches prod-tip fetch added an if/try/except branch. Extracted the
fetch-with-fallback into _ensure_prod_fetched (degan+fetch paths), moved
_run_git to the module-level import. Local make quality green (all gates
incl xenon/vulture/deptry/import-linter/foundation-check).
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* [W9-5a] Tooltip sweep foundation: HelpTip helper + shared cryptic badges
HelpTip: DRY wrapper over the verbose 3-element Radix Tooltip pattern so a
broad sweep stays a one-line wrap per site (falsy label short-circuits to the
bare child). Unit-tested (3 cases).
TaskStatusBadge + AgentStateBadge: the panel's most cryptic, most-frequent
elements (15 task lifecycle states, 11 agent states) had no explanation
anywhere. Add a per-state tooltip via HelpTip, with the canonical text in one
description map and exported as taskStatusDescription / agentStateDescription
so the per-view inline renderers (kanban, task header) reuse it instead of
re-declaring. This is part 1 of the W9-5 tooltip sweep; the per-view inline
surfaces follow in subsequent PRs.
* [W9-5b] Per-view tooltip sweep: decode cryptic badges, icon-only buttons, status dots
35 HelpTip additions across 22 panel components, reusing the W9-5a
helper plus taskStatusDescription/agentStateDescription. Tipped:
task-id/commit-hash/branch/PR badges, severity/origin/status badges,
priority (P0-P3), migration/shared flags, MegaTask umbrella badge,
Review Gate / For Resumption / Confidential note badges, icon-only
view/delete/edit/clear/show-hide buttons, semver bump + gate-state
badges, ahead-not-pushed badge. Skipped self-explanatory labeled
buttons and elements already carrying title=.
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
Backend: GET /git/file reads a file at a branch tip (read_file_at_branch)
and slices it to a line window — explicit start/end, a line+context center,
or the whole file capped at 2000 lines. _compute_file_range is the pure
helper (unit-tested).
Frontend: useGitFile hook + CodeSnippet (styled <pre>, line numbers, active-
line highlight — matches git-diff-viewer, no shiki). Wired into FindingCard
so each file:line finding shows the surrounding source. Fail-open: a missing
file renders a muted hint, never breaks the card.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
Backend: ProjectSummaryResponse gains task_counts (done/active/blocked) + ci_watch_enabled. ProjectService.task_counts_for_projects does one GROUP BY project_id over TaskTable for every distinct project_id in the list (a project with no tasks is absent — route falls back to None). ci_watch_enabled is read straight off the Project row (already a column) — a 0-cost schema extension, honest signal that CI-watch is armed, no live-conclusion fan-out. project_to_summary takes an optional task_counts. No migration.
Frontend: ProjectTable gains a Tasks column (done/active/blocked + health dot, amber at-risk when blocked>0) and a CI-Watch badge under the project name when ci_watch_enabled. Both desktop Table and mobile ResponsiveTableCard variants. Mock projects carry the new shape (two sample repos).
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
Backend: ProductSummaryResponse gains cells: [{team, project_id, project_name}] and progress: {done, active, blocked}. ProductService.progress_for_products does one grouped query over tasks for every distinct project_id any product references, summed per product (monorepo case dedups a project once per product via a seen set). list_all eager-loads cells + each cell's project (selectinload + joinedload) so product_to_summary reads project.name without an N+1. No migration — reads existing tasks.status + product_projects.
Frontend: ProductTable renders a Cells column (team badges + project names, Unmapped when empty) and a Progress column (done/active/blocked counts + a health dot: amber at-risk when blocked>0). Both desktop Table and mobile ResponsiveTableCard variants. Mock products carry the new shape.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
Backend: optional agent_slug filter on GET /usage/time-series + UsageService.get_time_series (AgentSpawnSessionTable.agent_slug column already exists — no migration).
Frontend: AgentActivityPanel on the agent detail page — a 7d per-agent token sparkline (recharts AreaChart) + a merged work-session/journal activity timeline. Work-sessions filter by the agent UUID (WorkSessionTable.agent_id is a UUID FK to agents.id), journals by slug. List grid left as-is (avoids 25-agent fan-out). Card last_active deferred (no live hook populates AgentMetrics).
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
Backend: widen usage _PeriodType to 24h/7d/30d/90d and add a 90d branch to
_parse_period (daily buckets already cover it). TestParsePeriod pins the
contract per window.
Frontend: UsagePeriod += 90d with a scaleFor helper (replacing 6 inline
ternaries) and 90 daily mock points. One generic SegmentedControl primitive
(reuses Radix Tabs) drives both the metrics time-window selector
(24h/7d/30d/90d) and the per-chart Chart/Table view toggle — one file, two
roles. The Token Usage & Costs tab drops 8 hardcoded '24h' hooks for a
single period state + selector; the stale '(24h)' cost-card parenthetical
goes too. The Performance landing tab gains a TaskStatusChart donut fed by
the status counts already on the page (no new hook). Agent/team bar charts
gain an inline table view.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
The bell showed a transient WebSocket-stream buffer count with no
read/ack actions; the persisted unread/ack state and the mark-read /
acknowledge / mark-all-read mutation hooks already existed
(use-notifications.ts) and powered the notifications page, but the bell
ignored them.
The bell now derives its badge from useNotifications().unread_count (the
real DB count, not the stream buffer), renders the recent items with
per-item Mark Read + Acknowledge + a header Mark all read, shows the
pending-ack count, and keeps the WS stream only for the connection
indicator (the NotificationAlerts sibling still owns the toast/chime).
The stream buffer is cleared on popover close so it can't grow unbounded
now that it's no longer displayed.
The mutations self-invalidate notificationKeys.all on success, so the
badge + popover refresh immediately after each action; useNotifications
also refetches every 30s.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* [W7] Add possibilities_matrix_enabled feature flag (default off)
* [W7] Add _work_appears_done predicate (status+commits+PR+ACs+no-open-findings)
* [W7] Add CI-green quality proxy for the fast path (local fallback on no-CI)
* [W7] Add work-already-done fast path in i_am_done (slimmed gates, no rich plan)
* [W7] Add WORK_ALREADY_DONE prompt state
* [W7] Make fast path mypy-clean (cast to helpers for _resolve_ci_status; typed mock locals)
* [W7] Extract _all_criteria_addressed to bring _work_appears_done under xenon B
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
docs/map is the agent-facing exhaustive codebase map (CLAUDE.md) but was never
RAG-indexed — only docs/rag was. Add it to OptimalService._auto_index_dirs so
every docs/map/*.md rides index_documentation (the generic _index_docs_directory
rglobs *.md and routes only the 'standards' subdir to the standards indexer) and
becomes roboco_kb_search-able. No map-specific branch needed.
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* feat(panel): add pickTab helper for validated URL tab params
* fix(panel): validate kanban view + KB tab params via pickTab
The bare `as T || default` cast only guarded null — a typo/invalid value
(?view=deev, ?tab=foo) passed through as an out-of-set TabValue, blanking the
active tab highlight and the content pane. pickTab validates against the
known set and falls back to the default on null/empty/invalid.
* fix(panel): highlight active sidebar footer link (exact match)
SidebarFooter had no isActive branch (SidebarNav does), so footer links
never highlighted. Added with EXACT match (pathname === href) — not
startsWith — so /settings does not also highlight on /settings/ai-providers.
Main nav keeps startsWith (longer hrefs need prefix matching); the two
intentionally differ.
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* fix(dispatch): prefilter sequence-held dev tasks before spawn
_spawn_pending_dev booted a full dev container for a pre-assigned pending task
that the assignee-blind sequence guard would refuse at the claim chokepoint (a
non-terminal lower-sequence same-parent sibling — not a declared dependency).
_blocked_by_earlier_lane_sibling is narrower (same dev's lane) and
_validate_task_for_spawn checks declared deps, not sequence siblings, so the
container spawned, the first claim hit _claim_blocked_by_sequence and was
refused, and the agent exited only to be re-spawned next tick — pure churn
until the predecessor went terminal.
Mirror the PM path's _pending_claim_blocked prefilter (the exact claim-gate
predicate, fails open) at the top of _spawn_pending_dev, before the narrower
per-dev lane probe. Reuses the helper so it can't drift from the chokepoint.
* fix(dispatch): prefilter sequence-held dev tasks before spawn
_spawn_pending_dev booted a full dev container for a pre-assigned pending task
that the assignee-blind sequence guard would refuse at the claim chokepoint (a
non-terminal lower-sequence same-parent sibling — not a declared dependency).
_blocked_by_earlier_lane_sibling is narrower (same dev's lane) and
_validate_task_for_spawn checks declared deps, not sequence siblings, so the
container spawned, the first claim hit _claim_blocked_by_sequence and was
refused, and the agent exited only to be re-spawned next tick — pure churn
until the predecessor went terminal.
Mirror the PM path's _pending_claim_blocked prefilter (the exact claim-gate
predicate, fails open) at the top of _spawn_pending_dev, before the narrower
per-dev lane probe. Reuses the helper so it can't drift from the chokepoint.
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* fix(prompts): point agents at Makefile, drop raw uv run instructions
backend.md:23-26 literally instructed raw uv run ruff/mypy/pytest (copied from
the human-facing CLAUDE.md), so agents bypassed the Makefile's UV_NO_SYNC=1 +
private UV_CACHE_DIR venv-corruption guard. Replace with make targets across
backend/developer/qa/cell_pm + a universal rule in base.md. Regenerate verbs.md
from the updated regen script (baked instruction now make foundation-check) and
align the Makefile drift message. Ships with the bash-guard deny in the next
commit so agents don't loop fighting the guard.
* feat(bash-guard): deny raw uv/pip/conda/poetry, point at Makefile
When a Makefile is present, deny raw uv run/uv pip/uv lock/add/remove, pip/pip3
install/uninstall, conda install/create/run, poetry run/install/add and remediate
to make quality/gate/lint/test. Skipped when no Makefile (Makefile-less projects
not blocked). ROBOCO_GUARD_SKIP_PM=1 (grok path) nudges exit 0 instead of the
run-canceling exit 2. Overrides the prior bare-uv-run-allowed stance by CEO
direction; the /app-targeted blocks above keep priority.
* feat(grok): deny raw uv/pip/conda/poetry via native --deny + PM-skip nudge
Add _RAW_PM_DENY (uv run/pip install/lock/add/remove, pip/pip3 install, conda
install/create/run, poetry run/install/add) to _deny_rules so grok's graceful
native --deny blocks raw package-manager commands (model adapts to make, run
continues — unlike a hook deny which cancels the run). The bash-guard hook
keeps the compound-command fallback (cd x && uv run) and nudges exit 0 there via
ROBOCO_GUARD_SKIP_PM=1 in the grok hook env, never canceling.
* test(bash-guard): align existing tests with W1 Makefile-gate policy
Raw uv run / pip install are now Makefile-gated (W1, CEO item #15), so two
existing bash-guard invariants reverse:
- test_allows_pytest_even_if_suite_uses_requests keeps its HTTP-injection
allow-path intent but uses bare `python -m pytest` (raw `uv run` is now
denied); the deny case is covered by test_bash_guard_makefile_guardrail.
- test_allows_pip_install_in_workspace -> test_denies_pip_install_when_makefile_
present: a workspace clone carries a Makefile, so bare pip install is now
denied -> agents use `make` / `uv sync --extra dev`. Makefile-less skips
stay covered.
Gate: 12994 passed, 439 skipped, 94.81% cov (DB env :55432 user renzof);
the lone flaky integration error passes in isolation (DB-state race, not W1).
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>