Files
roboco/docs/rag/architecture/video-engine.md
T
aa15dc40cc Feature/video artifact verification (#537)
* fix(release): CI wait polls the prod rung; escape the header tooltip apostrophe

get_latest_ci_conclusion defaults to the ladder's head rung, so
wait_for_ci searched slave for a release commit that lives on master
and timed out after 40 minutes with the run already green. The wait
now passes the prod branch explicitly. Also fixes the
react/no-unescaped-entities error that turned master's Panel CI red.

* fix(panel,video): dead dialog triggers behind tooltips; dotted composition ids render

HelpTip nested inside a Dialog/AlertDialog trigger puts the trigger's
click handler on the Tooltip root, which renders no DOM — the agents
Spawn item and the KB Reindex-All / Delete-index confirms were dead.
Tooltips now wrap the triggers. The video renderer accepts interior
single dots in composition ids (release-0.25.0) with '..' still
unrepresentable, and propose_video refuses an unrenderable id at
authoring time.

* fix(dispatch): restart-safe PM review turns

A leaf task in awaiting_pm_review had no periodic pickup: the closure
dispatcher bailed on childless tasks and skipped PR-bearing review
tasks as already-promoted, assuming the submit-time PM session was
still alive — an assumption every restart breaks. Proven live on the
docs-sync leaf after the 0.25.0 redeploy, which also dependency-blocked
its sibling dev task. Childless awaiting_pm_review tasks now flow to
the PM's review turn, and the merge turn respawns its PM when none is
active.

* feat(video): verify the rendered artifact, not the source

The 14s release-0.25.0 cut shipped with only one of four scenes visibly
registering: the dev authored DOM, the smoke asserted DOM, QA read code —
nobody consumed the rendered MP4 before the CEO did. Close that loop, and
the reject loop behind it:

- sidecar frames mode: POST /render with frames=1..32 renders the cut,
  ffprobes the REAL duration, extracts midpoint-sampled keyframe PNGs
  (timestamps in filenames), streams a tar.gz back with X-Video-Duration
- request_render do-verb (developer/QA, request_sandbox's shape): renders
  the caller's ACTUAL composition — dev's own worktree (head_sha/dirty
  provenance), QA a read-only git-archive export of the assembled branch —
  extracts frames to the container-shared .previews/ path, stamps the
  render_preview marker, returns the paths as envelope evidence
- gate: i_am_done on a source=video task refuses without a stamped
  render_preview (Requirement.RENDER_VERIFIED; canonical source string
  moved to foundation as markers.VIDEO_TASK_SOURCE; mirrored in the
  possibilities-matrix fast path so it cannot bypass the check)
- QA claim_review evidence carries video_context (composition id, the
  dev's preview, a re-render instruction) so review checks output
- dev spawn prompt block + a 4th authoring AC order Read-every-frame
  verification before submitting
- reject -> re-author: a CEO reject with a reason opens a fresh authoring
  task carrying the verbatim feedback + a revise-in-place pointer at the
  existing composition (best-effort, never fails the reject) — rejection
  feedback no longer dies on the cancelled draft

E2E: rendered the committed release-0.25.0 composition through the new
frames mode locally — the returned keyframes show exactly the reported
failure (blank frame at 5.8s, only 'Env ladder' by 12.8s), the check the
fleet was missing.

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-07-16 19:49:26 +02:00

6.8 KiB

Video Engine (HyperFrames)

What It Is

RoboCo can author and post short marketing videos for the company's X (Twitter) and TikTok accounts — a release announcement, a feature spotlight, or an on-demand CEO request — implemented in roboco/services/video_engine.py (VideoEngine), roboco/services/video_post_service.py (VideoPostService), and the motion/ composition package, rendered by a credential-free video-renderer sidecar. Every clip is HELD for an explicit, per-clip CEO approval; nothing is ever posted automatically. It mirrors the XEngine "detect → originate a CEO-gated artifact → hold" shape, extended with a render pass between authoring and the held draft.

Enable/Disable

Variable Default Effect
ROBOCO_VIDEO_ENGINE_ENABLED false Master switch. Off = no video-authoring task is ever opened and no render/post happens. Panel-toggleable (Settings → Feature Flags).
ROBOCO_VIDEO_ON_RELEASE false Sub-switch: open an authoring task when a release publishes. Off even with the master switch on.
ROBOCO_VIDEO_ON_SPOTLIGHT false Sub-switch: open an authoring task when the CEO approves a feature-spotlight draft that requests one. Off even with the master switch on.

Even when enabled, distribution requires an explicit per-clip CEO approval — a second independent gate beyond the flag.

Three trigger sources

Release videos (event-driven). A release publish fires the video-on-release trigger — gated by ROBOCO_VIDEO_ON_RELEASE independently of the master switch.

Feature-spotlight videos (event-driven). The CEO's approval of a feature-spotlight draft that requests a video fires the video-on-spotlight trigger — gated by ROBOCO_VIDEO_ON_SPOTLIGHT.

On-demand CEO request (panel-triggered). POST /api/video/request (CEO-role-gated) opens an authoring task directly — the CEO's escape hatch independent of the two automatic triggers.

All three open a normal, assigned UX/UI authoring task (balanced across the two ux-devs) rather than a held draft — the dev builds a HyperFrames HTML composition under motion/compositions/<id>/ (one <id>/vertical.html + <id>/square.html carrying the HyperFrames render params on <html>) and proposes its composition id + per-platform captions via the team-gated propose_video do-tool, then ships it through the standard commit/PR/QA/doc/review lifecycle.

Artifact verification (request_render)

Authoring is gated on the RENDERED artifact, not just its source: the request_render do-tool (developer/QA) renders the caller's actual composition through the sidecar and extracts evenly spaced keyframe PNGs to a container-shared .previews/ path, returning their absolute paths in the envelope's evidence.frames. The agent must Read every frame and verify each scene/feature from the brief appears fully and legibly — a 14-second cut that only ever shows its first scene is exactly what this catches. A developer renders their own working tree (worktree-aware, head_sha/dirty provenance stamped); QA renders a read-only git archive export of the assembled branch — never a working tree. A successful render stamps the task's render_preview marker, and i_am_done on a video-authoring task refuses without it (Requirement.RENDER_VERIFIED, mirrored in the possibilities-matrix fast path), so no video task can complete on a source-only self-review. QA's claim_review evidence carries a video_context block (composition id + the dev's stamped preview + an instruction to re-render the branch state) so the reviewer checks output, not source.

Render loop and the sidecar

Once the authoring task completes, an orchestrator render loop (_video_render_loop, interval ROBOCO_VIDEO_RENDER_INTERVAL_SECONDS) tars the merged motion/ source and POSTs it to the video-renderer sidecar (ROBOCO_VIDEO_RENDERER_BASE_URL). The sidecar is credential-free and git-free — it reads only what it's POSTed, bundles the composition, renders both the 9:16 and 1:1 MP4 cuts (ROBOCO_VIDEO_RENDER_TIMEOUT_SECONDS per render), and streams the bytes back. The orchestrator writes the MP4s to ROBOCO_VIDEO_OUTPUT_DIR (bind-mounted in all three compose files so renders survive container recreation) — the sidecar never writes to disk directly. The sidecar renders via @hyperframes/producer — agent-authored HTML rendered in headless Chrome plus system ffmpeg for the seek-driven, frame-deterministic MP4 cut. HyperFrames was chosen for its Apache-2.0 license (no React/TSX toolchain lock-in for the authoring agents), agent-authored HTML as the composition substrate, and seek-driven deterministic renders that match the post-pipeline shape without a node-bundled renderer.

Ownership and the CEO gate

The render pass materializes a held video_post draft — mirroring the X-post/release-proposal shape: team=main_pm, assigned_to=secretary-1, confirmed_by_human=False (HELD — skipped by every dispatcher, never delivered to an agent). The CEO previews (an axios-blob fetch that carries X-Agent-ID/X-Agent-Role), edits, approves, or rejects each draft in a new panel video queue.

Endpoint Effect
GET /api/video/posts List every held video draft awaiting decision.
POST /api/video/posts/{task_id}/approve Post the rendered clip to X (native video, v2 media upload) and/or TikTok (inbox upload). Idempotent — approving an already-posted draft is a no-op.
POST /api/video/posts/{task_id}/reject Cancel the draft with a reason — and route the feedback back into the flow: a non-empty reason opens a fresh authoring task (same occasion, brief = the CEO's verbatim feedback + a revise-in-place pointer at the existing composition), so a rejection is a rework loop, not a dead end.
POST /api/video/request Open an on-demand authoring task (CEO-only).

Approval runs under a Redis heartbeat-renewed lock so a double-click can't double-post; the task is marked COMPLETED under the same lock before it releases.

Credentials

TikTok's OAuth2 secrets are Fernet-encrypted (ROBOCO_ENCRYPTION_KEY) alongside the existing X credentials, entered in the panel only — never in .env or an agent-visible setting. Every unconfigured leg (renderer, X, TikTok) degrades to a graceful no-op rather than a crash — the same graceful-null pattern as the X NullXClient.

Media route confinement

The panel's media fetch for video previews is confined to a single media route that carries the agent-identity headers and streams bytes from ROBOCO_VIDEO_OUTPUT_DIR — it cannot traverse the filesystem.

  • docs/rag/architecture/x-engine.md — the held-draft marketing engine this mirrors and extends
  • docs/rag/architecture/config-reference.md — full env var table
  • docs/rag/roles/ceo.md — the CEO approval queues