mirror of
https://github.com/rennf93/roboco.git
synced 2026-08-03 07:23:24 +02:00
* feat(video): Phase A — VideoEngine origination spine + held-source gates
New default-off engine skeleton: opens a UX/UI authoring task (source=video, assigned to a ux-dev, LOW complexity to clear the dev-needs-subtasks guard) and materializes a held CEO-approval draft (source=video_post). Excludes video_post from all three held-source skip sites; adds the video_draft marker, six config flags, and the feature-flag entries. Origination + gate behavior unit-tested.
* refactor(orchestrator): fold _dispatch_dev_work skip chain into a helper
The per-source if/continue chain grew past xenon's --max-absolute B when the video_post held source joined it. Extract _is_non_dev_dispatch_source (every held-CEO source plus the two Board exploration sources) so the dev loop's skip is one flat call. Behavior-identical.
* feat(video): Phase B — propose_video do-tool (metadata-only, team-gated)
UX/UI dev records a video's composition ref + per-platform captions onto the authoring task's video_draft marker. Team-gated at runtime via _caller_team (Role.DEVELOPER can't tell a ux-dev from a be-dev). Resolves the caller's ACTIVE task via get_active_task_for_agent, not an oldest-first scan that would clobber a second open video task. Metadata only, no render. Wired through do_server + route + schema; added to _DEV_DO.
* feat(video): Phase D — render loop + RemotionRenderer client
Orchestrator-async _video_render_loop renders a completed authoring task's merged composition to MP4 (vertical + square) via the remotion-renderer sidecar and materializes the held video_post draft. RemotionRenderer tars the read-clone's motion/ source, POSTs it, and saves the returned MP4 bytes to a TASK-scoped local path (no shared volume; a composition is reused across videos so a composition-scoped path would clobber an earlier draft). Render failures bounded-retry (read-clone catch-up window, transient sidecar) up to a cap, then terminal-fail. Client tested vs a mock transport; loop vs a mock renderer + real DB.
* feat(video): Phase C — release / spotlight / on-demand video triggers
Three entry points open a UX/UI video-authoring task via VideoEngine.open_video_task: (1) a published release drafts a companion video — best-effort in ReleaseProposalService.approve, never fails the publish; script from the CHANGELOG via the local model with a template fallback. (2) propose_feature_spotlight gains optional wants_video/video_script — best-effort, gated on video_on_spotlight, default-off leaves the spotlight flow byte-for-byte unchanged. (3) POST /video/request (CEO-only) for an on-demand brief, with clean disabled/not_opened responses. All gated on video_engine_enabled.
* fix(video): savepoint-isolate video-task inserts (F042 poisoned session)
The best-effort try/except around open_video_task (release-publish + spotlight hooks) swallowed the Python exception, but a DBAPI error at the insert flush left the shared session must-rollback — so the caller's next commit (release finalize / request boundary) threw PendingRollbackError: the release stuck 'pending' after actually publishing, or the spotlight draft + HTTP response were lost. Wrap both inserts (open_video_task, _originate_video_post) in a begin_nested savepoint (the repo's established F042 pattern) so a DB error rolls back only the insert. open_video_task returns None (every caller already handles it); _originate_video_post propagates to the render loop's handler. Regression test: an insert FK error returns None with the session left usable. Dormant while the flags were off; armed on the NAS.
* feat(video): Phase G — motion/ package + remotion-renderer sidecar + compose
In-repo Remotion v4 motion/ package (ReleaseAnnouncement composition; calculateMetadata returns 1080x1920 vertical / 1080x1080 square from inputProps.orientation) + a credential-free remotion-renderer sidecar: untar the POSTed motion/ source, bundle (LRU-cached per source sha), selectComposition + renderMedia h264, stream the MP4 bytes back — matching the RemotionRenderer client contract. docker/remotion.Dockerfile on Debian (Chrome apt deps, build-time Chrome pre-warm, ffmpeg bundled in @remotion/renderer). Wired into both compose files (roboco_default only, shm_size 1gb, /health check) + the release publish matrix. Verified via a real local render of both cuts; the Debian docker build is the CEO's to run.
* chore(video): D-hardening — video_post source_task_id + render-loop docstring
Add a source_task_id back-reference to the video_post held-draft marker (traceability from a draft to its authoring task; also makes the render loop's two-key idempotency check wireable later). Fix the render-loop test's stale docstring ('never retried' -> bounded-retry). Both from the Phase D critic's non-blocking follow-ups.
* feat(video): Phase E1 — VideoPostService + heartbeat mutex (approve->post)
CEO-approve->post service: heartbeat-renewed Redis mutex (fail-closed, grace=ttl-2*heartbeat), re-read-in-lock double-post guard, per-platform durable commits (asyncio.shield-ed, settle-before-rollback on lock-loss), all writes inside the lock (captions validated pre-lock, applied in-lock — no stale whole-column clobber), idempotent, per-platform retry-skip. Poster interfaces (X/TikTok, mocked here). Reject + list-held-drafts. Survived 3 adversarial rounds; residual = a crash in the poster->commit window (CEO-gated low-freq, documented).
* fix(video): G-hardening — renderer leaks + Share Tech Mono brand font
Sidecar: give bundle() an explicit outDir tracked + deleted on LRU eviction (was leaking ~19MB remotion-webpack-bundle-* per source); res.on('close') cleanup so an aborted/retried download no longer leaks its remotion-out-* MP4 dir. Fonts: vendor Share Tech Mono (roboco-website brand font) as the display face (self-hosted woff2, 400-weight, headline fontWeight 700->400 to avoid faux-bold) + self-hosted Inter body — no gstatic fetch at render time (lsof-verified). Extras: composition_id whitelist (400) + Multer error middleware (400/413).
* feat(video): Phase E2 — X v2 + TikTok posters, tiktok_credentials, routes
LiveXVideoPoster (X v2 chunked media upload: init/append/finalize/STATUS-poll -> tweet w/ media_ids, OAuth1 signer reused). LiveTikTokPoster (OAuth2 inbox: init -> chunked PUT with asymmetric final chunk -> status-fetch; 401 -> refresh_token grant, rotated token persisted). tiktok_credentials Fernet singleton + migration 062 (single head). Routes: CEO approve/reject + list held drafts + write-only tiktok creds, wiring real posters into VideoPostService. Residual: a lock-loss right after a token-refresh flush can discard the rotated token (same rare CEO-gated class as the documented post->commit window).
* feat(video): Phase F — panel video-post queue + TikTok creds card + flags
video-post-queue.tsx: <video> MP4 preview with 9:16/1:1 cut switch, per-platform editable captions (280/2200 counters, over-limit disables approve), approve/reject, Request-a-video dialog. tiktok-credentials-card.tsx (4 write-only OAuth2 fields). feature-flags-card inlines TikTokCredentialsForm under video_engine_enabled. Mounted in command-center. tsc/eslint clean, 273 panel tests green. NOTE: needs the GET /video/posts/{id}/media route + mp4_paths on VideoPostResponse (folded into H) for the preview source.
* feat(video): Phase H — media route + e2e smoke + NAS arming + docs
GET /video/posts/{id}/media?cut= (CEO-gated FileResponse of the rendered MP4; closes the panel preview gap) + mp4_paths on VideoPostResponse. e2e smoke tests/e2e_smoke/test_video_pipeline.py (full flow, sidecar+X/TikTok mocked; asserts dispatcher skips, render-loop materialize, propose_video team-gate, approve idempotency). NAS arming: docker-compose.yml/.yaml ROBOCO_VIDEO_ENGINE_ENABLED/ON_RELEASE/ON_SPOTLIGHT default-on (.yaml resynced to .yml); registry stays off. CLAUDE.md video-engine section + CHANGELOG. Fixed 2 pre-existing route-test pollution leaks. Full suite 11763 passed.
* fix(video): auth-carrying preview, media route confinement, VideoPost type drift
Three fixes along the video preview path:
1. panel video preview auth: the <video> element was pointed straight at
GET /video/posts/{id}/media, but a native <video src> GET carries none
of axios's X-Agent-ID/X-Agent-Role headers — so in the default
header-trust deployment the request 401s. Fetch the cut via
videoApi.getMediaBlob (axios, responseType: blob) and drive <video>
off a URL.createObjectURL result instead. The object URL is revoked
on cut-change (the previous cut's URL) and on unmount, so neither
cut switches nor row teardown leak blob URLs.
2. backend media route confinement: GET /video/posts/{id}/media now
resolves mp4_path and refuses it with 404 when it falls outside
settings.video_output_dir. Defense-in-depth against any future
writer of mp4_paths serving files from arbitrary disk locations.
3. panel VideoPost type/comment drift: added mp4_paths to the
VideoPost interface (the committed VideoPostResponse already
carries it), and corrected the stale comment on videoMediaUrl
that claimed no route served the rendered bytes — the route has
existed since the media endpoint landed; the comment now describes
why getMediaBlob exists instead of a direct <video src>.
* Persist rendered videos to data in physical storage.
* ++
* docs(video): 0.18.0 CHANGELOG entry + RAG + map reference for video engine
- Move the video engine bullet from [Unreleased] into [0.18.0] and note
the ROBOCO_VIDEO_OUTPUT_DIR bind-mount persistence.
- Add docs/rag/architecture/video-engine.md (mirrors x-engine.md shape:
enable/disable, three triggers, render loop + sidecar, CEO gate, media
route confinement, credentials).
- Reference the video render loop in docs/map/orchestrator.md's engine list.
* chore(video): re-bump to 0.19.0 + sync registry compose defaults
Version was wrongly bumped to 0.18.0; 0.18.0 is an already-released
section. Restore its 2026-07-04 date and move the video-engine CHANGELOG
bullet into a new [0.19.0] - 2026-07-05 section above it. Bump
pyproject.toml, roboco/__init__.py, roboco/config.py (app_version),
panel/package.json, and the motion/README inputProps example to 0.19.0.
docker-compose.registry.yml: add ROBOCO_VIDEO_ENGINE_ENABLED /
_VIDEO_ON_RELEASE / _VIDEO_ON_SPOTLIGHT defaulted false (NAS arms them
true), and comment out the video-renders bind mount with a short note
so the public registry image ships video off by default. Structural
sync with docker-compose.yml maintained.
* fix(video): rate-limit /render + reflow motion/README
CodeQL flagged js/missing-rate-limiting on the renderer /render route.
The sidecar is container-network-only with one trusted caller (the
orchestrator, which renders cuts serially), so this limiter is a
retry-storm ceiling (30/min, well above legit render rate), not the
primary control. Also reflows motion/README.md hard-wrapped prose that
failed the markdown quality gate.
* fix(build): finish pnpm 11 migration + regen verb tables
The panel Docker image build failed on `pnpm install --frozen-lockfile`:
node:22-alpine's corepack resolved to its bundled pnpm 11, but
panel/package.json pinned packageManager to pnpm@10.25.0, and pnpm 11
refuses to run against that pin. The Dockerfiles were already written for
pnpm 11 (comments, CI=true, strictDepBuilds); the package.json pin was the
stale outlier. Finish the migration instead of working around it:
- panel/package.json: packageManager pnpm@10.25.0 -> pnpm@11.10.0; drop the
`pnpm` field (pnpm 11 ignores it — build approval lives in
panel/pnpm-workspace.yaml's allowBuilds). Lockfile unchanged (pnpm 11
accepts it as-is); frozen-lockfile verified.
- remotion-renderer/package.json: pin packageManager pnpm@11.10.0 for
determinism (was relying on corepack's implicit default); engines.node
>=22.13 (pnpm 11 requirement).
- docker/panel.Dockerfile + docker/remotion.Dockerfile: `corepack prepare
pnpm@11.10.0 --activate` so the build uses the pinned version explicitly
instead of trusting corepack's bundled default (which a future
node:22-alpine could change).
- .github/workflows/panel-ci.yml: Node 20 -> 22 (pnpm 11 requires
Node >=22.13; Node 20 fails the engines check).
Also regenerate agents/prompts/_generated/{developer,head_marketing,verbs}.md
— the video engine added propose_video and extended propose_feature_spotlight
(wants_video, video_script) but the verb tables weren't refreshed, failing
the foundation-check quality gate.
* chore(build): approve esbuild build script in remotion pnpm-workspace.yaml
pnpm 11 generated this file with a placeholder ('set this to true or false')
during install; resolve it to true so local dev of the renderer doesn't
re-prompt. esbuild's postinstall only verifies the prebuilt platform binary
(@esbuild/<platform> is installed as an optional dep), so approving it is
safe and silences the ERR_PNPM_IGNORED_BUILDS warning.
* fix(build): copy pnpm-workspace.yaml into panel + remotion images
pnpm 11 hard-errors with [ERR_PNPM_IGNORED_BUILDS] (exit 1) when a
dependency ships a postinstall script that isn't approved in
allowBuilds. Both Dockerfiles copied only package.json + pnpm-lock.yaml,
so the build-approval map in pnpm-workspace.yaml never made it into the
image — the remotion image build died on esbuild@0.28.1's postinstall.
Copy pnpm-workspace.yaml alongside the manifests in both images. In
panel, this also drops the --config.strictDepBuilds=false workaround:
with sharp and unrs-resolver now approved, their postinstalls run and
install the platform-specific binaries (previously skipped, leaving
sharp without its @img/sharp-* binary at runtime).
Verified locally: remotion + panel `pnpm install --frozen-lockfile`
exit 0 with the workspace file present; both exit 1 without it.
* fix(perf): offload conventions + release-readiness blocking I/O off the event loop
The orchestrator runs uvicorn and the orchestration background loops on a
single shared event loop, so any sync I/O anywhere — even inside a background
loop — blocks API responsiveness for its duration. Two call sites were missing
asyncio.to_thread wrappers:
- ConventionsService.get_map/health/restore called the sync _resolve
(`git rev-parse`), _read_committed_standard (file read + yaml parse), and
_derive (filesystem walk via derive_from_scan) inline. Reachable from
GET /api/projects/{id}/conventions and from the agent spawn-prepare path.
- ReleaseManagerEngine._production_assess called gather_snapshot inline —
multiple `subprocess.run` git calls + a filesystem walk, running inside the
release-manager background loop.
Wrap each blocking call in asyncio.to_thread at the async boundary. No
signature changes; helpers stay sync. Verified: targeted tests pass
(196 passed, 36 DB-skipped), ruff + format clean.
These were the only responsiveness gaps surfaced by the concurrency audit —
the rest of the heavy paths (agent spawn via `docker run -d`, video render
loop, git ops via the 16-worker ThreadPoolExecutor, workspace subprocess
calls) already offload correctly. No API/worker container split needed.
* feat(storage): add MinIO config + dep + compose (no-op, default-off)
Chunk 1 of the MinIO video-storage plan (§1, §2, §6). No behavior change:
minio_endpoint defaults to empty = disabled, the existing FileResponse serve
path is untouched (chunk 4 wires the serve path; chunk 2 adds the client).
- pyproject.toml: add `minio` (minio-py) to dependencies; regenerate uv.lock
(resolves minio v7.2.20 + pycryptodome transitive).
- roboco/config.py: add 5 settings fields after video_output_dir
(minio_endpoint/_access_key/_secret_key/_bucket/_region). Plain str Fields
matching the existing ROBOCO_ENCRYPTION_KEY style; no SecretStr, no
presign_ttl_seconds (YAGNI — we don't presign in phase 1).
- docker-compose.yml: add `minio` service (data network only, named
minio-data volume, host ports 19000/19001 for debugging, mc healthcheck)
and a one-shot `minio-init` service mirroring the ollama-init pattern
(mc alias set + mb -p, idempotent via || true). Add ROBOCO_MINIO_* env to
the orchestrator env block (endpoint, access/secret key, bucket, region).
- docker-compose.registry.yml: intentionally omit the minio/minio-init
services and leave ROBOCO_MINIO_* unset (NAS default-on, registry
default-off — the established pattern); comment added to the orchestrator
env block noting the omission.
* docs(storage): 0.19.0 CHANGELOG + RAG + map reference for MinIO chunk 1
Backfills the release-polish docs for MinIO chunk 1 (§10 of the plan):
- docker-compose.yaml synced to docker-compose.yml (the two NAS compose files
must stay byte-identical; .yml was edited in chunk 1, .yaml was stale).
- CHANGELOG [0.19.0]: Added (MinIO scaffolding) + Fixed (event-loop I/O offload).
- docs/rag/architecture/minio-storage.md: RAG doc mirroring video-engine.md.
- docs/map/deployment-tooling.md: one-line storage reference.
* MinIO chunk 2: minio_client module (singleton + unconfigured guard) (#309)
* feat(storage): minio_client module (singleton + unconfigured guard)
Chunk 2 of the MinIO plan (§3). roboco/services/minio_client.py adds:
- get_client(): singleton minio-py Minio from settings; returns None when
minio_endpoint is empty (the disabled path used by the chunk 3/4 guards).
Parses http://... endpoint into host:port + secure flag.
- put_object(bytes, key): no-ops when unconfigured; otherwise PUTs to
settings.minio_bucket with ContentType video/mp4.
- get_object_stream(key): yields object bytes for StreamingResponse; lets
S3Error propagate so the serve route (chunk 4) can fall back to disk.
Sync calls — every call site wraps in asyncio.to_thread (chunks 3/4). One
unit test covers the unconfigured guard + endpoint scheme parsing (mocks,
no real MinIO). Not yet wired into remotion_client._save or the media route.
* MinIO chunk 3: wire write path (remotion_client._save PUT) (#310)
* feat(storage): wire MinIO write path in remotion_client._save
Chunk 3 of the MinIO plan (§3). After the local mp4 write, _save PUTs the bytes
to MinIO under key = Path(mp4_path).name (already {render_key}-{orientation}.mp4),
guarded by minio_client.get_client() (None when minio_endpoint empty) and
wrapped in asyncio.to_thread. Local disk stays the source of truth for the
poster publish path (x_video_client/tiktok_client read mp4_path from disk);
the PUT is additive. _save still returns the local path str — mp4_paths,
marker, and schema unchanged. Disabled (local-only) when MinIO unconfigured.
One test: asserts put_object is called with the basename key when configured
and the local file is still written; existing test stays green via the
unconfigured-default path. Mocks only.
* fix(storage): make MinIO PUT non-fatal in remotion_client._save
A configured-but-down MinIO made put_object raise inside the worker thread,
failing the render and retry-looping a task whose local file was already
written. Local disk is the source of truth and the serve route falls back to
FileResponse on S3Error, so a failed durable-copy PUT must never fail the
render — log and continue; the next render re-attempts the PUT.
Adds test_save_swallows_minio_put_failure (PUT raises -> _save still returns
the local path and the local file is written). Extends the CHANGELOG write-
path bullet with the non-fatal guarantee.
* MinIO chunk 4: serve path (StreamingResponse + FileResponse fallback) (#311)
* feat(storage): serve MinIO via the media route (StreamingResponse + FileResponse fallback)
Chunk 4 of the MinIO plan (§4 — the crux). GET /api/video/posts/{id}/media
derives key = Path(mp4_path).name and, when minio_endpoint is set, returns a
StreamingResponse over minio_client.get_object_stream(key), keeping
_require_ceo so auth stays end-to-end (no presigned URLs). Falls back to
FileResponse on S3Error (old render not in MinIO) or when MinIO is
unconfigured — the panel's axios-blob flow is unchanged (same URL, headers,
body, just chunked). The confinement check is kept as defense-in-depth (the
key is a basename so traversal is impossible, but the check is cheap and
protects the poster path).
Two integration tests: configured serve path streams from a stubbed
get_object_stream (CEO 200, non-CEO 403); unconfigured fallback serves the
local file via FileResponse. Mocks only — no real MinIO.
* fix(storage): eager stat_object probe so the MinIO serve fallback actually fires
The chunk-4 route wrapped StreamingResponse(get_object_stream(key), ...) in a
try/except, but get_object_stream is a lazy generator — its client.get_object
call runs on the first next(), i.e. AFTER the route returned and Starlette
started streaming. An S3Error (NoSuchKey / MinIO down) there is uncatchable;
the try/except caught nothing and the FileResponse fallback never triggered.
Add minio_client.stat_object(key): an eager existence/readiness probe that
runs INSIDE the route's try/except, so a missing object or down MinIO raises
before the StreamingResponse starts and the fallback serves the local file.
stat-then-get is two round trips; a mid-stream failure after a successful stat
is a rare race the CEO can retry (documented ceiling).
Tests: the configured test now stubs stat_object; a new test asserts the
S3Error fallback serves the local file via FileResponse and that
get_object_stream is never called. RAG doc updated to record the eager-probe
correctness detail + the non-fatal PUT.
* docs(rag): mark MinIO deployment note landed (chunk 5) (#312)
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
* fix(video): offload minio stat_object off the event loop
stat_object was called inline in the async media route, blocking the
shared event loop for one sync urllib3 round-trip per preview request —
contradicting minio_client's own 'every call site wraps in to_thread'
docstring and this PR's perf-fix theme. Wrap in asyncio.to_thread; the
try/except still catches S3Error (to_thread re-raises) so the
FileResponse fallback is unchanged. Also add the trailing newline to
the minio-storage RAG doc.
* Fix red CI
* Make CI green
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
376 lines
18 KiB
YAML
376 lines
18 KiB
YAML
# ============================================================================
|
|
# RoboCo — pre-built (registry) deployment
|
|
# ============================================================================
|
|
# This is the "pull and run" compose for USERS: it runs the images the release
|
|
# workflow publishes (GHCR + Docker Hub) instead of building from source, so a
|
|
# host needs neither the repo's build context nor a build toolchain.
|
|
#
|
|
# 1. Copy `.env.example` to `.env` and fill in the required secrets.
|
|
# 2. docker compose -f docker-compose.registry.yml pull
|
|
# 3. docker compose -f docker-compose.registry.yml up -d
|
|
#
|
|
# Pick the registry + version with two env vars (defaults shown):
|
|
# ROBOCO_REGISTRY=ghcr.io/rennf93 # or docker.io/renzof93
|
|
# ROBOCO_VERSION=latest # or a pinned release, e.g. 0.5.0
|
|
#
|
|
# The orchestrator spawns agent containers itself; ROBOCO_AGENT_IMAGE_REGISTRY
|
|
# + ROBOCO_AGENT_IMAGE_TAG below tell it to spawn the SAME pre-built agent
|
|
# images (it pulls any it doesn't already have). The one-shot agent-*-image
|
|
# services exist only so `docker compose pull` fetches every agent image up
|
|
# front; they pull then exit.
|
|
#
|
|
# NOTE: this file is the registry counterpart of docker-compose.yml — when you
|
|
# add or change a service there, mirror it here. The infra services
|
|
# (postgres/redis/ollama/nginx) are byte-identical to the build compose.
|
|
# ============================================================================
|
|
services:
|
|
postgres:
|
|
image: pgvector/pgvector:pg16
|
|
container_name: roboco-postgres
|
|
restart: unless-stopped
|
|
# data-only network: agent containers (on roboco_default) cannot reach
|
|
# the production DB; only the multi-homed orchestrator can. Host port
|
|
# publishing (15432) is unaffected — roboco_data is a normal bridge.
|
|
networks:
|
|
- data
|
|
environment:
|
|
POSTGRES_USER: roboco
|
|
POSTGRES_PASSWORD: roboco
|
|
POSTGRES_DB: roboco
|
|
ports:
|
|
- "15432:5432"
|
|
volumes:
|
|
- ${ROBOCO_DATA_DIR:-./data}/postgres:/var/lib/postgresql/data
|
|
healthcheck:
|
|
test: ["CMD-SHELL", "pg_isready -U roboco -d roboco"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 5
|
|
|
|
redis:
|
|
image: redis:8-alpine
|
|
container_name: roboco-redis
|
|
restart: unless-stopped
|
|
# data-only network — see postgres. Redis has no auth, so network
|
|
# membership is its ONLY containment against agent containers.
|
|
networks:
|
|
- data
|
|
command: redis-server --appendonly yes
|
|
ports:
|
|
- "16379:6379"
|
|
volumes:
|
|
- ${ROBOCO_DATA_DIR:-./data}/redis:/data
|
|
healthcheck:
|
|
test: ["CMD", "redis-cli", "ping"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 5
|
|
|
|
ollama:
|
|
image: ollama/ollama:latest
|
|
container_name: roboco-ollama
|
|
restart: unless-stopped
|
|
environment:
|
|
OLLAMA_API_KEY: ${OLLAMA_API_KEY:-}
|
|
ports:
|
|
- "11435:11434"
|
|
volumes:
|
|
- ${ROBOCO_DATA_DIR:-./data}/ollama:/root/.ollama
|
|
healthcheck:
|
|
test: ["CMD", "ollama", "list"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 5
|
|
start_period: 10s
|
|
|
|
ollama-init:
|
|
image: curlimages/curl:latest
|
|
container_name: roboco-ollama-init
|
|
depends_on:
|
|
ollama:
|
|
condition: service_healthy
|
|
restart: "no"
|
|
entrypoint: ["/bin/sh", "-c"]
|
|
command:
|
|
- |
|
|
# NO `set -e`: pulls are best-effort. Cached models persist in the ollama
|
|
# volume and MUST survive a slow/unreachable ollama registry — a manifest
|
|
# re-check failure there must never down a fully-cached deployment (it
|
|
# used to: a degraded registry made `ollama pull` fail under set -e, which
|
|
# blocked the orchestrator's service_completed_successfully gate). Success
|
|
# is gated on the models being PRESENT (verify step), not on the pull.
|
|
# Note: $$ escapes $ for docker-compose variable substitution.
|
|
echo "=== Pulling embedding model (qwen3-embedding:0.6b) — best-effort ==="
|
|
curl -sN http://ollama:11434/api/pull -d '{"name":"qwen3-embedding:0.6b"}' | while read -r line; do
|
|
status=$$(echo "$$line" | grep -o '"status":"[^"]*"' | cut -d'"' -f4)
|
|
[ -n "$$status" ] && echo " $$status"
|
|
done || echo " (pull failed — relying on the cached model)"
|
|
echo "=== Pulling LLM model (glm-5.2:cloud) — best-effort ==="
|
|
curl -sN http://ollama:11434/api/pull -d '{"name":"glm-5.2:cloud"}' | while read -r line; do
|
|
status=$$(echo "$$line" | grep -o '"status":"[^"]*"' | cut -d'"' -f4)
|
|
[ -n "$$status" ] && echo " $$status"
|
|
done || echo " (pull failed — relying on the cached model)"
|
|
echo "=== Verifying models are present (the real success gate) ==="
|
|
curl -sf http://ollama:11434/api/tags | grep -q "qwen3-embedding" || { echo "FATAL: qwen3-embedding missing and could not be pulled"; exit 1; }
|
|
curl -sf http://ollama:11434/api/tags | grep -q "glm-5.2" || { echo "FATAL: glm-5.2 missing and could not be pulled"; exit 1; }
|
|
echo "=== All models ready! ==="
|
|
|
|
# --------------------------------------------------------------------------
|
|
# Remotion Renderer - sidecar for the video-generation engine (bundles +
|
|
# renders motion/ compositions to MP4). Harmless when idle — it only
|
|
# renders when POSTed; the video_engine_* feature flags (off by default in
|
|
# this compose) gate whether the orchestrator ever sends it work. No DB
|
|
# access, no git, no credentials.
|
|
# --------------------------------------------------------------------------
|
|
remotion-renderer:
|
|
image: ${ROBOCO_REGISTRY:-ghcr.io/rennf93}/roboco-remotion:${ROBOCO_VERSION:-latest}
|
|
container_name: roboco-remotion
|
|
restart: unless-stopped
|
|
networks:
|
|
- default
|
|
# Chrome headless rendering can crash under Docker's default 64MB
|
|
# /dev/shm ("Chrome crashed"); give it real shared memory.
|
|
shm_size: "1gb"
|
|
healthcheck:
|
|
test: ["CMD", "node", "-e", "fetch('http://127.0.0.1:3001/health').then((r) => process.exit(r.ok ? 0 : 1)).catch(() => process.exit(1))"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 5
|
|
start_period: 10s
|
|
|
|
# --------------------------------------------------------------------------
|
|
# Agent image pre-pull (one-shot). Each pulls its published image then exits,
|
|
# so `docker compose pull` fetches every agent image up front. The
|
|
# orchestrator spawns these same images at runtime.
|
|
# --------------------------------------------------------------------------
|
|
agent-base-image:
|
|
image: ${ROBOCO_REGISTRY:-ghcr.io/rennf93}/roboco-agent-base:${ROBOCO_VERSION:-latest}
|
|
entrypoint: ["/bin/sh", "-c", "echo 'agent-base image present'"]
|
|
restart: "no"
|
|
|
|
agent-pm-image:
|
|
image: ${ROBOCO_REGISTRY:-ghcr.io/rennf93}/roboco-agent-pm:${ROBOCO_VERSION:-latest}
|
|
entrypoint: ["/bin/sh", "-c", "echo 'agent-pm image present'"]
|
|
restart: "no"
|
|
|
|
agent-dev-be-image:
|
|
image: ${ROBOCO_REGISTRY:-ghcr.io/rennf93}/roboco-agent-dev-be:${ROBOCO_VERSION:-latest}
|
|
entrypoint: ["/bin/sh", "-c", "echo 'agent-dev-be image present'"]
|
|
restart: "no"
|
|
|
|
agent-dev-fe-image:
|
|
image: ${ROBOCO_REGISTRY:-ghcr.io/rennf93}/roboco-agent-dev-fe:${ROBOCO_VERSION:-latest}
|
|
entrypoint: ["/bin/sh", "-c", "echo 'agent-dev-fe image present'"]
|
|
restart: "no"
|
|
|
|
agent-qa-be-image:
|
|
image: ${ROBOCO_REGISTRY:-ghcr.io/rennf93}/roboco-agent-qa-be:${ROBOCO_VERSION:-latest}
|
|
entrypoint: ["/bin/sh", "-c", "echo 'agent-qa-be image present'"]
|
|
restart: "no"
|
|
|
|
agent-qa-fe-image:
|
|
image: ${ROBOCO_REGISTRY:-ghcr.io/rennf93}/roboco-agent-qa-fe:${ROBOCO_VERSION:-latest}
|
|
entrypoint: ["/bin/sh", "-c", "echo 'agent-qa-fe image present'"]
|
|
restart: "no"
|
|
|
|
agent-ux-image:
|
|
image: ${ROBOCO_REGISTRY:-ghcr.io/rennf93}/roboco-agent-ux:${ROBOCO_VERSION:-latest}
|
|
entrypoint: ["/bin/sh", "-c", "echo 'agent-ux image present'"]
|
|
restart: "no"
|
|
|
|
agent-doc-image:
|
|
image: ${ROBOCO_REGISTRY:-ghcr.io/rennf93}/roboco-agent-doc:${ROBOCO_VERSION:-latest}
|
|
entrypoint: ["/bin/sh", "-c", "echo 'agent-doc image present'"]
|
|
restart: "no"
|
|
|
|
agent-prompter-image:
|
|
image: ${ROBOCO_REGISTRY:-ghcr.io/rennf93}/roboco-agent-prompter:${ROBOCO_VERSION:-latest}
|
|
entrypoint: ["/bin/sh", "-c", "echo 'agent-prompter image present'"]
|
|
restart: "no"
|
|
|
|
agent-secretary-image:
|
|
image: ${ROBOCO_REGISTRY:-ghcr.io/rennf93}/roboco-agent-secretary:${ROBOCO_VERSION:-latest}
|
|
entrypoint: ["/bin/sh", "-c", "echo 'agent-secretary image present'"]
|
|
restart: "no"
|
|
|
|
agent-pr-reviewer-image:
|
|
image: ${ROBOCO_REGISTRY:-ghcr.io/rennf93}/roboco-agent-pr-reviewer:${ROBOCO_VERSION:-latest}
|
|
entrypoint: ["/bin/sh", "-c", "echo 'agent-pr-reviewer image present'"]
|
|
restart: "no"
|
|
|
|
agent-grok-image:
|
|
image: ${ROBOCO_REGISTRY:-ghcr.io/rennf93}/roboco-agent-grok:${ROBOCO_VERSION:-latest}
|
|
entrypoint: ["/bin/sh", "-c", "echo 'agent-grok image present'"]
|
|
restart: "no"
|
|
|
|
agent-grok-prompter-image:
|
|
image: ${ROBOCO_REGISTRY:-ghcr.io/rennf93}/roboco-agent-grok-prompter:${ROBOCO_VERSION:-latest}
|
|
entrypoint: ["/bin/sh", "-c", "echo 'agent-grok-prompter image present'"]
|
|
restart: "no"
|
|
|
|
agent-grok-secretary-image:
|
|
image: ${ROBOCO_REGISTRY:-ghcr.io/rennf93}/roboco-agent-grok-secretary:${ROBOCO_VERSION:-latest}
|
|
entrypoint: ["/bin/sh", "-c", "echo 'agent-grok-secretary image present'"]
|
|
restart: "no"
|
|
|
|
# --------------------------------------------------------------------------
|
|
# Orchestrator — API server + agent spawner
|
|
# --------------------------------------------------------------------------
|
|
orchestrator:
|
|
image: ${ROBOCO_REGISTRY:-ghcr.io/rennf93}/roboco-orchestrator:${ROBOCO_VERSION:-latest}
|
|
container_name: roboco-orchestrator
|
|
restart: unless-stopped
|
|
# Multi-homed: the agent mesh (default) for spawned agents / panel /
|
|
# ollama, plus the data network for postgres/redis.
|
|
networks:
|
|
- default
|
|
- data
|
|
ports:
|
|
- "8000:8000"
|
|
environment:
|
|
# DB network isolation is LIVE in this file (postgres/redis on the
|
|
# data-only network): suppress the legacy prod-creds gate-env
|
|
# injection — agents can't reach roboco-postgres anyway. This flag
|
|
# must always travel with the networks: topology.
|
|
ROBOCO_DB_NETWORK_ISOLATED: ${ROBOCO_DB_NETWORK_ISOLATED:-true}
|
|
ROBOCO_DATABASE_HOST: roboco-postgres
|
|
ROBOCO_DATABASE_PORT: 5432
|
|
ROBOCO_DATABASE_USER: roboco
|
|
ROBOCO_DATABASE_PASSWORD: roboco
|
|
ROBOCO_DATABASE_NAME: roboco
|
|
ROBOCO_REDIS_HOST: roboco-redis
|
|
ROBOCO_REDIS_PORT: 6379
|
|
ROBOCO_HOST: 0.0.0.0
|
|
ROBOCO_PORT: 8000
|
|
ROBOCO_ENCRYPTION_KEY: ${ROBOCO_ENCRYPTION_KEY:?ROBOCO_ENCRYPTION_KEY is required}
|
|
ROBOCO_AGENT_AUTH_SECRET: ${ROBOCO_AGENT_AUTH_SECRET:?ROBOCO_AGENT_AUTH_SECRET is required}
|
|
ROBOCO_AGENT_AUTH_REQUIRED: ${ROBOCO_AGENT_AUTH_REQUIRED:-false}
|
|
ROBOCO_LOCAL_LLM_BASE_URL: http://roboco-ollama:11434/v1
|
|
ROBOCO_LOCAL_LLM_MODEL: glm-5.2:cloud
|
|
ROBOCO_DEFAULT_EMBEDDING_MODEL: qwen3-embedding:0.6b
|
|
ROBOCO_OLLAMA_BASE_URL: http://roboco-ollama:11434
|
|
# Remotion renderer sidecar (use container name). The video engine
|
|
# itself is default-off (video_engine_enabled); this just points the
|
|
# client at the sidecar for when it's armed.
|
|
ROBOCO_REMOTION_BASE_URL: http://roboco-remotion:3001
|
|
# Spawn the PRE-BUILT agent images from the same registry instead of
|
|
# building them from source (the orchestrator pulls any it lacks).
|
|
ROBOCO_AGENT_IMAGE_REGISTRY: ${ROBOCO_REGISTRY:-ghcr.io/rennf93}
|
|
ROBOCO_AGENT_IMAGE_TAG: ${ROBOCO_VERSION:-latest}
|
|
# Host paths for spawning agent containers (Docker-in-Docker). MUST be
|
|
# absolute paths on the host. Default to this compose project's ./data.
|
|
ROBOCO_HOST_PROJECT_DIR: ${ROBOCO_HOST_PROJECT_DIR:-/opt/roboco}
|
|
ROBOCO_HOST_CLAUDE_DIR: ${ROBOCO_HOST_CLAUDE_DIR:-${HOME}/.claude}
|
|
# SuperGrok auth (host ~/.grok) for Grok-CLI agents — the orchestrator
|
|
# mounts <dir>/auth.json into each Grok agent. Run `grok login` on the host.
|
|
ROBOCO_HOST_GROK_DIR: ${ROBOCO_HOST_GROK_DIR:-${HOME}/.grok}
|
|
ROBOCO_HOST_DATA_DIR: ${ROBOCO_HOST_DATA_DIR:-/opt/roboco/data}
|
|
# Reachable base URL for commit-trailer links — set to your host's LAN
|
|
# address or domain so the links in commit bodies resolve.
|
|
ROBOCO_PUBLIC_BASE_URL: ${ROBOCO_PUBLIC_BASE_URL:-http://localhost:8000}
|
|
ROBOCO_ENVIRONMENT: production
|
|
ROBOCO_CLAIM_STALE_SECONDS: "1800"
|
|
ROBOCO_STALE_CLAIM_REAP_SECONDS: "1800"
|
|
# Every optional feature ships OFF in this user-facing compose (the build
|
|
# compose arms them for the personal deploy). External-PR review included —
|
|
# arm it (and any other subsystem) via .env or Settings → Feature Flags.
|
|
ROBOCO_EXTERNAL_PR_ENABLED: ${ROBOCO_EXTERNAL_PR_ENABLED:-false}
|
|
ROBOCO_EXTERNAL_PR_REQUIRE_HUMAN_CONFIRM: ${ROBOCO_EXTERNAL_PR_REQUIRE_HUMAN_CONFIRM:-true}
|
|
# ROBOCO_EXTERNAL_PR_POLL_INTERVAL_SECONDS: "300"
|
|
# ROBOCO_EXTERNAL_PR_AUTHOR_ALLOWLIST: '["corey"]' # empty = every external PR
|
|
# Production self-healing ("engine 4"). RoboCo watches its OWN repo CI and,
|
|
# when red, notifies the CEO and (with originate on) opens a PENDING fix
|
|
# task that STOPS for the CEO's Approve-&-Start — it never self-deploys.
|
|
# Both toggles default OFF (arm from Settings -> Feature Flags). Set
|
|
# PROJECT_SLUG to the registered project that IS RoboCo; CI_WORKFLOW scopes
|
|
# the signal to the real CI workflow (RoboCo has several workflows, so the
|
|
# unscoped "latest run" would be unreliable).
|
|
ROBOCO_SELF_HEAL_ENABLED: ${ROBOCO_SELF_HEAL_ENABLED:-false}
|
|
ROBOCO_SELF_HEAL_ORIGINATE_ENABLED: ${ROBOCO_SELF_HEAL_ORIGINATE_ENABLED:-false}
|
|
ROBOCO_SELF_HEAL_PROJECT_SLUG: ${ROBOCO_SELF_HEAL_PROJECT_SLUG:-roboco-api}
|
|
ROBOCO_SELF_HEAL_CI_WORKFLOW: ${ROBOCO_SELF_HEAL_CI_WORKFLOW:-ci.yml}
|
|
# Video engine (Remotion): NAS-default-on, OFF in public registry —
|
|
# heavier optional feature. Arm via .env + uncomment the video-renders
|
|
# bind mount above so renders persist.
|
|
ROBOCO_VIDEO_ENGINE_ENABLED: ${ROBOCO_VIDEO_ENGINE_ENABLED:-false}
|
|
ROBOCO_VIDEO_ON_RELEASE: ${ROBOCO_VIDEO_ON_RELEASE:-false}
|
|
ROBOCO_VIDEO_ON_SPOTLIGHT: ${ROBOCO_VIDEO_ON_SPOTLIGHT:-false}
|
|
# MinIO object storage for rendered videos is intentionally omitted from
|
|
# this registry compose (NAS default-on, registry default-off — the
|
|
# established pattern). The minio/minio-init services are absent here and
|
|
# ROBOCO_MINIO_* is left unset, so minio_endpoint defaults to empty and the
|
|
# media route falls back to FileResponse. Arm via a custom override file.
|
|
volumes:
|
|
- /var/run/docker.sock:/var/run/docker.sock
|
|
- ${CLAUDE_AUTH_DIR:-${HOME}/.claude}:/root/.claude
|
|
# SuperGrok auth — mount host ~/.grok at the SAME host path the orchestrator
|
|
# hands each Grok agent's `-v`, so its auth.json exists() check passes here
|
|
# AND the agent bind resolves on the host. Read-WRITE: the orchestrator
|
|
# auto-refreshes the ~6h token in place (grok_auth.refresh_if_stale) so
|
|
# agents never mount a dead credential; the agent's own mount stays RO.
|
|
- ${ROBOCO_HOST_GROK_DIR:-${HOME}/.grok}:${ROBOCO_HOST_GROK_DIR:-${HOME}/.grok}
|
|
- ${ROBOCO_DATA_DIR:-./data}/mcp-configs:/app/mcp-configs
|
|
- ${ROBOCO_DATA_DIR:-./data}/prompts-generated:/app/prompts-generated
|
|
- ${ROBOCO_DATA_DIR:-./data}/agent-settings:/app/agent-settings
|
|
- ${ROBOCO_DATA_DIR:-./data}/workspaces:/data/workspaces
|
|
# Per-agent GROK usage capture (usage.json -> finalizer).
|
|
- ${ROBOCO_DATA_DIR:-./data}/grok-usage:/data/grok-usage
|
|
- ${ROBOCO_DATA_DIR:-./data}/logs:/data/logs
|
|
# video engine: NAS-only bind mount, off by default in public registry.
|
|
# Uncomment + set ROBOCO_VIDEO_ENGINE_ENABLED=true in .env to arm.
|
|
# - ${ROBOCO_DATA_DIR:-./data}/video-renders:/data/video-renders
|
|
- ${ROBOCO_DATA_DIR:-./data}/briefings:/app/briefings
|
|
- ${ROBOCO_DATA_DIR:-./data}/manifests:/app/manifests
|
|
depends_on:
|
|
postgres:
|
|
condition: service_healthy
|
|
redis:
|
|
condition: service_healthy
|
|
ollama:
|
|
condition: service_healthy
|
|
ollama-init:
|
|
condition: service_completed_successfully
|
|
agent-base-image:
|
|
condition: service_completed_successfully
|
|
|
|
# --------------------------------------------------------------------------
|
|
# Next.js control panel (fronted by nginx; not exposed directly)
|
|
# --------------------------------------------------------------------------
|
|
panel:
|
|
image: ${ROBOCO_REGISTRY:-ghcr.io/rennf93}/roboco-panel:${ROBOCO_VERSION:-latest}
|
|
container_name: roboco-panel
|
|
restart: unless-stopped
|
|
expose:
|
|
- "3000"
|
|
depends_on:
|
|
- orchestrator
|
|
|
|
# --------------------------------------------------------------------------
|
|
# Nginx — single entry point on port 3000
|
|
# --------------------------------------------------------------------------
|
|
nginx:
|
|
image: nginx:alpine
|
|
container_name: roboco-nginx
|
|
restart: unless-stopped
|
|
ports:
|
|
- "3000:80"
|
|
environment:
|
|
ROBOCO_PANEL_AGENT_TOKEN: ${ROBOCO_PANEL_AGENT_TOKEN:-}
|
|
NGINX_ENVSUBST_FILTER: "^ROBOCO_"
|
|
volumes:
|
|
- ./docker/nginx.conf:/etc/nginx/templates/default.conf.template:ro
|
|
depends_on:
|
|
- panel
|
|
- orchestrator
|
|
|
|
networks:
|
|
default:
|
|
name: roboco_default
|
|
# DB isolation: postgres/redis live ONLY here; the orchestrator is the
|
|
# only service homed on both. Spawned agent containers + sandbox sidecars
|
|
# join roboco_default (AGENT_NETWORK) and cannot resolve or reach the
|
|
# production DB/Redis. Normal bridge (not internal) so host-published
|
|
# ports keep working.
|
|
data:
|
|
name: roboco_data
|