Files
roboco/docs/deploy/deployment.md
T
a8cb2470ba v0.15.0: Metrics granularity — per-member / per-task / org + CEO scorecards (#289)
* feat(metrics): capture per-session turns + tool_calls (phase 1)

Persist LLM iterations (turns) and tool invocations per agent spawn session,
the raw signal the granular per-member performance metrics build on (real
effort/iterations vs wall-clock).

- sum_transcript_usage returns a 5-tuple adding turns = unique assistant
  message-id count; _usage_from_transcript + _resolve_active_tokens updated to
  the 5-tuple (active-tokens keeps its 4-tuple contract by slicing).
- SDK: _SessionState.turns, set by /usage/sync; /usage/status (TokenUsageStatus)
  now carries turns + tool_calls (= total_calls).
- orchestrator: new _resolve_final_turns_tools (SDK primary, transcript fallback
  for turns only; Grok -> 0/0) wired into _finalize_spawn_session, which writes
  turns + tool_calls to agent_spawn_sessions.
- migration 055 adds turns + tool_calls (BigInteger DEFAULT 0 -> historical/Grok
  rows read 0, surfaced as n/a). Verified real alembic upgrade/downgrade.

Part of metrics-granularity (v0.15.0); recon-adjusted plan on disk.

* feat(metrics): pure compute_stage_effort helper (phase 2, part 1)

Foundation-layer overlap math (no DB): split each task status window into
active (merged wall-clock overlap of spawn stints — concurrent stints counted
once, so active <= window) vs wait (queue/review idle). Distinct from summed
effort. The per-task metrics service will feed it audit-log windows + spawn
stints. 9 unit tests (disjoint/nested/partial/merged/clamped/zero/multi-window).

* feat(metrics): per-task live metrics + GET /metrics/task/{id} (phase 2)

TaskMetrics dataclass + MetricsService.get_task_metrics: summed spawn effort
(vs wall-clock), turns/tool_calls/tokens/cost, per-stage active-vs-wait
(compute_stage_effort over audit windows x spawn stints), and who-caused-rework
(revision_count + named qa/pr fail events). Open stints and the open final
stage window close at completed_at for a terminal task (else now), so stages
don't grow past completion. Exposed at GET /dashboard/metrics/task/{task_id}
(404 if absent). Real-PG tests (compose/none/in-flight) + route tests (200/404).

* feat(metrics): CEO-as-member scorecard + ceo_reject audit regression (phase 3)

The human CEO is a measured member, read purely from audit_log (agent_role='ceo'
serializes from the CEO StrEnum): approval dwell (awaiting_ceo_approval -> a CEO
decision, incl. the coordination-root reject that lands in pending), unblock
dwell (blocked -> a CEO revive), and god-mode action count (every CEO-attributed
transition). CeoScorecard + MetricsService.get_ceo_scorecard (p50/p90 via
PERCENTILE_CONT, expanding IN for the decision sets) + GET
/dashboard/metrics/member/ceo (declared before any future member/{id} route).

The ceo_reject coordination-root audit gap the plan meant to close was already
closed by the gap-sweep (routes through admin_set_status -> agent_role='ceo'
audit); locked with a regression assertion in the existing coordination-reject
test. Real-PG tests: approval/unblock/godmode, non-ceo exclusion, empty->zeros.

* feat(metrics): audit instrumentation for escalations/blocked-others/idle (phase 4a)

The three extra per-member metrics that had no data source get durable,
in-session audit events (additive; never gate the underlying action):
- apply_escalation -> task.escalated (details.escalator_slug) on both the
  normal block path and the pool-divert path -> escalations count.
- _unblock_dependents -> task.unblocked_dependents (details.count) on the
  completed BLOCKER task, captured before the dependency edges are pruned ->
  blocked-others count (sweeper attributes to the blocker's owner).
- mark_agent_idle -> agent.idle (details.agent_slug) -> idle/utilization (the
  sweeper pairs an idle mark to the member's next spawn for idle duration).
(QA pass-rate needs no new event — reuses task.awaiting_documentation[qa] +
task.qa_fail.) Real-PG tests for each; 111 transition tests still green.

* feat(metrics): member_performance_daily rollup table + migration 056 (phase 4b)

The per-member scorecard rollup: one row per (date, member_kind, agent_slug),
CEO as a first-class member_kind='ceo' row (agent_slug='' NOT NULL so the
NULL-distinct UNIQUE keeps it unique). Full column set + the four CEO-approved
extras (qa_reviews_total/passed, escalations, blocked_others, idle_seconds) plus
blocked_seconds. Overwrite-upsert on (date, member_kind, agent_slug) for an
idempotent sweep. Migration 056 verified real up/down (24 cols, 4 indexes).

* feat(metrics): _sweep_member_performance rollup sweeper (phase 4c)

The daily per-member rollup sweep (mirrors _sweep_daily_rollup): a trailing
7-day, idempotent overwrite-upsert wired into _run_sweep. One focused query per
metric merges into a (date, agent_slug) accumulator — spawn effort/turns/tokens/
cost, completed/first-pass/revisions-received, revisions-caused (qa/pr fails),
QA pass-rate (passed + total), escalations (by escalator_slug), blocked-others
(unblocked_dependents by blocker owner), idle_seconds (idle mark -> next spawn),
blocked_seconds (blocked dwell) — plus one CEO row/day (approval/unblock dwell +
god-mode). Real-PG test asserts every facet + idempotency (a 2nd sweep
overwrites, never doubles); spawn-day != completion-day split is by-design.

* feat(metrics): member/org rollup scorecards + endpoints + live overlay (phase 5)

MemberScorecard + OrgScorecard with derived rates (FPY, effort-throughput,
turns/tool-calls per task, QA pass-rate, utilization) — all division-guarded to
None. get_member_scorecard reads member_performance_daily by slug and overlays
the member's live in-flight (non-terminal) tasks' effort via get_task_metrics
(disjoint by status: completion counts stay rollup-only, overlay only enriches
effort/turns/cost; includes_live_inflight flags it). get_org_scorecard
aggregates the cell (?team=) or whole org. Routes: GET /metrics/member/{agent_id}
(404 if absent, after the ceo literal route) + GET /metrics/org?team=. Real-PG
tests (derived rates, overlay no double-count, guards, org) + route tests.

* feat(metrics): granular CEO completion notification (phase 6)

There was no CEO completion notification at all (EventType.TASK_COMPLETED was
defined but never emitted). Add notify_ceo_of_completion in
NotificationDeliveryService — a granular body (real effort vs wall-clock +
stints/turns/tool-calls/revisions[QA/PR]/cost from get_task_metrics; degrades to
wall-clock-only, turns 'n/a', when there are no spawn sessions). Reuses the
existing ALERT type (no enum migration; the notificationtype PG enum is fixed at
001). ceo_approve now emits TASK_COMPLETED + fires the notification (best-effort
via _notify_completion — never blocks completion); complete() emits
TASK_COMPLETED too (closes the dead-code gap; the WS bridge can forward it).
Pure formatter tests + real-PG notification test.

* [metrics-granularity] Phase 7: panel Scorecards tab + dashboard overview

Add the CEO-facing metrics surfaces for the granularity feature:

- New "Scorecards" tab on the Metrics page: org rollup headline, the
  CEO-as-member card (approval/unblock dwell + god-mode count), and a
  per-member table (completed, first-pass yield, active effort, turns/task,
  QA pass-rate, escalations, blocked-others, utilization). Each member row
  self-fetches its rollup scorecard; live in-flight rows carry a "live" badge.
- New dashboard overview card (ScorecardOverviewPanel): org-wide 30-day
  headline (completed, FPY, throughput/hr, active effort, cost) deep-linking
  into the Scorecards tab.
- Plumbing: TaskMetrics/MemberScorecard/OrgScorecard/CeoScorecard types,
  observability API client methods + empty fallbacks, and the four
  useCeoScorecard/useMemberScorecard/useOrgScorecard/useTaskMetrics hooks.

Panel gate green: tsc, eslint, prettier, vitest (175 tests, +6 new).

* [metrics-granularity] test: make completion-notification robust to shared-DB CEO

test_notify_ceo_of_completion_creates_alert errored in the full suite (passed
in isolation): the session-scoped test DB is shared across the run, and the
sibling real-DB board-gate test commits a role=CEO agent (slug="ceo") without
cleanup — so my env fixture's hardcoded slug="ceo" insert hit a unique-constraint
violation, and a second role=CEO row would also make _get_ceo_agent()'s
scalar_one_or_none() raise. Reuse an existing CEO when present (the singleton the
production system actually has), else create one with a unique slug. Order-
independent. Also reflow test_metrics_instrumentation.py to ruff format.

* chore(release): 0.15.0

Metrics granularity: per-member/per-task/org + CEO-as-member scorecards,
turn/tool-call capture (migration 055), member_performance_daily rollup
(migration 056) with QA pass-rate / escalations / blocked-others / utilization,
per-task active-vs-wait metrics, granular completion notification, panel
Scorecards tab + dashboard Performance card, and the ceo_reject audit fix.

Version bump across the canonical set + CHANGELOG.

* [metrics-granularity] fix pre-tag audit findings (overlay double-count + panel error states)

Adversarial review before the v0.15.0 tag surfaced two real logical gaps:

- MAJOR (backend): the live in-flight overlay re-summed ALL sessions of every
  non-terminal task via get_task_metrics, but _msweep_spawn already rolls up
  every CLOSED session regardless of task status — so a closed session on a
  still-open task was counted twice (rollup + overlay), permanently inflating a
  member's effort/turns/tokens/cost on the common reap/respawn path. The overlay
  now sums only OPEN sessions (ended_at IS NULL), which the closed-only rollup
  can never contain — disjoint by construction. A just-closed session lands in
  the rollup on the next ~60s sweep (no gap of note). Aggregated in SQL to mirror
  _msweep_spawn. Regression test reproduces the double-count (turns 10→5).

- MAJOR (panel): the four new scorecard surfaces used `isLoading || !data` with
  no isError branch, so a failed query span forever on a skeleton. They now
  surface a load error. Tests added.

Also: OrgSummary active-effort formatting no longer round-trips hours→seconds→
hours; dashboard grid uses xl:grid-cols-4 (was 2xl) so 4 panels show at 1280px;
corrected the inaccurate "NULL distinct" CEO-row uniqueness comment (agent_slug
is NOT NULL; the '' tuple is simply distinct from agent rows).

make quality GREEN (cov 95.31%); panel GREEN (vitest 178).

* [metrics-granularity] fix: decode bytes stream message-id before XCLAIM

StreamEventBus._recover_stream passed the pending message id to XCLAIM via
str() on the raw bytes the client returns (redis client has no
decode_responses), producing "b'1782066556728-0'". Redis rejects that with
"Unrecognized XCLAIM option", so pending-message recovery threw on every
reclaim tick and unacked messages from crashed/slow consumers were never
reclaimed (leaking in the PEL on every stream, spamming the error log). Decode
via the existing _to_str helper — the fix the sibling claim path already uses.

Pre-existing in v0.14.0 (unrelated to metrics granularity); folded into this
release per CEO. TDD regression test + CHANGELOG entry. make quality GREEN.

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-07-01 05:18:45 +02:00

12 KiB

Production deploy

This is the operator reference for running RoboCo on a NAS or server. If you just want it up on your laptop, the install quickstart is faster — this page assumes you've done that once and now want the durable, server-side setup: the compose files, the host mounts agents need, where data lives, how to back it up, and how to harden it.

!!! warning "Trusted network only" RoboCo is built for a private LAN or homelab. Do not expose it directly to the public internet. nginx is the single entry point, but the orchestrator's WebSocket streams and (in header-trust mode) its API assume a trusted network. Put it behind your own VPN if you need remote access.

The three compose files

There are three tracked compose files, and they are not interchangeable:

File What it does Needs a build toolchain?
docker-compose.yml Builds every image from the Dockerfiles in docker/. Yes
docker-compose.yaml Byte-identical to docker-compose.yml. Yes
docker-compose.registry.yml Pulls and runs the pre-built published images. No

docker-compose.yml and docker-compose.yaml are the same file under two names — Docker Compose picks up either, and the NAS deployment runs the .yaml. If you fork RoboCo and change a service, keep all three in sync.

Which one to run

For a server you don't intend to hack on, run the registry file — it pulls finished images and needs no source tree or compiler on the host:

docker compose -f docker-compose.registry.yml pull
docker compose -f docker-compose.registry.yml up -d

Two variables choose what you pull (defaults shown):

ROBOCO_REGISTRY=ghcr.io/rennf93   # or docker.io/renzof93
ROBOCO_VERSION=latest             # or a pinned release, e.g. 0.15.0

The orchestrator then spawns the matching pre-built agent images on demand (it reads ROBOCO_AGENT_IMAGE_REGISTRY / ROBOCO_AGENT_IMAGE_TAG, which the registry compose wires to the same registry and version). Pin ROBOCO_VERSION to a release tag in production so an upstream latest push can't silently change your fleet.

Build from source only when you're modifying RoboCo:

docker compose up -d   # builds on first run

!!! note "Agent images are build/pull-only services" The agent-*-image services in every compose file are one-shot stubs — they exist so docker compose build/pull materializes each per-role agent image up front. They never run as long-lived containers. The orchestrator spawns the actual agent containers itself, on demand, over the mounted Docker socket, and tears them down when their work is done.

The single origin

nginx (docker/nginx.conf, rendered from an envsubst template) is the only externally-exposed service. It listens on localhost:3000 and routes by path:

flowchart LR
    B[Browser :3000] --> N[nginx]
    N -->|/| P[panel:3000]
    N -->|/api/, /ws/, /health, /ready| O[orchestrator:8000]
Path Upstream
/api/, /ws/, /health, /ready roboco-orchestrator:8000
everything else roboco-panel:3000

The browser only ever sees one origin (:3000), so there's no CORS to configure — the panel uses relative /api and /ws URLs and lets nginx dispatch. The panel container is never published directly; you reach it only through nginx. /ws/ also gets a long (86400s) read timeout so live sockets stay open.

The backing services do publish host ports for direct inspection — Postgres on 15432, Redis on 16379, Ollama on 11435, and the orchestrator on 8000. You don't route browser traffic at these; they're there for psql, redis-cli, and the like.

Required host-path mounts

The orchestrator is Docker-in-Docker: it mounts /var/run/docker.sock and spawns agent containers itself. Because those agent bind-mounts resolve on the host daemon (not inside the orchestrator container), several paths must be given as absolute host paths — the orchestrator passes them straight through to docker run -v for each agent.

Variable What it points at Compose default
ROBOCO_HOST_PROJECT_DIR The RoboCo project directory on the host. /volume1/roboco
ROBOCO_HOST_CLAUDE_DIR / CLAUDE_AUTH_DIR The host ~/.claude Claude Code auth dir, mounted into the orchestrator and each agent. /home/renzof/.claude / ${HOME}/.claude
ROBOCO_HOST_DATA_DIR The host data dir handed to agents for shared volumes (workspaces, logs, grok-usage). /volume1/roboco/data
ROBOCO_DATA_DIR Host root for all persistent volumes mounted into the backing services and orchestrator (see below). ./data
ROBOCO_HOST_GROK_DIR Host ~/.grok SuperGrok auth — only needed if you run any agent on Grok. /home/renzof/.grok

!!! danger "These must be real, absolute host paths" A relative path or a path that only exists inside the orchestrator container will make agent spawns fail, because the host Docker daemon resolves the bind. On a NAS the project and data dirs usually live on the RAID volume (e.g. /volume1/roboco and /volume1/roboco/data).

The host ~/.grok is mounted read-write into the orchestrator (it rewrites the short-lived token in place to keep agents from hanging on an expired login) and read-only into each Grok agent. Run grok login on the host once before enabling Grok. Provider routing and the Grok runtime are covered in the models section.

Data persistence and backup

Everything durable lives under ROBOCO_DATA_DIR (default ./data). On a server, point this at a RAID volume:

ROBOCO_DATA_DIR=/volume1/roboco/data
Subdirectory Holds
postgres/ The entire database — tasks, projects, work sessions, journals, encrypted git tokens, the pgvector store.
redis/ Append-only cache, sessions, rate-limit + event-bus state.
ollama/ The local model cache (embedding model + local LLM) — large, but re-pullable.
workspaces/ Each agent's git clone of each project.
logs/ Per-agent run logs.
mcp-configs/, prompts-generated/, agent-settings/, briefings/, manifests/ Per-agent spawn artifacts the orchestrator writes.
grok-usage/ Per-agent Grok cost/usage capture.

For backup, the load-bearing directory is postgres/ (everything that isn't re-derivable). ollama/ and workspaces/ are reconstructible — Ollama re-pulls models, agents re-clone repos — so they're optional in a backup. Take Postgres backups with pg_dump against the published port rather than copying the live data directory:

pg_dump -h localhost -p 15432 -U roboco roboco > roboco-backup.sql

!!! danger "Back up ROBOCO_ENCRYPTION_KEY with the database" Every per-project GitHub token in the database is Fernet-encrypted with ROBOCO_ENCRYPTION_KEY. A database backup is useless without the key. If you lose or change the key, every stored token becomes undecryptable and must be re-entered project by project. Store the key with your secrets, keep it stable across restarts, and never commit .env.

Secure mode

On a trusted LAN RoboCo runs in header-trust mode by default (ROBOCO_AGENT_AUTH_REQUIRED=false): callers are identified by role headers, no token required. That's the intended homelab setup.

To harden it so one agent can't spoof another's role, turn on fail-closed auth:

ROBOCO_AGENT_AUTH_REQUIRED=true
ROBOCO_AGENT_AUTH_SECRET=<your HMAC secret>     # already required for docker compose
ROBOCO_PANEL_AGENT_TOKEN=<from make panel-token>

With auth required, every API call must carry a valid X-Agent-Token. The panel runs in your browser and can't hold the signing secret, so nginx injects the CEO's token for it: generate the token with make panel-token (it signs one using your ROBOCO_AGENT_AUTH_SECRET), put it in ROBOCO_PANEL_AGENT_TOKEN, and nginx adds it as X-Agent-Token on /api and /ws. The panel keeps working; the secret never reaches the browser.

ROBOCO_ENCRYPTION_KEY and ROBOCO_AGENT_AUTH_SECRET are both required for any docker compose run — the orchestrator service block guards them with compose :? so the stack refuses to start if either is unset. See Security for the full sandboxing model and the env reference for every knob.

Startup sequence

depends_on conditions enforce a strict boot order; the effective sequence is:

flowchart LR
    PG[postgres] --> OL[ollama]
    RD[redis] --> OL
    OL --> OI[ollama-init]
    OI --> OR[orchestrator]
    AB[agent-base-image] --> OR
    OR --> PN[panel]
    PN --> NG[nginx]
    OR --> NG
  • postgres / redis / ollama must each pass their healthcheck (pg_isready, redis-cli ping, ollama list) before anything downstream starts.
  • ollama-init is a one-shot that best-effort pulls the embedding model and the local LLM, then gates success on the models being present — a degraded model registry can't take down a fully-cached deployment.
  • orchestrator waits for postgres + redis + ollama healthy, ollama-init completed, and agent-base-image built. On startup it runs the database migrations itself (idempotently, to head) and indexes its knowledge base — you do not run alembic upgrade head by hand for the compose path.
  • panel waits for the orchestrator; nginx waits for both.

First boot is the slow one: the model pulls (the LLM is a couple of minutes) plus knowledge-base indexing. Watch it come up:

docker compose logs -f orchestrator
curl http://localhost:8000/health
docker ps --filter name=roboco

When the orchestrator reports serving, open http://localhost:3000. A boot that hangs is almost always waiting on ollama-init (model pull) or a healthcheck — check docker compose ps to see which service is still starting. Migration and data details are in Data & migrations; recurring boot symptoms are in Common issues.

Operator-relevant Makefile targets

The Makefile drives the host developer workflow (uv-based, for hacking on RoboCo itself) — it is separate from the Docker stack and needs uv on the host. The handful that matter operationally:

Target Does
make panel-token Prints a signed CEO token for ROBOCO_PANEL_AGENT_TOKEN (secure mode).
make infra Brings up only postgres + redis (make infra-down stops them) — for host-side dev against the backing services.
make migrate Runs alembic upgrade head on the host (the compose stack self-migrates; this is the host-dev path).
make run Runs the API + orchestrator on the host (no --reload); make api is the reload dev server, make dev runs both.
make quality The full merge gate: ruff format-check + lint, mypy, pytest with 80% coverage floor, complexity, security, dependency, and migration checks.
make serve-docs Serves this documentation locally with mkdocs serve.
make status / make logs Orchestrator status / recent logs against a running instance.

Run make help for the full list.

Next