mirror of
https://github.com/rennf93/roboco.git
synced 2026-08-03 07:23:24 +02:00
* fix(gateway): push the branch before QA handoff so reviewers see the latest commits The commit content tool commits locally without pushing; only open_pr pushed the branch. On the first submission that was fine, but a fix committed while addressing needs_revision never reached origin (open_pr is skipped once the PR exists), so QA — which reviews the remote PR branch — re-reviewed the stale remote and re-failed the task on every cycle, a loop that never converged. i_am_done now pushes the task branch (idempotent; a no-op when nothing is unpushed) as part of the shared submit gate, covering both the normal and resume-from-verifying paths. A push failure blocks the handoff with a clear remediation rather than parking the task in awaiting_qa with commits that exist only in the developer's local workspace. * fix(orchestrator): don't reap a stale claim while the agent's container is alive The stale-claim reaper released any claimed/in_progress task whose last_heartbeat_at exceeded the TTL. The heartbeat only updates on certain gateway calls, so a developer deep in a long edit/test cycle outran the TTL and had its claim reaped mid-work — churning the task and risking a double spawn against the still-running container. The reaper now skips a task whose assignee still holds a live (ACTIVE) agent instance, trusting container liveness — the ground truth — over the heartbeat proxy. The check is defensive on missing fields so a heartbeat-only caller (and the reaper's existing unit tests) behave exactly as before. * fix(gateway): refuse to unblock a task while a dependency is unfinished A PM unblock on a dependency-gated task moved it straight to in_progress, overriding the dependency — letting a dependent proceed without its upstream's work (e.g. a frontend task built before its UX design lands). A dependency block is meant to clear on its own via _unblock_dependents the moment the upstream reaches a terminal state. unblock now refuses while any dependency is still non-terminal, returning a clear remediation that the block resolves automatically. Manual unblock remains available for genuine, non-dependency blockers. * fix(gateway): release a dependency-blocked claim to pending instead of looping A task that reached claimed/in_progress with an unfinished dependency was left in that state when the claim guard rejected, so the orchestrator's respawn loop kept reviving its assignee — which could make no progress — burning work for nothing. The claim guard now releases such a task back to pending. claimed -> blocked is not a legal transition, so pending — held by the dispatch dependency filter — is the lifecycle-correct resting state: the respawn loop ignores pending tasks, and _unblock_dependents re-dispatches it once the upstream reaches a terminal state. release_dependency_blocked_claim shares a _force_unclaim_to_pending core with unclaim_for_reaper so both record a truthful work-session abandon reason. * feat(security): warn at startup in header-trust mode + document the auth posture When ROBOCO_AGENT_AUTH_REQUIRED is not enabled the API accepts the X-Agent-Id / X-Agent-Role headers without a signed token, so any client that can reach it may act as any role (including 'ceo'). The API now logs a clear warning at startup in this mode, and the README gains a Security section documenting the auth posture and how to harden it. Acceptable only on a trusted private network — do not expose the API to untrusted networks. * fix(workspace): scope the refresh fetch to current + default branch ensure_workspace's healthy short-circuit ran an all-refs 'git fetch origin' to keep every origin/<branch> ref current. On a monorepo with many accumulated feature/* branches that exceeds the refresh timeout, the fetch silently fails, and the workspace keeps a stale base — so an agent builds on an out-of-date branch. The refresh now fetches only the workspace's current branch and the repo's default branch (resolved via origin/HEAD), with --no-tags --prune: it transfers near-nothing and can't time out. Readers need their own branch and the default; the integration branch is refreshed at branch-creation time. * fix(git): refresh a dependency-blocked task's branch off the current integration tip A cross-cell dependent (e.g. a frontend task waiting on the UX design) was branched off a base captured before its upstream merged into the integration branch, and the branch was never re-synced — so the agent built on a stale snapshot with none of the upstream's work. Two changes close the gap: - release_dependency_blocked_claim now clears branch_name, so the re-claim (after the dependency clears) re-runs branch creation. - create_branch, when the branch is already on disk with no commits of its own, resets it onto the freshly-pulled base — the dependent now builds on the current integration tip. A branch carrying real commits is left untouched, so no work is discarded; the cell->leaf cascade carries the upstream down to the dev branch automatically. * refactor(gateway): drop the sibling-sequence claim guard Sibling sequence no longer gates a claim. Cross-cell ordering is enforced by task dependencies — a cell task that depends on another is held until its upstream reaches a terminal state, a stronger, status-aware gate than the sequence-number check. That check was dormant in practice anyway: every fan-out child carries sequence 0, on which the guard short-circuited. `sequence` stays a sibling-ordering / dispatch-priority field (list_pending ordering and the panel). Removes sibling_sequence_guard and its _earlier_blocking_sibling helper, the now-unused skip_sequence parameter threaded through the claim verbs, and the sibling fetch that fed it. * feat(gateway): sort a cross-cell dependent after its upstream When the frontend cell task is wired to depend on its UX/UI sibling, set its sequence to the upstream's sequence + 1 so it sorts after the design it waits on — list_pending ordering and the panel now show UX ahead of the implementation it gates, in either delegation order. Adds TaskService.set_sequence (the sibling-ordering field is a service write; it carries no claim-gating semantics — dependencies gate claims). * feat(gateway): make the backend cell depend on UX too UX/UI design defines the screens and API contracts both implementation cells build against, so the backend cell — not just the frontend — waits on the UX/UI cell task in a product fan-out and sorts after it. Wires in either delegation order: a backend task delegated after UX gets the dependency directly; a UX task delegated after a still-pending backend sibling retro-wires it. Mirrors the existing frontend wiring (_depend_backend_on_ux and _depend_pending_backends_on_ux). Backend is held by the same dependency gate, so it costs no extra dispatch churn. * fix(websocket): forward notification acks instead of logging them incomplete The bridge handler serves both notification.sent and notification.acked, but acked events carry `agent_id` (the acking agent) rather than `recipient_id`, so every acknowledgement tripped the missing-field guard and logged "Incomplete notification event" instead of reaching the panel. Accept either field as the recipient. * feat(api): hint the full UUID when a truncated task id fails validation Agents copy the 8-character task prefix the system shows them (the commit prefix, task summaries) and send it as task_id, which fails UUID validation with an opaque "invalid length" 422 and wastes a call. The request-validation handler now detects a task_id UUID error and attaches a `remediate` hint telling the agent to retry with the full 36-character UUID from its task envelope. * fix(audit): record the blocked transition when a task is escalated Escalation sets a task to blocked by writing task.status directly, which bypassed the validated transition helper and so never emitted a task.blocked audit row — the lifecycle moved but the Auditor saw nothing. Extract the audit emit from the central transition helper into _emit_status_transition_audit and call it from the escalate path, capturing the prior status and outgoing owner before reassignment so the row is attributed correctly. * fix(docs): stop doubling the docs path so design specs index into RAG The documenter sometimes hands a doc path already rooted at docs/, and joining it onto DOCS_BASE_PATH (/app/docs) produced /app/docs/docs/..., so the file was never found and the spec never indexed — the frontend cell could not retrieve the UX design over RAG. Normalize the path before joining: trust an absolute path, otherwise strip a single redundant leading docs/ segment. * feat(security): let the control panel authenticate in secure mode With ROBOCO_AGENT_AUTH_REQUIRED=true every request must carry a valid HMAC token, which locked the human control panel out — it sends role headers but no token. nginx, the only trusted hop between the browser and the API, now injects the CEO token on /api and /ws, so the browser never holds the signing secret. The injected value is just the existing per-agent token issued for the CEO identity (issue_panel_token), so the token-verification path is unchanged. An empty value (dev/header-trust mode) renders to no header. `make panel-token` prints the value; set it as ROBOCO_PANEL_AGENT_TOKEN in .env before enabling secure mode. .env.example and the README Security section document the flow. * chore(compose): consolidate the two compose files into one docker-compose.yml and docker-compose.yaml had diverged: .yml — the file Docker actually uses — carried ROBOCO_PUBLIC_BASE_URL but was missing the /app/manifests bind-mount, while .yaml had the manifests mount but not the base URL. Merge the union into docker-compose.yml and delete the duplicate so there is one source of truth and no "multiple config files" warning. This activates the manifests mount in the deployed file: without it the orchestrator writes per-agent tool manifests to its ephemeral container fs, they never reach the host for the daemon to bind-mount, and agents fall back to all-verbs registration. Drop the stale .yaml reference from the config.py docstring, the labeler, and the CI path filters. --------- Co-authored-by: Renn F <rennf93@users.noreply.github.com>
330 lines
13 KiB
YAML
330 lines
13 KiB
YAML
services:
|
|
# ==========================================================================
|
|
# PostgreSQL - Primary Database with pgvector for RAG
|
|
# ==========================================================================
|
|
postgres:
|
|
image: pgvector/pgvector:pg16
|
|
container_name: roboco-postgres
|
|
restart: unless-stopped
|
|
environment:
|
|
POSTGRES_USER: roboco
|
|
POSTGRES_PASSWORD: roboco
|
|
POSTGRES_DB: roboco
|
|
ports:
|
|
- "15432:5432"
|
|
volumes:
|
|
- ${ROBOCO_DATA_DIR:-./data}/postgres:/var/lib/postgresql/data
|
|
healthcheck:
|
|
test: ["CMD-SHELL", "pg_isready -U roboco -d roboco"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 5
|
|
|
|
# ==========================================================================
|
|
# Redis - Cache, Sessions, Event Bus
|
|
# ==========================================================================
|
|
redis:
|
|
image: redis:8-alpine
|
|
container_name: roboco-redis
|
|
restart: unless-stopped
|
|
command: redis-server --appendonly yes
|
|
ports:
|
|
- "16379:6379"
|
|
volumes:
|
|
- ${ROBOCO_DATA_DIR:-./data}/redis:/data
|
|
healthcheck:
|
|
test: ["CMD", "redis-cli", "ping"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 5
|
|
|
|
# ==========================================================================
|
|
# Ollama - Local LLM and Embedding Server
|
|
# ==========================================================================
|
|
ollama:
|
|
image: ollama/ollama:latest
|
|
container_name: roboco-ollama
|
|
restart: unless-stopped
|
|
environment:
|
|
OLLAMA_API_KEY: ${OLLAMA_API_KEY}
|
|
ports:
|
|
- "11435:11434"
|
|
volumes:
|
|
- ${ROBOCO_DATA_DIR:-./data}/ollama:/root/.ollama
|
|
healthcheck:
|
|
# Use ollama CLI (guaranteed available) to check if server is responding
|
|
test: ["CMD", "ollama", "list"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 5
|
|
start_period: 10s
|
|
|
|
# Ollama model puller - pulls required models on startup
|
|
# Uses streaming curl to wait for full model download
|
|
ollama-init:
|
|
image: curlimages/curl:latest
|
|
container_name: roboco-ollama-init
|
|
depends_on:
|
|
ollama:
|
|
condition: service_healthy
|
|
restart: "no"
|
|
entrypoint: ["/bin/sh", "-c"]
|
|
command:
|
|
- |
|
|
set -e
|
|
echo "=== Pulling embedding model (qwen3-embedding:0.6b) ==="
|
|
# Ollama /api/pull streams JSON lines until complete - consume full stream
|
|
# Note: $$ escapes $ for docker-compose variable substitution
|
|
curl -sN http://ollama:11434/api/pull -d '{"name":"qwen3-embedding:0.6b"}' | while read -r line; do
|
|
status=$$(echo "$$line" | grep -o '"status":"[^"]*"' | cut -d'"' -f4)
|
|
[ -n "$$status" ] && echo " $$status"
|
|
done
|
|
echo "=== Pulling LLM model (glm-5:cloud) ==="
|
|
curl -sN http://ollama:11434/api/pull -d '{"name":"glm-5:cloud"}' | while read -r line; do
|
|
status=$$(echo "$$line" | grep -o '"status":"[^"]*"' | cut -d'"' -f4)
|
|
[ -n "$$status" ] && echo " $$status"
|
|
done
|
|
echo "=== Verifying models are available ==="
|
|
curl -sf http://ollama:11434/api/tags | grep -q "qwen3-embedding" && echo " qwen3-embedding: OK"
|
|
curl -sf http://ollama:11434/api/tags | grep -q "glm-5" && echo " glm-5: OK"
|
|
echo "=== All models ready! ==="
|
|
|
|
# ==========================================================================
|
|
# Agent Base Image Builder (specialized images built on-demand by orchestrator)
|
|
# ==========================================================================
|
|
agent-base-image:
|
|
build:
|
|
context: .
|
|
dockerfile: docker/agent-base.Dockerfile
|
|
image: roboco-agent-base
|
|
container_name: roboco-agent-base-builder
|
|
entrypoint: ["/bin/sh", "-c", "echo 'Agent base image built successfully'"]
|
|
restart: "no"
|
|
|
|
# ==========================================================================
|
|
# Agent PM Image Builder (specialized image built on-demand by orchestrator)
|
|
# ==========================================================================
|
|
agent-pm-image:
|
|
build:
|
|
context: .
|
|
dockerfile: docker/agent-pm.Dockerfile
|
|
image: roboco-agent-pm
|
|
entrypoint: ["/bin/sh", "-c", "echo 'Agent PM image built'"]
|
|
restart: "no"
|
|
depends_on:
|
|
- agent-base-image
|
|
|
|
# ==========================================================================
|
|
# Agent Backend Dev Image Builder (specialized image built on-demand by orchestrator)
|
|
# ==========================================================================
|
|
agent-dev-be-image:
|
|
build:
|
|
context: .
|
|
dockerfile: docker/agent-dev-be.Dockerfile
|
|
image: roboco-agent-dev-be
|
|
entrypoint: ["/bin/sh", "-c", "echo 'Agent Backend Dev image built'"]
|
|
restart: "no"
|
|
depends_on:
|
|
- agent-base-image
|
|
|
|
# ==========================================================================
|
|
# Agent Frontend Dev Image Builder (specialized image built on-demand by orchestrator)
|
|
# ==========================================================================
|
|
agent-dev-fe-image:
|
|
build:
|
|
context: .
|
|
dockerfile: docker/agent-dev-fe.Dockerfile
|
|
image: roboco-agent-dev-fe
|
|
entrypoint: ["/bin/sh", "-c", "echo 'Agent Frontend Dev image built'"]
|
|
restart: "no"
|
|
depends_on:
|
|
- agent-base-image
|
|
|
|
# ==========================================================================
|
|
# Agent Backend QA Image Builder (specialized image built on-demand by orchestrator)
|
|
# ==========================================================================
|
|
agent-qa-be-image:
|
|
build:
|
|
context: .
|
|
dockerfile: docker/agent-qa-be.Dockerfile
|
|
image: roboco-agent-qa-be
|
|
entrypoint: ["/bin/sh", "-c", 'echo "Agent Backend QA image built"']
|
|
restart: "no"
|
|
depends_on:
|
|
- agent-base-image
|
|
|
|
# ==========================================================================
|
|
# Agent Frontend QA Image Builder (specialized image built on-demand by orchestrator)
|
|
# ==========================================================================
|
|
agent-qa-fe-image:
|
|
build:
|
|
context: .
|
|
dockerfile: docker/agent-qa-fe.Dockerfile
|
|
image: roboco-agent-qa-fe
|
|
entrypoint: ["/bin/sh", "-c", 'echo "Agent Frontend QA image built"']
|
|
restart: "no"
|
|
depends_on:
|
|
- agent-base-image
|
|
|
|
# ==========================================================================
|
|
# Agent UX/UI Dev Image Builder (specialized image built on-demand by orchestrator)
|
|
# ==========================================================================
|
|
agent-ux-image:
|
|
build:
|
|
context: .
|
|
dockerfile: docker/agent-ux.Dockerfile
|
|
image: roboco-agent-ux
|
|
entrypoint: ["/bin/sh", "-c", 'echo "Agent UX/UI image built"']
|
|
restart: "no"
|
|
depends_on:
|
|
- agent-base-image
|
|
|
|
# ==========================================================================
|
|
# Agent Documenter Image Builder (specialized image built on-demand by orchestrator)
|
|
# ==========================================================================
|
|
agent-doc-image:
|
|
build:
|
|
context: .
|
|
dockerfile: docker/agent-doc.Dockerfile
|
|
image: roboco-agent-doc
|
|
entrypoint: ["/bin/sh", "-c", 'echo "Agent Documenter image built"']
|
|
restart: "no"
|
|
depends_on:
|
|
- agent-base-image
|
|
|
|
# ==========================================================================
|
|
# Orchestrator - API Server + Agent Spawner
|
|
# ==========================================================================
|
|
orchestrator:
|
|
build:
|
|
context: .
|
|
dockerfile: docker/orchestrator.Dockerfile
|
|
image: roboco-orchestrator
|
|
container_name: roboco-orchestrator
|
|
restart: unless-stopped
|
|
ports:
|
|
- "8000:8000"
|
|
environment:
|
|
# Database (use container name, not localhost)
|
|
ROBOCO_DATABASE_HOST: roboco-postgres
|
|
ROBOCO_DATABASE_PORT: 5432
|
|
ROBOCO_DATABASE_USER: roboco
|
|
ROBOCO_DATABASE_PASSWORD: roboco
|
|
ROBOCO_DATABASE_NAME: roboco
|
|
# Redis (use container name)
|
|
ROBOCO_REDIS_HOST: roboco-redis
|
|
ROBOCO_REDIS_PORT: 6379
|
|
# API
|
|
ROBOCO_HOST: 0.0.0.0
|
|
ROBOCO_PORT: 8000
|
|
ROBOCO_ENCRYPTION_KEY: ${ROBOCO_ENCRYPTION_KEY:?ROBOCO_ENCRYPTION_KEY is required}
|
|
# HMAC secret for agent auth tokens. Orchestrator signs tokens
|
|
# per-agent at spawn; API middleware verifies them. Generate with:
|
|
# python -c 'import secrets; print(secrets.token_hex(32))'
|
|
ROBOCO_AGENT_AUTH_SECRET: ${ROBOCO_AGENT_AUTH_SECRET:?ROBOCO_AGENT_AUTH_SECRET is required}
|
|
# Set to "true" to require tokens on every API call (fail-closed).
|
|
# Leave unset/false during rollout so the panel + curl still work.
|
|
ROBOCO_AGENT_AUTH_REQUIRED: ${ROBOCO_AGENT_AUTH_REQUIRED:-false}
|
|
# Ollama (use container name)
|
|
ROBOCO_LOCAL_LLM_BASE_URL: http://roboco-ollama:11434/v1
|
|
ROBOCO_LOCAL_LLM_MODEL: glm-5:cloud
|
|
ROBOCO_DEFAULT_EMBEDDING_MODEL: qwen3-embedding:0.6b
|
|
ROBOCO_OLLAMA_BASE_URL: http://roboco-ollama:11434
|
|
# Host paths for spawning agent containers (required for Docker-in-Docker)
|
|
# IMPORTANT: These must be ABSOLUTE paths on the host filesystem
|
|
ROBOCO_HOST_PROJECT_DIR: ${ROBOCO_HOST_PROJECT_DIR:-/volume1/roboco}
|
|
ROBOCO_HOST_CLAUDE_DIR: ${ROBOCO_HOST_CLAUDE_DIR:-/home/renzof/.claude}
|
|
ROBOCO_HOST_DATA_DIR: ${ROBOCO_HOST_DATA_DIR:-/volume1/roboco/data}
|
|
# Public base URL for commit-trailer links. Default 127.0.0.1 produces
|
|
# unusable links in commit message bodies; set to NAS LAN IP so
|
|
# f"{api_base}/tasks/{task_id}" renders a reachable URL.
|
|
ROBOCO_PUBLIC_BASE_URL: "http://192.168.50.111:8000"
|
|
# Production environment selects structlog's JSONRenderer (machine-
|
|
# parseable logs) over the dev ConsoleRenderer.
|
|
ROBOCO_ENVIRONMENT: production
|
|
volumes:
|
|
# Docker socket - allows spawning agent containers
|
|
- /var/run/docker.sock:/var/run/docker.sock
|
|
# Claude Code auth - mount your ~/.claude directory
|
|
- ${CLAUDE_AUTH_DIR:-/home/renzof/.claude}:/root/.claude
|
|
# Shared config directory for MCP configs (writable)
|
|
- ${ROBOCO_DATA_DIR:-./data}/mcp-configs:/app/mcp-configs
|
|
# Generated prompts directory - composed at runtime from layers
|
|
- ${ROBOCO_DATA_DIR:-./data}/prompts-generated:/app/prompts-generated
|
|
# Per-agent Claude settings (generated at spawn time)
|
|
- ${ROBOCO_DATA_DIR:-./data}/agent-settings:/app/agent-settings
|
|
# Agent workspaces (git clones) - persisted across restarts
|
|
- ${ROBOCO_DATA_DIR:-./data}/workspaces:/data/workspaces
|
|
# Persistent logs — survive `docker compose down/up`. Orchestrator and
|
|
# each spawned agent write structured logs here so we can audit past
|
|
# runs instead of relying on ephemeral `docker logs`.
|
|
- ${ROBOCO_DATA_DIR:-./data}/logs:/data/logs
|
|
# Per-agent SessionStart briefings (pre-rendered task context)
|
|
- ${ROBOCO_DATA_DIR:-./data}/briefings:/app/briefings
|
|
# Per-agent spawn manifests (role-scoped tool list) — written by the
|
|
# orchestrator to /app/manifests, bind-mounted into each agent container
|
|
# as /app/tool-manifest.json. Without this mount the file is written to
|
|
# the orchestrator's ephemeral fs, never reaches the host, and agents
|
|
# fall back to all-verbs registration.
|
|
- ${ROBOCO_DATA_DIR:-./data}/manifests:/app/manifests
|
|
depends_on:
|
|
postgres:
|
|
condition: service_healthy
|
|
redis:
|
|
condition: service_healthy
|
|
ollama:
|
|
condition: service_healthy
|
|
ollama-init:
|
|
condition: service_completed_successfully
|
|
agent-base-image:
|
|
condition: service_completed_successfully
|
|
# Default agents to spawn (override in .env or command line)
|
|
# command: ["--spawn", "main-pm", "be-dev-1", "be-qa"]
|
|
|
|
# ==========================================================================
|
|
# Next.js Control Panel (Frontend)
|
|
# ==========================================================================
|
|
# Not exposed directly - nginx is the single entry point (port 3000).
|
|
# API/WS traffic is proxied to the orchestrator, everything else to panel.
|
|
panel:
|
|
build:
|
|
context: .
|
|
dockerfile: docker/panel.Dockerfile
|
|
image: roboco-panel
|
|
container_name: roboco-panel
|
|
restart: unless-stopped
|
|
expose:
|
|
- "3000"
|
|
depends_on:
|
|
- orchestrator
|
|
|
|
# ==========================================================================
|
|
# Nginx - Reverse proxy fronting panel + orchestrator
|
|
# ==========================================================================
|
|
# Single entry point on port 3000 so the browser hits one origin and we
|
|
# don't need CORS. /api/* and /ws/* go to the orchestrator, everything
|
|
# else goes to the Next.js panel.
|
|
nginx:
|
|
image: nginx:alpine
|
|
container_name: roboco-nginx
|
|
restart: unless-stopped
|
|
ports:
|
|
- "3000:80"
|
|
environment:
|
|
# Rendered into the proxy config by the nginx image's envsubst
|
|
# entrypoint so the human panel authenticates in secure mode. The
|
|
# filter limits substitution to ROBOCO_* vars, leaving nginx's own
|
|
# $host / $remote_addr runtime variables untouched. Get the value
|
|
# with `make panel-token`.
|
|
ROBOCO_PANEL_AGENT_TOKEN: ${ROBOCO_PANEL_AGENT_TOKEN:-}
|
|
NGINX_ENVSUBST_FILTER: "^ROBOCO_"
|
|
volumes:
|
|
- ./docker/nginx.conf:/etc/nginx/templates/default.conf.template:ro
|
|
depends_on:
|
|
- panel
|
|
- orchestrator
|
|
|
|
networks:
|
|
default:
|
|
name: roboco_default
|