* feat(metrics): capture per-session turns + tool_calls (phase 1)
Persist LLM iterations (turns) and tool invocations per agent spawn session,
the raw signal the granular per-member performance metrics build on (real
effort/iterations vs wall-clock).
- sum_transcript_usage returns a 5-tuple adding turns = unique assistant
message-id count; _usage_from_transcript + _resolve_active_tokens updated to
the 5-tuple (active-tokens keeps its 4-tuple contract by slicing).
- SDK: _SessionState.turns, set by /usage/sync; /usage/status (TokenUsageStatus)
now carries turns + tool_calls (= total_calls).
- orchestrator: new _resolve_final_turns_tools (SDK primary, transcript fallback
for turns only; Grok -> 0/0) wired into _finalize_spawn_session, which writes
turns + tool_calls to agent_spawn_sessions.
- migration 055 adds turns + tool_calls (BigInteger DEFAULT 0 -> historical/Grok
rows read 0, surfaced as n/a). Verified real alembic upgrade/downgrade.
Part of metrics-granularity (v0.15.0); recon-adjusted plan on disk.
* feat(metrics): pure compute_stage_effort helper (phase 2, part 1)
Foundation-layer overlap math (no DB): split each task status window into
active (merged wall-clock overlap of spawn stints — concurrent stints counted
once, so active <= window) vs wait (queue/review idle). Distinct from summed
effort. The per-task metrics service will feed it audit-log windows + spawn
stints. 9 unit tests (disjoint/nested/partial/merged/clamped/zero/multi-window).
* feat(metrics): per-task live metrics + GET /metrics/task/{id} (phase 2)
TaskMetrics dataclass + MetricsService.get_task_metrics: summed spawn effort
(vs wall-clock), turns/tool_calls/tokens/cost, per-stage active-vs-wait
(compute_stage_effort over audit windows x spawn stints), and who-caused-rework
(revision_count + named qa/pr fail events). Open stints and the open final
stage window close at completed_at for a terminal task (else now), so stages
don't grow past completion. Exposed at GET /dashboard/metrics/task/{task_id}
(404 if absent). Real-PG tests (compose/none/in-flight) + route tests (200/404).
* feat(metrics): CEO-as-member scorecard + ceo_reject audit regression (phase 3)
The human CEO is a measured member, read purely from audit_log (agent_role='ceo'
serializes from the CEO StrEnum): approval dwell (awaiting_ceo_approval -> a CEO
decision, incl. the coordination-root reject that lands in pending), unblock
dwell (blocked -> a CEO revive), and god-mode action count (every CEO-attributed
transition). CeoScorecard + MetricsService.get_ceo_scorecard (p50/p90 via
PERCENTILE_CONT, expanding IN for the decision sets) + GET
/dashboard/metrics/member/ceo (declared before any future member/{id} route).
The ceo_reject coordination-root audit gap the plan meant to close was already
closed by the gap-sweep (routes through admin_set_status -> agent_role='ceo'
audit); locked with a regression assertion in the existing coordination-reject
test. Real-PG tests: approval/unblock/godmode, non-ceo exclusion, empty->zeros.
* feat(metrics): audit instrumentation for escalations/blocked-others/idle (phase 4a)
The three extra per-member metrics that had no data source get durable,
in-session audit events (additive; never gate the underlying action):
- apply_escalation -> task.escalated (details.escalator_slug) on both the
normal block path and the pool-divert path -> escalations count.
- _unblock_dependents -> task.unblocked_dependents (details.count) on the
completed BLOCKER task, captured before the dependency edges are pruned ->
blocked-others count (sweeper attributes to the blocker's owner).
- mark_agent_idle -> agent.idle (details.agent_slug) -> idle/utilization (the
sweeper pairs an idle mark to the member's next spawn for idle duration).
(QA pass-rate needs no new event — reuses task.awaiting_documentation[qa] +
task.qa_fail.) Real-PG tests for each; 111 transition tests still green.
* feat(metrics): member_performance_daily rollup table + migration 056 (phase 4b)
The per-member scorecard rollup: one row per (date, member_kind, agent_slug),
CEO as a first-class member_kind='ceo' row (agent_slug='' NOT NULL so the
NULL-distinct UNIQUE keeps it unique). Full column set + the four CEO-approved
extras (qa_reviews_total/passed, escalations, blocked_others, idle_seconds) plus
blocked_seconds. Overwrite-upsert on (date, member_kind, agent_slug) for an
idempotent sweep. Migration 056 verified real up/down (24 cols, 4 indexes).
* feat(metrics): _sweep_member_performance rollup sweeper (phase 4c)
The daily per-member rollup sweep (mirrors _sweep_daily_rollup): a trailing
7-day, idempotent overwrite-upsert wired into _run_sweep. One focused query per
metric merges into a (date, agent_slug) accumulator — spawn effort/turns/tokens/
cost, completed/first-pass/revisions-received, revisions-caused (qa/pr fails),
QA pass-rate (passed + total), escalations (by escalator_slug), blocked-others
(unblocked_dependents by blocker owner), idle_seconds (idle mark -> next spawn),
blocked_seconds (blocked dwell) — plus one CEO row/day (approval/unblock dwell +
god-mode). Real-PG test asserts every facet + idempotency (a 2nd sweep
overwrites, never doubles); spawn-day != completion-day split is by-design.
* feat(metrics): member/org rollup scorecards + endpoints + live overlay (phase 5)
MemberScorecard + OrgScorecard with derived rates (FPY, effort-throughput,
turns/tool-calls per task, QA pass-rate, utilization) — all division-guarded to
None. get_member_scorecard reads member_performance_daily by slug and overlays
the member's live in-flight (non-terminal) tasks' effort via get_task_metrics
(disjoint by status: completion counts stay rollup-only, overlay only enriches
effort/turns/cost; includes_live_inflight flags it). get_org_scorecard
aggregates the cell (?team=) or whole org. Routes: GET /metrics/member/{agent_id}
(404 if absent, after the ceo literal route) + GET /metrics/org?team=. Real-PG
tests (derived rates, overlay no double-count, guards, org) + route tests.
* feat(metrics): granular CEO completion notification (phase 6)
There was no CEO completion notification at all (EventType.TASK_COMPLETED was
defined but never emitted). Add notify_ceo_of_completion in
NotificationDeliveryService — a granular body (real effort vs wall-clock +
stints/turns/tool-calls/revisions[QA/PR]/cost from get_task_metrics; degrades to
wall-clock-only, turns 'n/a', when there are no spawn sessions). Reuses the
existing ALERT type (no enum migration; the notificationtype PG enum is fixed at
001). ceo_approve now emits TASK_COMPLETED + fires the notification (best-effort
via _notify_completion — never blocks completion); complete() emits
TASK_COMPLETED too (closes the dead-code gap; the WS bridge can forward it).
Pure formatter tests + real-PG notification test.
* [metrics-granularity] Phase 7: panel Scorecards tab + dashboard overview
Add the CEO-facing metrics surfaces for the granularity feature:
- New "Scorecards" tab on the Metrics page: org rollup headline, the
CEO-as-member card (approval/unblock dwell + god-mode count), and a
per-member table (completed, first-pass yield, active effort, turns/task,
QA pass-rate, escalations, blocked-others, utilization). Each member row
self-fetches its rollup scorecard; live in-flight rows carry a "live" badge.
- New dashboard overview card (ScorecardOverviewPanel): org-wide 30-day
headline (completed, FPY, throughput/hr, active effort, cost) deep-linking
into the Scorecards tab.
- Plumbing: TaskMetrics/MemberScorecard/OrgScorecard/CeoScorecard types,
observability API client methods + empty fallbacks, and the four
useCeoScorecard/useMemberScorecard/useOrgScorecard/useTaskMetrics hooks.
Panel gate green: tsc, eslint, prettier, vitest (175 tests, +6 new).
* [metrics-granularity] test: make completion-notification robust to shared-DB CEO
test_notify_ceo_of_completion_creates_alert errored in the full suite (passed
in isolation): the session-scoped test DB is shared across the run, and the
sibling real-DB board-gate test commits a role=CEO agent (slug="ceo") without
cleanup — so my env fixture's hardcoded slug="ceo" insert hit a unique-constraint
violation, and a second role=CEO row would also make _get_ceo_agent()'s
scalar_one_or_none() raise. Reuse an existing CEO when present (the singleton the
production system actually has), else create one with a unique slug. Order-
independent. Also reflow test_metrics_instrumentation.py to ruff format.
* chore(release): 0.15.0
Metrics granularity: per-member/per-task/org + CEO-as-member scorecards,
turn/tool-call capture (migration 055), member_performance_daily rollup
(migration 056) with QA pass-rate / escalations / blocked-others / utilization,
per-task active-vs-wait metrics, granular completion notification, panel
Scorecards tab + dashboard Performance card, and the ceo_reject audit fix.
Version bump across the canonical set + CHANGELOG.
* [metrics-granularity] fix pre-tag audit findings (overlay double-count + panel error states)
Adversarial review before the v0.15.0 tag surfaced two real logical gaps:
- MAJOR (backend): the live in-flight overlay re-summed ALL sessions of every
non-terminal task via get_task_metrics, but _msweep_spawn already rolls up
every CLOSED session regardless of task status — so a closed session on a
still-open task was counted twice (rollup + overlay), permanently inflating a
member's effort/turns/tokens/cost on the common reap/respawn path. The overlay
now sums only OPEN sessions (ended_at IS NULL), which the closed-only rollup
can never contain — disjoint by construction. A just-closed session lands in
the rollup on the next ~60s sweep (no gap of note). Aggregated in SQL to mirror
_msweep_spawn. Regression test reproduces the double-count (turns 10→5).
- MAJOR (panel): the four new scorecard surfaces used `isLoading || !data` with
no isError branch, so a failed query span forever on a skeleton. They now
surface a load error. Tests added.
Also: OrgSummary active-effort formatting no longer round-trips hours→seconds→
hours; dashboard grid uses xl:grid-cols-4 (was 2xl) so 4 panels show at 1280px;
corrected the inaccurate "NULL distinct" CEO-row uniqueness comment (agent_slug
is NOT NULL; the '' tuple is simply distinct from agent rows).
make quality GREEN (cov 95.31%); panel GREEN (vitest 178).
* [metrics-granularity] fix: decode bytes stream message-id before XCLAIM
StreamEventBus._recover_stream passed the pending message id to XCLAIM via
str() on the raw bytes the client returns (redis client has no
decode_responses), producing "b'1782066556728-0'". Redis rejects that with
"Unrecognized XCLAIM option", so pending-message recovery threw on every
reclaim tick and unacked messages from crashed/slow consumers were never
reclaimed (leaking in the PEL on every stream, spamming the error log). Decode
via the existing _to_str helper — the fix the sibling claim path already uses.
Pre-existing in v0.14.0 (unrelated to metrics granularity); folded into this
release per CEO. TDD regression test + CHANGELOG entry. make quality GREEN.
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
RoboCo
AI Agents Company - A virtual organization of 25 AI agents + 1 human CEO, designed to operate as a complete software development workforce.
▶ Watch the 26-min intro what it is, a walkthrough, and how to use it |
▶ Watch the 2.5-hour build session a conversation → a shipped feature |
Watch the full 2:33 walkthrough (.mp4) →
Warning
RoboCo is early-stage, work-in-progress software (v0). It's under active development, runs in a homelab, and will have rough edges, breaking changes, and bugs. It is not production-ready and the API/database schema are not stable yet. Treat it as a working prototype to explore and build on — please don't expose it to the public internet as-is. Issues and PRs very welcome.
Tip
📚 Full documentation: docs.roboco.tech — install & first run, the company model, a page-by-page panel reference, model providers, the optional subsystems, deployment, and the API.
Overview
RoboCo implements a structured organizational hierarchy with formal communication protocols, task management, and quality controls. The system enables a single human (CEO) to orchestrate complex multi-project development at scale.
CEO (You, the human)
│
├── Intake (on-demand interviewer: chats only with you to draft a task)
├── Secretary (on-demand chief-of-staff: reads company state, runs gated directives)
├── PR Reviewer (read-only main reviewer: inbound external/fork + internal PRs, and the root→master in-path gate)
│
└── Board (3 agents)
├── Product Owner
├── Head of Marketing
└── Auditor (silent observer, reports to you)
│
└── Main PM (coordinates all cells)
│
├── Backend Cell (6 agents: 2 Devs, 1 QA, 1 PM, 1 Documenter, 1 PR Reviewer)
├── Frontend Cell (6 agents: 2 Devs, 1 QA, 1 PM, 1 Documenter, 1 PR Reviewer)
└── UX/UI Cell (6 agents: 2 Devs, 1 QA, 1 PM, 1 Documenter, 1 PR Reviewer)
The 25 agents = Intake + Secretary + PR Reviewer + the Board (3) + Main PM + the three 6-agent cells (18). Agents run on Anthropic Claude by default, or on xAI Grok (the official grok CLI on a SuperGrok subscription) — see the provider note under Configuration.
How it works
You hand a task to the company; it runs through a real build → review → document → merge pipeline and comes back to you to approve.
One full loop, put simply:
- You give the Board a task — they review it. The Product Owner and Head of Marketing turn your ask into requirements and acceptance criteria.
- You approve — the Main PM starts the work. A notification asks for your Approve & Start decision; approve, and the Main PM breaks it into per-cell subtasks.
- Each cell's PM delegates, supports, and triages its developers (UX/UI, Frontend, Backend).
- Developers build it, QA verifies and gates it, Documenters keep the books.
- Cell PMs merge their PRs into the Main PM's branch.
- The Main PM opens the final PR and notifies you "It's done!" — you approve and merge, or send it back for rework. (Only you ever merge to
master.)
— Full circle —
See the full walkthrough, with screenshots →
Or watch the full panel walkthrough (video) →
Project Structure
roboco/
├── roboco/ # Main Python package
│ ├── api/ # FastAPI routes & schemas
│ │ ├── routes/ # API endpoints (tasks, git, agents, etc.)
│ │ └── schemas/ # Pydantic request/response models
│ ├── services/ # Business logic services
│ │ ├── task.py # Task lifecycle management
│ │ ├── workspace.py # Multi-agent workspace management
│ │ ├── messaging.py # Agent communication
│ │ └── optimal.py # RAG/Knowledge base (in-house pgvector)
│ ├── models/ # Pydantic domain models
│ ├── db/ # SQLAlchemy ORM & session
│ ├── enforcement/ # Task lifecycle state machine
│ ├── runtime/ # Orchestrator for agent spawning
│ ├── agents/ # Agent base classes
│ ├── mcp/ # MCP server implementations
│ └── config.py # Application configuration
├── agents/
│ └── prompts/ # Agent system prompts (roles, teams, identities)
├── docs/
│ ├── how-to/ # Visual walkthrough — 5-chapter guide (start at README.md)
│ └── rag/ # Agent knowledge base (indexed into RAG)
├── alembic/ # Database migrations
├── CLAUDE.md # Claude Code guidance
├── docker-compose.yml # Full stack, built from source
└── docker-compose.registry.yml # Full stack, pulled from the image registry
Running RoboCo
You need Docker + Docker Compose and a Claude Code auth directory on the host (~/.claude, mounted into the orchestrator so agents can reach the model). Copy .env.example to .env and set at least ROBOCO_ENCRYPTION_KEY and ROBOCO_AGENT_AUTH_SECRET (that file shows how to generate each). However you start it, the whole company is reachable at one origin: http://localhost:3000.
Optional — run agents on xAI Grok instead of Claude. RoboCo can spawn agents on xAI's official grok CLI authenticated by a SuperGrok subscription (no metered API key). Run grok login once on the host and point ROBOCO_HOST_GROK_DIR at the resulting ~/.grok so it mounts into Grok agents; the orchestrator keeps the ~6h token refreshed for you. See the Grok block in .env.example (ROBOCO_HOST_GROK_DIR, ROBOCO_GROK_AGENT_IMAGE, ROBOCO_GROK_CLI_MODEL, ROBOCO_GROK_REASONING_EFFORT).
Option 1 — Run the pre-built images (quickest)
Every release publishes all RoboCo images to both the GitHub Container Registry and Docker Hub, so you can run the full stack without building anything. Use the registry compose:
git clone https://github.com/rennf93/roboco.git && cd roboco
cp .env.example .env # then edit in your secrets
docker compose -f docker-compose.registry.yml pull
docker compose -f docker-compose.registry.yml up -d
Choose the registry and version with two env vars (defaults shown):
ROBOCO_REGISTRY=ghcr.io/rennf93 # or docker.io/renzof93
ROBOCO_VERSION=latest # or a pinned release, e.g. 0.15.0
The orchestrator spawns the matching pre-built agent images on demand — no build toolchain or source compile on your host.
Option 2 — Build from source
The same full stack, built locally from the Dockerfiles instead of pulled:
git clone https://github.com/rennf93/roboco.git && cd roboco
cp .env.example .env # then edit in your secrets
docker compose up -d # builds images on first run, then starts everything
Option 3 — Local development (no full stack)
For hacking on the code itself, run only the backing services in Docker and the API on your host. RoboCo's own code requires Python 3.13+ (uv will fetch it if needed):
uv sync
docker compose up -d postgres redis ollama # backing services only
uv run alembic upgrade head # migrate the database
uv run python -m roboco.cli # API + orchestrator
# Or just the API without the orchestrator:
uv run uvicorn roboco.api.app:app --reload --host 0.0.0.0 --port 8000
Configuration
Key environment variables (see roboco/config.py for all options):
# API Server
ROBOCO_HOST=0.0.0.0
ROBOCO_PORT=8000
# Database
ROBOCO_DATABASE_HOST=localhost
ROBOCO_DATABASE_PORT=5432
ROBOCO_DATABASE_NAME=roboco
# Workspaces (Multi-Agent Git)
ROBOCO_WORKSPACES_ROOT=/data/workspaces
ROBOCO_WORKSPACE_AUTO_CLONE=true
# RAG/LLM
ROBOCO_LOCAL_LLM_BASE_URL=http://roboco-ollama:11434/v1
ROBOCO_LOCAL_LLM_MODEL=glm-5.2:cloud
# Feature flags (default-off unless noted; toggle from Settings → Feature Flags)
ROBOCO_CONVENTIONS_ENABLED=false # per-project architectural conventions standard
ROBOCO_TOOLCHAIN_MATCH_ENABLED=false # build each target project under its own Python
ROBOCO_OVERLOAD_BREAK_ENABLED=true # park a provider on a persistent model-API overload
Multi-Agent Workspace Structure
Each agent gets their own git clone for parallel development:
{ROBOCO_WORKSPACES_ROOT}/
└── {project-slug}/
└── {team}/
└── {agent-slug}/
└── [git repository]
Example:
/data/workspaces/roboco/backend/be-dev-1/
/data/workspaces/roboco/backend/be-dev-2/
Task Lifecycle
backlog → pending → claimed → in_progress → verifying → awaiting_qa
↓ ↓ ↓ ↓
cancelled blocked needs_revision awaiting_documentation
paused ↓
awaiting_pm_review
↓
awaiting_ceo_approval
↓
completed
Assembled, PR-bearing tasks pass through one extra stage — the in-path PR-review gate — before the PM merges:
in_progress → awaiting_pr_review → awaiting_pm_review
(submit_up / (pr_pass)
submit_root) (pr_fail → needs_revision)
The cell PM's submit_up (cell→root PR) and the Main PM's submit_root (root→master PR) open the assembled PR and enter the gate; a PR reviewer pr_passes it on to the PM merge or pr_fails it back. Leaf dev tasks (reviewed by QA) and branchless coordination roots skip the gate.
API Endpoints
Domain routes are mounted under /api:
| Route Group | Description |
|---|---|
/api/tasks |
Task CRUD, lifecycle, claiming |
/api/agents |
Agent management |
/api/git |
Git operations (status, commit, push, PR) |
/api/sessions |
Communication sessions |
/api/messages |
Agent messages |
/api/projects |
Project (repo) management |
/api/work-sessions |
Git work session tracking |
/api/optimal |
RAG/Knowledge base queries |
/api/journals |
Agent journals/reflections |
/api/orchestrator/status |
Orchestrator / dispatcher status |
The agent gateway verbs are served separately under /api/v1/flow/{role}/{verb} (intent verbs) and /api/v1/do (content tools) — see the Agent Gateway.
Development
# Install dev dependencies
uv sync --all-extras
# Run tests
uv run pytest
# Format and lint
uv run ruff format .
uv run ruff check .
uv run mypy roboco/
# Type checking
uv run mypy roboco/
Core Principles
- Everything is a task - All work is tracked and documented
- No work without a task - Create task record first
- No task without acceptance criteria - How do we know it's done?
- No closure without documentation - Future agents need context
- Communication is constant - Stream reasoning, log everything
- The Auditor sees all - Quality monitored silently
- CEO approves major changes - Human-in-the-loop for critical decisions
Technology Stack
| Layer | Technology |
|---|---|
| API Framework | FastAPI |
| Database | PostgreSQL + SQLAlchemy (async) |
| Vector Store | PostgreSQL + pgvector (in-house engine) |
| Cache/Queue | Redis |
| RAG Engine | in-house (asyncpg + pgvector, hybrid retrieval) |
| Embeddings | qwen3-embedding:0.6b (Ollama) |
| Local LLM | Ollama (glm-5.2:cloud) |
| Cloud LLM | Claude API (Anthropic) + xAI Grok (official grok CLI, SuperGrok subscription) |
| Package Manager | uv |
Status
Core Infrastructure (Complete)
- Data models (Pydantic)
- Database ORM (SQLAlchemy async)
- Task lifecycle state machine
- Multi-agent workspace management
- Agent prompts (25 agents)
- Messaging API
- Task API with full lifecycle
- Git operations API
- RAG/Knowledge base (in-house pgvector engine)
- Agent orchestrator
- CEO approval workflow
- Pluggable agent providers (Claude Code + xAI Grok on the official
grokCLI) - Inbound PR review (read-only PR-reviewer + CEO supersede/dismiss queue)
- Self-healing CI loop for RoboCo's own repo (default-off, CEO-gated)
- Business Goals tab with a live Company Scorecard (delivery, spend-vs-budget, lead time)
In Progress
- Frontend panel (vendored under
panel/, served through nginx on :3000) - Full agent autonomy testing
Security
Important
Do not expose RoboCo to the public internet as-is. It is designed to run on a trusted private network (homelab / LAN).
Agent authentication. Requests identify the caller with X-Agent-Id / X-Agent-Role headers. The orchestrator issues each spawned agent an HMAC token (X-Agent-Token, signed with ROBOCO_AGENT_AUTH_SECRET) that binds its id, role and team. Token enforcement is gated by ROBOCO_AGENT_AUTH_REQUIRED:
ROBOCO_AGENT_AUTH_REQUIREDunset/false (default): header-trust mode — the role headers are accepted without a token, so any client that can reach the API may claim any role (includingceo). The API logs a warning at startup in this mode. Acceptable only on a trusted network.ROBOCO_AGENT_AUTH_REQUIRED=true: every request must carry a valid token; an agent cannot spoof another agent's role. The control panel keeps working because nginx — the only trusted hop between the browser and the API — injects the CEO token (X-Agent-Token) on/apiand/ws, so the browser never holds the signing secret. Generate that token withmake panel-tokenand set it asROBOCO_PANEL_AGENT_TOKENin.envbefore enabling secure mode.
WebSocket streams. Token enforcement is currently REST-only. The /ws/* endpoints authenticate by agent_id query param at most and do not yet validate X-Agent-Token, even in secure mode — nginx injects the token so the panel works, but a direct WebSocket connection that bypasses nginx is not rejected. In particular the operator stream /ws/system (rate-limit lifecycle + token-usage snapshots for the dashboard) is unauthenticated. These streams are read-only — no control surface, secrets, or task content — but treat the orchestrator port as trusted-network-only until WebSocket auth lands.
Secrets (the Fernet ROBOCO_ENCRYPTION_KEY, GitHub PATs) live encrypted in the database and in gitignored env files — never in the repo. Per-project git tokens are Fernet-encrypted at rest and never returned by the API.
License
Copyright (c) 2026 Renzo Franceschini
RoboCo is licensed under the GNU Affero General Public License v3.0 (AGPL-3.0). See LICENSE for the full text.
The AGPL's network-use clause (section 13) means that if you run a modified version of RoboCo as a network service, you must make your modified source available to its users. This keeps the project open while preventing closed, hosted re-distributions.
Contributing
Contributions are welcome. All contributors must sign the Contributor License Agreement (CLA.md) — this is automated on your first pull request. See CONTRIBUTING.md for the workflow and why the CLA exists.