* feat(tg): P0 — dev mock bridge + Telegram-native foundations
Mini App V4 phase 0. The (tg) shell gains the groundwork every later
phase builds on:
- Dev mock bridge: outside Telegram, a development build falls back to a
no-op WebApp object and skips the webapp-auth POST (the regular panel
session cookie authorizes API calls), so the cockpit is workable in a
plain browser. Production keeps the "Open from Telegram" wall.
- Telegram theme adoption: themeParams map onto the shadcn CSS variables
scoped to #tg-shell (desktop dashboard untouched), colorScheme drives
the dark class, themeChanged re-applies live. Non-hex values are
dropped at the trust boundary.
- Viewport/swipe correctness: shell height rides Telegram's own
--tg-viewport-stable-height (100dvh fallback), vertical swipe-to-close
disabled so list scrolling can't dismiss the app.
- Native chrome bindings: TgWebAppProvider context plus useMainButton /
useBackButton declarative hooks and a null-safe haptics helper —
consumers never touch window.Telegram directly.
* feat(tg): P1 — Today home tab + one-round-trip /telegram/today brief
Mini App V4 phase 1: the cockpit now opens on a "Today" brief answering
"does anything need me?" in one glance.
Backend: GET /api/telegram/today (CEO-gated, rate-limited) returns the
whole brief in one round trip via the new TgCockpitService — needs-you
items (awaiting-CEO + blocked tasks capped for a phone screen, held-draft
counts across release/X/video/roadmap queues), a fleet snapshot with
per-agent current-task titles, today's spend from the day rollup
(degrading to zeros on a usage hiccup, mirroring the CEO overview), and
ship state. Deliberately DB-only: no live GitHub calls, no readiness
snapshot (that path clones), no orchestrator singleton — the CI
red/green proxy is the set of open ci_watch fix tasks.
Panel: TgTodayTab is the new default tab (Gauge icon) — needs-you rows
and draft chips deep-link into the tab that acts on them (with a haptic
tap), fleet/spend/ship render as dense cards, 45s refetch until the P3
WebSocket wiring lands.
* fix(tg): dev mock engages when the CDN bridge loads outside Telegram
Live browser smoke caught it: a bare tab still loads telegram-web-app.js,
so window.Telegram.WebApp EXISTS outside Telegram — just with empty
initData. The dev fallback keyed on a null bridge only, so a dev browser
went down the real-auth path and posted empty initData instead of
mounting the mock. The fallback now treats bridge-with-no-initData the
same as no bridge (a real Telegram launch always carries initData);
production behavior is unchanged.
* feat(tg): P2 — native approvals card stack
Mini App V4 phase 2: the Approvals tab stops stacking the four desktop
queue cards and becomes a phone-native flow — one normalized list across
release proposal / X drafts / video drafts / proposed roadmap items, and
a full-context detail per item:
- Release: version/bump/gate badges, changelog draft, gaps, migration
notes, in-flight + failed-execute banners; approve runs the fail-closed
executor, reject requires a substantive change request (10 chars).
- X: editable body with the live 280 counter, replied-to mention quoted;
approve sends the edited body only when actually edited.
- Video: cut-toggled player (blob-fetched through the authed client — a
bare <video src> would 401), per-platform caption edits with 280/2200
counters; approve sends only checked-in edits.
- Roadmap: the PO's full pitch (description, rationale, ACs); approve
materializes into the backlog per item.
The detail's primary action rides Telegram's native MainButton and back
navigation rides the BackButton, with visible fallbacks outside Telegram
(dev mock, old clients). Haptics fire on outcomes. An acted-on item
vanishes from the refetched queue, popping back to the list by
construction. A failed queue source is surfaced ("list may be
incomplete" / "couldn't load") instead of masquerading as an empty
queue — caught live in the browser smoke.
* feat(tg): dev demo mode — /tg?demo=1 renders canned cockpit data
Development-only: with the flag param present, the Today brief and the
four approval queues resolve typed fixtures (dynamically imported, so
production bundles never carry them) instead of hitting the backend —
the cockpit is fully browsable with zero stack running. Mutations still
go to the real API and fail loudly; it's a showroom, not a simulator.
* feat(tg): P3 — cockpit rides /ws/system live
Chat adopts the desktop A2A invalidate-on-frame idiom over the shared
ref-counted /ws/system socket: every a2a.message frame refreshes the
conversation list and the affected thread, missed-frame gaps are healed
by a reconnect refetch, and the 10s thread poll turns off entirely while
the socket is up (it remains the fallback). The Today brief refreshes on
each USAGE_SNAPSHOT push so the spend line tracks the sweeper live, with
the 45s poll as the socket-down fallback. No new sockets, no backend
changes — the WS gate already accepts the cloud-auth session cookie.
* feat(tg): P4 — deterministic bot command tier + self-syncing menu
Mini App V4 phase 4 (deterministic half): three new bot commands beside
/status /queue /task —
- /agents: who's mid-task right now, from the same TgCockpitService
fleet snapshot the Today brief renders (now public `fleet()`).
- /usage: today's spend from the day rollup.
- /blocked: awaiting-you + blocked tasks, deep-linked into the panel,
capped per section, titles HTML-escaped.
BOT_COMMANDS is the single registry driving /help AND a once-per-process
Bot API setMyCommands sync on the first poll cycle (new client method,
best-effort), so the Telegram command menu can never drift from what the
code implements. The interactive tier (/secretary, /newtask riding a
live Intake interview in-thread) is specced but not in this commit.
* feat(tg): direction-C styling pass — Telegram palette, RoboCo voice
The cockpit stops wearing default-shadcn and gets its own visual
language on top of the P0 themeParams bridge (colors stay CSS-variable
driven, so inside Telegram everything still adopts the user's theme):
- Shared primitives (components/tg/ui.tsx): TgSection grouped cards with
tracked micro-label headers, TgRow list rows (44px targets, press
feedback, 1/2-line clamp), TgRowIcon glyph tiles, TgStat tabular-nums
figures. Every tab composes the same three, so density and rhythm are
identical across the surface.
- Shell renders a centered 430px column (sm:border-x) — the phone UI no
longer stretches across a desktop dev browser.
- Tab bar: tighter type, active stroke-weight shift, backdrop blur.
- Today: needs-you count badge, divided task rows with inline blocked
marker, fleet as mono-named rows, spend/ship as stat tiles.
- Approvals rows as icon-tile cards; detail header gains the kind glyph.
- Inbox/Chat rows aligned to the same card language.
* feat(tg): P5 — /secretary and /newtask live-chat bridges
The bot's interactive tier: both commands bridge the CEO's Telegram chat
into the same in-process runtimes the panel drives — the persistent
Secretary container and the scoped Intake interview.
There is no synchronous send→reply seam (replies land on the session's
single-consumer relay queue), so each bridged session runs one long-lived
consumer task (roboco/services/telegram_bridge.py) that drains
PrompterLiveRegistry.stream and pushes one Telegram message per completed
turn. While a session is live, plain chat text IS the conversation;
/end closes it.
/newtask resolves the intake scope (single project auto-picked, multiple
offered as a tap-to-pick keyboard holding the initial text), and the
interview happens in-thread. A draft proposal renders as a card with
Send-to-Board / Discard buttons: confirm routes through the normal
board-review path (PrompterService.confirm_live_draft, route=board) and
PARKS the session — board feedback later streams straight back into the
same thread, closing the redraft loop from the phone. MegaTask batches
still confirm in the panel only.
The consumer's open stream arms the registry's 60s keepalive, so the
bridge runs its own idle TTL (same setting, parked sessions exempt).
State is per-process in-memory by design (the _PENDING_REPLIES posture);
intake/secretary containers are process-wide singletons, so a bridged
session preempts a live panel session of the same kind by construction.
* feat(tg): cockpit skin — RoboCo dark deck with a constant amber accent
The cockpit no longer inherits the dashboard's white default outside
Telegram: #tg-shell carries its own standing skin (deep slate surfaces,
amber primary) so the Mini App looks like RoboCo everywhere. Inside
Telegram the themeParams bridge now overrides SURFACE tokens only —
background/card/text/hint/border repaint to the user's Telegram theme
while --primary/--ring stay RoboCo amber: Telegram's surfaces, RoboCo's
voice. Demo fixtures also rewritten to neutral content (they previously
depicted unbuilt forge work and already-shipped roadmap items as live).
---------
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
RoboCo
AI Agents Company - A virtual organization of 25 AI agents + 1 human CEO, designed to operate as a complete software development workforce.
▶ Watch the 26-min intro what it is, a walkthrough, and how to use it |
▶ Watch the 2.5-hour build session a conversation → a shipped feature |
Watch the full 2:33 walkthrough (.mp4) →
Warning
RoboCo is early-stage, work-in-progress software (v0). It's under active development, runs in a homelab, and will have rough edges, breaking changes, and bugs. It is not production-ready and the API/database schema are not stable yet. Treat it as a working prototype to explore and build on — please don't expose it to the public internet as-is. Issues and PRs very welcome.
Tip
📚 Full documentation: docs.roboco.tech — install & first run, the company model, a page-by-page panel reference, model providers, the optional subsystems, deployment, and the API.
Overview
RoboCo implements a structured organizational hierarchy with formal communication protocols, task management, and quality controls. The system enables a single human (CEO) to orchestrate complex multi-project development at scale.
CEO (You, the human)
│
├── Intake (on-demand interviewer: chats only with you to draft a task)
├── Secretary (on-demand chief-of-staff: reads company state, runs gated directives)
├── PR Reviewer (read-only main reviewer: inbound external/fork + internal PRs, and the root→master in-path gate)
│
└── Board (3 agents)
├── Product Owner
├── Head of Marketing
└── Auditor (silent observer, reports to you)
│
└── Main PM (coordinates all cells)
│
├── Backend Cell (6 agents: 2 Devs, 1 QA, 1 PM, 1 Documenter, 1 PR Reviewer)
├── Frontend Cell (6 agents: 2 Devs, 1 QA, 1 PM, 1 Documenter, 1 PR Reviewer)
└── UX/UI Cell (6 agents: 2 Devs, 1 QA, 1 PM, 1 Documenter, 1 PR Reviewer)
The 25 agents = Intake + Secretary + PR Reviewer + the Board (3) + Main PM + the three 6-agent cells (18). Agents run on Anthropic Claude by default, or on xAI Grok (the official grok CLI on a SuperGrok subscription) — see the provider note under Configuration.
How it works
You hand a task to the company; it runs through a real build → review → document → merge pipeline and comes back to you to approve.
One full loop, put simply:
- You give the Board a task — they review it. The Product Owner and Head of Marketing turn your ask into requirements and acceptance criteria.
- You approve — the Main PM starts the work. A notification asks for your Approve & Start decision; approve, and the Main PM breaks it into per-cell subtasks.
- Each cell's PM delegates, supports, and triages its developers (UX/UI, Frontend, Backend).
- Developers build it, QA verifies and gates it, Documenters keep the books.
- Cell PMs merge their PRs into the Main PM's branch.
- The Main PM opens the final PR and notifies you "It's done!" — you approve and merge, or send it back for rework. (Only you ever merge to
master.)
— Full circle —
See the full walkthrough, with screenshots →
Or watch the full panel walkthrough (video) →
Project Structure
roboco/
├── roboco/ # Main Python package
│ ├── api/ # FastAPI routes & schemas
│ │ ├── routes/ # API endpoints (tasks, git, agents, etc.)
│ │ └── schemas/ # Pydantic request/response models
│ ├── services/ # Business logic services
│ │ ├── task.py # Task lifecycle management
│ │ ├── workspace.py # Multi-agent workspace management
│ │ ├── messaging.py # Agent communication
│ │ └── optimal.py # RAG/Knowledge base (in-house pgvector)
│ ├── models/ # Pydantic domain models
│ ├── db/ # SQLAlchemy ORM & session
│ ├── enforcement/ # Task lifecycle state machine
│ ├── runtime/ # Orchestrator for agent spawning
│ ├── agents/ # Agent base classes
│ ├── mcp/ # MCP server implementations
│ └── config.py # Application configuration
├── agents/
│ └── prompts/ # Agent system prompts (roles, teams, identities)
├── docs/
│ ├── rag/ # Agent knowledge base (indexed into RAG)
│ └── map/ # Exhaustive codebase map (agent-facing)
├── alembic/ # Database migrations
├── CLAUDE.md # Claude Code guidance
├── docker-compose.yml # Full stack, built from source
└── docker-compose.registry.yml # Full stack, pulled from the image registry
Running RoboCo
You need Docker + Docker Compose and a Claude Code auth directory on the host (~/.claude, mounted into the orchestrator so agents can reach the model). Copy .env.example to .env and set at least ROBOCO_ENCRYPTION_KEY and ROBOCO_AGENT_AUTH_SECRET (that file shows how to generate each). However you start it, the whole company is reachable at one origin: http://localhost:3000.
Optional — run agents on xAI Grok instead of Claude. RoboCo can spawn agents on xAI's official grok CLI authenticated by a SuperGrok subscription (no metered API key). Run grok login once on the host and point ROBOCO_HOST_GROK_DIR at the resulting ~/.grok so it mounts into Grok agents; the orchestrator keeps the ~6h token refreshed for you. See the Grok block in .env.example (ROBOCO_HOST_GROK_DIR, ROBOCO_GROK_AGENT_IMAGE, ROBOCO_GROK_CLI_MODEL, ROBOCO_GROK_REASONING_EFFORT).
Option 1 — Run the pre-built images (quickest)
Every release publishes all RoboCo images to both the GitHub Container Registry and Docker Hub, so you can run the full stack without building anything. Use the registry compose:
git clone https://github.com/rennf93/roboco.git && cd roboco
cp .env.example .env # then edit in your secrets
docker compose -f docker-compose.registry.yml pull
docker compose -f docker-compose.registry.yml up -d
Choose the registry and version with two env vars (defaults shown):
ROBOCO_REGISTRY=ghcr.io/rennf93 # or docker.io/renzof93
ROBOCO_VERSION=latest # or a pinned release, e.g. 0.15.0
The orchestrator spawns the matching pre-built agent images on demand — no build toolchain or source compile on your host.
Option 2 — Build from source
The same full stack, built locally from the Dockerfiles instead of pulled:
git clone https://github.com/rennf93/roboco.git && cd roboco
cp .env.example .env # then edit in your secrets
docker compose up -d # builds images on first run, then starts everything
Option 3 — Local development (no full stack)
For hacking on the code itself, run only the backing services in Docker and the API on your host. RoboCo's own code requires Python 3.13+ (uv will fetch it if needed):
uv sync
docker compose up -d postgres redis ollama # backing services only
uv run alembic upgrade head # migrate the database
uv run python -m roboco.cli # API + orchestrator
# Or just the API without the orchestrator:
uv run uvicorn roboco.api.app:app --reload --host 0.0.0.0 --port 8000
Configuration
Key environment variables (see roboco/config.py for all options):
# API Server
ROBOCO_HOST=0.0.0.0
ROBOCO_PORT=8000
# Database
ROBOCO_DATABASE_HOST=localhost
ROBOCO_DATABASE_PORT=5432
ROBOCO_DATABASE_NAME=roboco
# Workspaces (Multi-Agent Git)
ROBOCO_WORKSPACES_ROOT=/data/workspaces
ROBOCO_WORKSPACE_AUTO_CLONE=true
# RAG/LLM
ROBOCO_LOCAL_LLM_BASE_URL=http://roboco-ollama:11434/v1
ROBOCO_LOCAL_LLM_MODEL=glm-5.2:cloud
# Feature flags (default-off unless noted; toggle from Settings → Feature Flags)
ROBOCO_CONVENTIONS_ENABLED=false # per-project architectural conventions standard
ROBOCO_TOOLCHAIN_MATCH_ENABLED=false # build each target project under its own Python
ROBOCO_OVERLOAD_BREAK_ENABLED=true # park a provider on a persistent model-API overload
ROBOCO_DOCS_SYNC_ENABLED=false # docs-divergence sync (release → docs-update task). Default-off; when on, a successful release publish originates one bounded, deduped docs-update task against the roboco-website project.
ROBOCO_DOCS_SYNC_MAX_OPEN_TASKS=3 # rolling cap on concurrently-open docs-sync tasks
ROBOCO_DOCS_SYNC_MAX_PER_CYCLE=1 # max docs-sync tasks originated per publish invocation
# Auditor scheduled sweeps (default 6 hours; 0 disables)
ROBOCO_AUDIT_INTERVAL_SECONDS=21600
Multi-Agent Workspace Structure
Each agent gets their own git clone for parallel development:
{ROBOCO_WORKSPACES_ROOT}/
└── {project-slug}/
└── {team}/
└── {agent-slug}/
└── [git repository]
Example:
/data/workspaces/roboco/backend/be-dev-1/
/data/workspaces/roboco/backend/be-dev-2/
Task Lifecycle
backlog → pending → claimed → in_progress → verifying → awaiting_qa
↓ ↓ ↓ ↓
cancelled blocked needs_revision awaiting_documentation
paused ↓
awaiting_pm_review
↓
awaiting_ceo_approval
↓
completed
Assembled, PR-bearing tasks pass through one extra stage — the in-path PR-review gate — before the PM merges:
in_progress → awaiting_pr_review → awaiting_pm_review
(submit_up / (pr_pass)
submit_root) (pr_fail → needs_revision)
The cell PM's submit_up (cell→root PR) and the Main PM's submit_root (root→master PR) open the assembled PR and enter the gate; a PR reviewer pr_passes it on to the PM merge or pr_fails it back. Leaf dev tasks (reviewed by QA) and branchless coordination roots skip the gate.
API Endpoints
Domain routes are mounted under /api:
| Route Group | Description |
|---|---|
/api/tasks |
Task CRUD, lifecycle, claiming |
/api/agents |
Agent management |
/api/git |
Git operations (status, commit, push, PR) |
/api/sessions |
Communication sessions |
/api/messages |
Agent messages |
/api/projects |
Project (repo) management |
/api/work-sessions |
Git work session tracking |
/api/optimal |
RAG/Knowledge base queries |
/api/journals |
Agent journals/reflections |
/api/orchestrator/status |
Orchestrator / dispatcher status |
The agent gateway verbs are served separately under /api/v1/flow/{role}/{verb} (intent verbs) and /api/v1/do (content tools) — see the Agent Gateway.
Development
# Install dev dependencies
uv sync --all-extras
# Run tests
uv run pytest
# Format and lint
uv run ruff format .
uv run ruff check .
uv run mypy roboco/
# Type checking
uv run mypy roboco/
Core Principles
- Everything is a task - All work is tracked and documented
- No work without a task - Create task record first
- No task without acceptance criteria - How do we know it's done?
- No closure without documentation - Future agents need context
- Communication is constant - Stream reasoning, log everything
- The Auditor sees all - Quality monitored silently
- CEO approves major changes - Human-in-the-loop for critical decisions
Technology Stack
| Layer | Technology |
|---|---|
| API Framework | FastAPI |
| Database | PostgreSQL + SQLAlchemy (async) |
| Vector Store | PostgreSQL + pgvector (in-house engine) |
| Cache/Queue | Redis |
| RAG Engine | in-house (asyncpg + pgvector, hybrid retrieval) |
| Embeddings | qwen3-embedding:0.6b (Ollama) |
| Local LLM | Ollama (glm-5.2:cloud) |
| Cloud LLM | Claude API (Anthropic) + xAI Grok (official grok CLI, SuperGrok subscription) |
| Package Manager | uv |
Status
Core Infrastructure (Complete)
- Data models (Pydantic)
- Database ORM (SQLAlchemy async)
- Task lifecycle state machine
- Multi-agent workspace management
- Agent prompts (25 agents)
- Messaging API
- Task API with full lifecycle
- Git operations API
- RAG/Knowledge base (in-house pgvector engine)
- Agent orchestrator
- CEO approval workflow
- Pluggable agent providers (Claude Code + xAI Grok on the official
grokCLI) - Inbound PR review (read-only PR-reviewer + CEO supersede/dismiss queue)
- Self-healing CI loop for RoboCo's own repo (default-off, CEO-gated)
- Business Goals tab with a live Company Scorecard (delivery, spend-vs-budget, lead time)
In Progress
- Frontend panel (vendored under
panel/, served through nginx on :3000) - Full agent autonomy testing
Security
Important
Do not expose RoboCo to the public internet as-is. It is designed to run on a trusted private network (homelab / LAN).
Agent authentication. Requests identify the caller with X-Agent-Id / X-Agent-Role headers. The orchestrator issues each spawned agent an HMAC token (X-Agent-Token, signed with ROBOCO_AGENT_AUTH_SECRET) that binds its id, role and team. Token enforcement is gated by ROBOCO_AGENT_AUTH_REQUIRED:
ROBOCO_AGENT_AUTH_REQUIREDunset/false (default): header-trust mode — the role headers are accepted without a token, so any client that can reach the API may claim any role (includingceo). The API logs a warning at startup in this mode. Acceptable only on a trusted network.ROBOCO_AGENT_AUTH_REQUIRED=true: every request must carry a valid token; an agent cannot spoof another agent's role. The control panel keeps working because nginx — the only trusted hop between the browser and the API — injects the CEO token (X-Agent-Token) on/apiand/ws, so the browser never holds the signing secret. Generate that token withmake panel-tokenand set it asROBOCO_PANEL_AGENT_TOKENin.envbefore enabling secure mode.
WebSocket streams. Token enforcement is currently REST-only. The /ws/* endpoints authenticate by agent_id query param at most and do not yet validate X-Agent-Token, even in secure mode — nginx injects the token so the panel works, but a direct WebSocket connection that bypasses nginx is not rejected. In particular the operator stream /ws/system (rate-limit lifecycle + token-usage snapshots for the dashboard) is unauthenticated. These streams are read-only — no control surface, secrets, or task content — but treat the orchestrator port as trusted-network-only until WebSocket auth lands.
Secrets (the Fernet ROBOCO_ENCRYPTION_KEY, GitHub PATs) live encrypted in the database and in gitignored env files — never in the repo. Per-project git tokens are Fernet-encrypted at rest and never returned by the API.
License
Copyright (c) 2026 Renzo Franceschini
RoboCo is licensed under the GNU Affero General Public License v3.0 (AGPL-3.0). See LICENSE for the full text.
The AGPL's network-use clause (section 13) means that if you run a modified version of RoboCo as a network service, you must make your modified source available to its users. This keeps the project open while preventing closed, hosted re-distributions.
Contributing
Contributions are welcome. All contributors must sign the Contributor License Agreement (CLA.md) — this is automated on your first pull request. See CONTRIBUTING.md for the workflow and why the CLA exists.