Commit Graph
571 Commits
Author SHA1 Message Date
Renn F 2089b9e765 feat(grok): cost-ceiling kill-switch (budget-guardrail parity)
Claude Code's per-agent token-budget hook fires against the SDK :9000 server;
opencode exposes NO usage/budget hook to a plugin (confirmed against its plugin
docs), so the budget kill-switch can't be a plugin/sidecar — the orchestrator
enforces it instead.

_enforce_grok_cost_budget runs each dispatch tick: for every ACTIVE GROK
container it reads cumulative cost from the opencode store (the Phase-2 reader)
and kills + evicts it past ROBOCO_GROK_MAX_COST_USD (0 = off), after which the
reaper releases the freed task. This also catches a runaway loop that keeps
firing verbs (so it evades the idle watchdog) but still burns cost.

Covers the budget/runaway-burn slice of guardrail parity. The remaining Claude
hooks (prompt-injection PRE-gate, stop-guard terminal-verb) have no blocking
opencode equivalent — opencode's message/stop hooks are observe-only — and the
interactive reasoning-variant has no opencode.json/serve knob (CLI-flag only);
both are pinned for a live probe rather than shipped as a guess.
2026-06-18 12:48:40 +02:00
Renn F 7681c47370 test(grok): mypy-clean the reaper watchdog + interactive spawn tests
The CI mypy scope (roboco/ tests/) flagged test-only typing issues my per-file
runs missed: direct method assignment (orch._remove_container = AsyncMock())
trips [method-assign], and a module-level dict[str,str] is invariant against
the dict[str, str|None] the run-spec expects.

- Use monkeypatch.setattr for _remove_container in the watchdog tests.
- Annotate the shared _HOSTS as dict[str, str | None].

Production code unchanged; mypy roboco/ tests/ is green.
2026-06-18 12:38:23 +02:00
Renn F f1981ca163 feat(grok): surface intake/secretary in the mix-mode picker; doc guardrail parity
- Panel: add intake-1 (prompter) and secretary-1 to the mix-mode per-agent
  routing list so an operator can assign Grok (or Claude) to the interactive
  roles from the UI; assigning a Grok model routes them to the opencode-serve
  image. tsc + eslint clean.
- opencode_config: correct the now-stale parity note — bash-guard is ported
  (secret-scrub.js) and usage/cost is captured (opencode store); the remaining
  gap is the budget/loop/stop/prompt-injection hooks, which need a sidecar
  plugin (open decision), with ROBOCO_GROK_BASH_PERMISSION as the interim gate.
2026-06-18 12:35:51 +02:00
Renn F 0d76ebf586 feat(grok): route interactive intake/secretary to opencode-serve images
Wire the GROK interactive path the in-place way (matching how the interactive
roles already choose ANTHROPIC_* per route), so a GROK route launches the
Grok-native opencode-serve image instead of the Claude SDK-driver image:

- _spawn_intake_container / _spawn_secretary_container pick the
  grok-prompter / grok-secretary image (ensuring the base→grok→interactive
  build chain) when the route is GROK, and stamp provider_type on the spec +
  AgentConfig so finalize routes usage to the opencode store.
- _build_intake_run_cmd / _build_secretary_run_cmd inject OPENAI_* + the
  opencode store mount + system-prompt env for GROK via a shared
  _append_interactive_provider_env, keeping ANTHROPIC_* for every other
  provider. The intake's minimal mounts (no gateway MCP) are preserved, so
  Grok intake matches the Claude intake's tool surface (the spec).
- Add a per-agent opencode store mount to the interactive host paths so
  interactive Grok usage/cost is captured like the one-shot path.

Removes the interim Phase-0 routing guard (the real path supersedes it) and
retires the unused AgentProvider.spawn_interactive/InteractiveSpawnSpec seam —
the interactive roles have a bespoke assembly that the one-shot provider
surface doesn't fit, so the fork lives in their own builders.

UNVERIFIED-LIVE: end-to-end intake/secretary chat on Grok needs the stack up +
opencode serve confirmed against grok-build-0.1.
2026-06-18 12:32:41 +02:00
Renn F 3753fbccc4 feat(grok): Grok-native interactive runtime (opencode serve) — container side
Builds the Grok analogue of the Claude intake/secretary live-session runtime,
satisfying the same IntakeSession seam so the existing IntakeDriver loop,
message source, relay, and StreamChunk panel contract are reused unchanged:

- OpencodeServeSession: a held-open `opencode serve` session (context persists
  across turns) where each human turn is one synchronous POST /session/:id/
  message; normalize_opencode_message maps the reply parts to text/thinking/
  tool_use/draft/turn_end chunks (draft via a propose_draft tool part or the
  fenced roboco-draft fallback). Doc-verified against opencode's server API.
- grok_intake_main / grok_secretary_main: container entrypoints mirroring the
  Claude mains but yielding an OpencodeServeSession; they render opencode.json
  (xAI provider + MCP + system prompt) first, then run the receiver + driver.
- roboco-agent-grok-prompter / -secretary images (FROM roboco-agent-grok) +
  their builder services in all three compose files.

UNVERIFIED-LIVE: the opencode serve flow + exact Part schema + draft path need
a live run against grok-build-0.1 (the part mapping is defensive). The
orchestrator wiring that routes a GROK intake/secretary route to these images
is the next step (a design decision is open — see the handoff notes).
2026-06-18 12:13:08 +02:00
Renn F 6f60d180e9 feat(grok): make interactive spawns first-class on AgentProvider (additive)
The AgentProvider ABC modelled only the one-shot lifecycle (spawn/stop/
health_check/remove), so the interactive intake/secretary roles could never
route through a provider. Add an opt-in interactive surface:

- supports_interactive class flag (default False).
- InteractiveSpawnSpec: the resolved AgentConfig + session id + role-specific
  image + optional HMAC token — everything a provider needs without importing
  orchestrator internals.
- spawn_interactive(spec): a non-abstract default that declines via
  ProviderError, so every existing one-shot provider is unchanged.

Pure scaffolding — no provider opts in yet (GrokProvider flips the flag when
its interactive driver lands). Zero behavioural change.
2026-06-18 11:57:42 +02:00
Renn F e36549f01f feat(grok): capture one-shot Grok usage/cost from the opencode store
A GROK agent runs opencode, not Claude Code: it has no SDK /usage/status
server and writes no Claude transcript, so _resolve_final_token_usage found
nothing and every Grok agent finalized at 0 tokens / $0 — the opencode_usage
reader existed but had no caller.

- Mount a per-agent opencode data dir ($DATA/opencode/<agent_id> →
  /home/agent/.local/share/opencode) so opencode.db is captured, and mount the
  same host dir into the orchestrator (/data/opencode) in all three compose
  files so the finalizer can read it back — the opencode analogue of the
  mounted Claude transcript.
- _resolve_final_token_usage branches on provider_type: GROK reads opencode.db
  via opencode_usage (reasoning folded into output, billed at the output rate)
  and skips the SDK/transcript path. A 0-token read logs a WARNING so a silent
  mount failure isn't mistaken for a real zero-cost run.
- ROBOCO_OPENCODE_DATA_DIR overrides the in-orchestrator path for local runs.
2026-06-18 11:54:58 +02:00
Renn F 7e077fe642 feat(grok): guard interactive roles from GROK routes (interim)
intake (prompter) and secretary run a held-open chat session driven by the
Claude Agent SDK. GROK has no interactive runtime yet, and a GROK route for
those slugs would be spawned with the route creds injected as ANTHROPIC_*
against api.x.ai/v1 — the wrong protocol — producing a silent, empty reply
(the blank intake we observed).

Downgrade a GROK route for intake-1/secretary-1 to the Anthropic default with
a logged warning. The one-shot delivery roles route to GROK unchanged. This
guard is replaced by the real interactive fork once the opencode interactive
driver lands.
2026-06-18 11:46:47 +02:00
Renn F b61148cdc9 feat(grok): reaper watchdog kills wedged opencode containers
The heartbeat reaper deliberately skips a task whose assignee holds a live
ACTIVE container, so a Claude agent deep in a long edit/test cycle isn't
churned out from under live work. A wedged opencode container breaks that
assumption: it stays ACTIVE while firing no gateway verb, so its heartbeat
never advances and the live-instance skip would shield its task forever — the
exact way the Grok pr_reviewer parked in_progress.

Add a longer grok-idle kill threshold (ROBOCO_GROK_IDLE_KILL_SECONDS, default
900s, well past the stream chunk timeout). A GROK instance idle past it is
force-removed (its logs dumped to disk first) and evicted from the instance
registry, so the same reaper pass then releases the task. Only GROK runtimes
are eligible — a quiet Claude agent keeps the heartbeat-skip protection.
2026-06-18 11:37:40 +02:00
Renn F ff8867dd1f fix(grok): stop opencode subagent-stream hang at the config layer
The Grok pr_reviewer wedged in_progress forever: opencode's default agent ran
with the subagent `task` tool enabled, spawned an Explore subagent on
grok-build-0.1 whose model call opened an SSE stream that went idle, and the
run hung with no timeout.

- Hard-disable opencode's subagent `task` tool in the generated opencode.json.
  No RoboCo role uses opencode-internal subagents — work flows through the
  gateway verbs — so removing the tool kills the hang trigger outright.
- Set provider.xai.options.timeout + chunkTimeout (operator-tunable via
  ROBOCO_GROK_REQUEST_TIMEOUT_MS / ROBOCO_GROK_CHUNK_TIMEOUT_MS) as the
  defence-in-depth backstop; chunkTimeout aborts an idle stream.
- Bundle the permission + timeout + subagent knobs into an OpencodeGuards
  dataclass (keeps the builder under the arg-count gate).
- Drop the dead ROBOCO_AGENT_TOOLS spawn env (it had no consumer); opencode
  tool restriction lives in the rendered config now.
2026-06-18 11:31:13 +02:00
Renn F 4720e4129d style(panel): show the Grok (xAI) key card above the Ollama card 2026-06-18 10:35:04 +02:00
Renn F 5725ec998b feat(grok): reasoning-effort by role (cut grok-build cost on cheap roles)
grok-build-0.1 reasons heavily by default and reasoning bills at the output
rate (a live "say ok" call emitted ~300 reasoning tokens, ~85% of its cost).
Confirmed live that opencode's `--variant minimal` cuts reasoning ~54%
(298 -> 136 tokens, same prompt).

GrokProvider now picks reasoning effort by role: code-quality roles (developer,
qa, pr_reviewer) keep full reasoning; coordination / docs / board roles
(cell_pm, main_pm, documenter, product_owner, head_marketing, auditor, prompter,
secretary) run "minimal". It's passed to opencode via the entrypoint's
`--variant`. Operators can force one effort for ALL grok agents with the
ROBOCO_GROK_REASONING_EFFORT env (minimal | high | max, or default/full).

Tests cover the role map, the env override, and the spawn env wiring.
2026-06-18 10:26:29 +02:00
Renn F 0e9bbc15db feat(grok): first-class xAI/Grok routing mode (UI + backend)
The Routing-mode toggle had Anthropic / Ollama / Self-Hosted / Mix but no way
to route the whole org to Grok. Add it end to end:

- backend: apply_mode("grok") + _apply_grok (GLOBAL default -> grok-build-0.1) +
  derive_mode "grok" detection; ApplyModeRequest/ModeResponse accept "grok".
- panel: a "Grok" routing-mode card (between Anthropic and Ollama, gated on the
  xAI key) + flipToGrok; a Grok group in the per-agent mix dropdown +
  catalogGrokOnly + a grok ProviderBadge variant; the mix-save key check and
  the AI-routing description now cover Grok.
- tests: integration derive_mode/apply_mode "grok" cases (+ grok provider row
  in the fixture).

Gated: ruff + mypy clean; panel typecheck + lint clean.
2026-06-18 10:13:14 +02:00
Renn F b0857915a1 fix(grok): correct opencode provider (Responses API), stdin, reasoning cost
A live opencode run against api.x.ai/v1 surfaced three real bugs:

1. Provider package — grok-build-0.1 is driven via the OpenAI Responses API
   (opencode calls model.responses()). @ai-sdk/openai-compatible is
   chat/completions only and errors "responses is not a function". Switch the
   generated opencode.json provider + the grok image to @ai-sdk/openai.
2. Headless hang — `opencode run` blocks after init without a TTY; close stdin
   (`< /dev/null`) in the entrypoint so it proceeds to the model call.
3. Reasoning-token cost — grok-build-0.1 is a reasoning model; reasoning tokens
   bill as output but opencode stores them in a separate column. cost_for_session
   folds tokens_reasoning into output (else ~22x undercount).

Verified end-to-end against a real session row (input=6120, output=1,
reasoning=226, cache_read=1856): our pricing reproduces opencode's stored USD
cost ($0.0069452) exactly. Tests anchored to that real row.
2026-06-18 09:57:20 +02:00
Renn F af6cad98d8 feat(grok): read opencode session usage for cost capture
Confirmed by inspecting a local opencode run: opencode persists per-session
usage in SQLite at ~/.local/share/opencode/opencode.db — the `session` table
carries cost + tokens_input/output/reasoning/cache_read/cache_write. xAI's
response usage object (prompt_tokens, completion_tokens,
prompt_tokens_details.cached_tokens, completion_tokens_details.reasoning_tokens)
maps directly onto those columns.

Add opencode_usage.read_session_usage / cost_for_session: read the opencode DB
and price the tokens via roboco.billing.pricing (our cost stays authoritative;
opencode's own `cost` column is kept for reference). Tested against a fixture DB
mirroring the real schema (single session, summed sessions, missing/empty DB).

Remaining wiring (for the live spawn): mount the opencode data dir on grok
spawn + call cost_for_session at reap to record the usage rollup.
2026-06-18 09:04:49 +02:00
Renn F 07f55117de feat(grok): price grok-build-0.1 + secret-scrub opencode plugin
- pricing.py: add grok-build-0.1 rates ($1/1M input, $0.20 cached, $2/1M
  output), verified against xAI's published pricing. Grok is a priced
  non-Anthropic model, so cost computes the moment usage is captured.
- secret-scrub.js: an opencode tool.execute.before plugin porting the
  security-critical bash-guard deny rules (git network ops, credential-file
  reads, /proc env, internal-host HTTP, roboco.* imports, ROBOCO_AGENT_ID
  forgery, env dumps, destructive rm) to the opencode runtime — restoring the
  guard the Claude Code hook can't provide there. Throwing denies the call
  (confirmed by opencode's env-protection example). Wired into the generated
  opencode.json plugin array + baked into the grok image.

Deny logic verified via node (9 deny + 5 allow cases). UNVALIDATED against a
live opencode runtime: confirm it fires in the live E2E spawn before a Grok
dev-agent touches a real repo; the bash permission is operator-tunable as a
second gate.

Cost CAPTURE (distinct from pricing) is intentionally NOT built yet: opencode's
plugin hooks expose model info but no token/usage object, so the capture path
is unconfirmed and needs the live spawn to settle.
2026-06-18 08:36:36 +02:00
Renn F a9b7b7ead5 fix(migration): commit the grok enum value before seeding (autocommit_block)
CI's "Apply database migrations" failed with asyncpg
UnsafeNewEnumValueUsageError: alembic runs the whole upgrade in a single
transaction, so migration 039's INSERT used 'grok' in the same transaction
that 038 added it — which Postgres forbids. Splitting into two migration
files did not help (one transaction spans both). Wrap the ALTER TYPE ADD
VALUE in op.get_context().autocommit_block() so the value commits before 039
(and any later migration) uses it. Still renders in offline --sql, so the
enum-migration-parity test is unaffected.
2026-06-18 07:30:09 +02:00
Renn F 5e364f6fdc ci(release): build + publish the roboco-agent-grok image
Add roboco-agent-grok to the release workflow's image build/publish map so
registry deploys carry the Grok runtime image (parity with every other
agent image). Split from the feature commit because pushing a workflow
change requires a workflow-scoped token.
2026-06-18 07:13:17 +02:00
Renn F 085414dcf5 feat(providers): native Grok runtime — opencode image, config gen, panel key
Complete the native Grok (xAI) path so grok-build-0.1 runs as a real
RoboCo agent, not just the provider seam.

- roboco-agent-grok image (docker/agent-grok.Dockerfile): FROM agent-base
  + opencode (the OpenAI-protocol runtime). One image serves every role;
  role behaviour comes from the mounted manifest / mcp-config / system
  prompt, exactly as on the Claude path.
- Entrypoint renders opencode.json at spawn from the GrokProvider env
  contract + the mounted Claude Code mcp-config.json
  (roboco.llm.providers.opencode_config): translates RoboCo's gateway
  servers (roboco-flow / roboco-do / ...) into opencode's mcp block,
  declares the xAI OpenAI-compatible provider + model, and wires
  permissions + instructions. Pure, unit-tested translation.
- Orchestrator registers GrokProvider with the registry-qualified image
  (_qualify_agent_image) so it resolves in local and registry deploys.
- Compose (both files + the registry compose) gain an agent-grok-image
  builder service.
- Panel: a Grok (xAI) API key card on the AI Providers page, plus the
  grok ModelProvider value.

KNOWN PARITY GAP (opencode runtime): the bash-guard PAT-scrub and the
transcript-based usage/cost capture are Claude Code hooks and do not
transfer to opencode. bash permission is operator-tunable
(ROBOCO_GROK_BASH_PERMISSION) so a deployment can fail closed until a
security/usage-parity opencode plugin lands. That plugin and live E2E
validation are the remaining work to finalize with xAI.
2026-06-18 07:12:57 +02:00
Renn F a956083f9f feat(providers): pluggable agent providers + Grok (xAI) backend
Add a roboco/llm/providers/ seam — an AgentProvider lifecycle ABC and a
ProviderRegistry keyed by ModelProvider — so the orchestrator can drive
agent backends other than Claude Code.

The first non-Claude backend is GrokProvider for xAI's grok-build-0.1.
xAI is OpenAI-compatible only (no Anthropic-Messages endpoint), so a Grok
agent runs an OpenAI-protocol runtime pointed at https://api.x.ai/v1
rather than the ANTHROPIC_BASE_URL injection the other providers use. It
reuses the orchestrator's existing mount/auth assembly, so it inherits the
same MCP gateway + tool-manifest wiring as every other agent by
construction, and passes its prompt via env (never an argv positional).

The change is purely additive: only GROK routes through the registry;
Anthropic / Ollama Cloud / self-hosted spawns run the existing
_spawn_container path unchanged.

Includes:
- ModelProvider.GROK (migration 038) + a seeded Grok provider row
  (migration 039) + a grok-build-0.1 catalog entry
- GET/PUT /api/providers/grok-key to store the xAI key (Fernet-encrypted,
  reusing the existing provider-key machinery)
- ClaudeCodeProvider reference adapter over the current spawn
- unit tests for the registry, GrokProvider (gateway wiring, no
  ANTHROPIC_* leak, prompt-injection safety, failure paths) and routing

The dedicated roboco-agent-grok image and the exact OpenAI-protocol CLI
invocation are the remaining piece to finalise with xAI.
2026-06-18 06:14:27 +02:00
Renn F 7209edefbe fix(panel): regenerate pnpm-lock.yaml for added vitest test deps
The Company Scorecard work (#212) added 6 test dependencies to
panel/package.json (vitest, @vitest/coverage-v8, jsdom, and three
@testing-library packages) without updating the lockfile, so every
--frozen-lockfile install (Docker panel-builder stage and CI) failed
with ERR_PNPM_OUTDATED_LOCKFILE. Regenerate the lockfile so it matches
package.json.
2026-06-18 03:56:18 +02:00
6007f47fc9 [ef7b7cb9] Add Company Scorecard to Business Goals tab (#212)
* [0c7a4732] feat(cockpit): add completed_30d and median_lead_time_hours to delivery summary (#207) (#210)

- Extend DeliverySummary schema with completed_30d: int = 0 and
  median_lead_time_hours: float | None = None fields
- Add TaskService.get_delivery_stats_30d() that queries tasks completed
  in the last 30 days and computes statistics.median of lead times
- Update CockpitService.summary() to source both new keys from
  get_delivery_stats_30d() and include them in the delivery dict
- Update tests: mock new method in _patch(), assert new fields in
  test_summary_aggregates, fix test_route_ok_for_ceo dict, add three
  new unit tests for get_delivery_stats_30d (empty, multi, single)

Co-authored-by: Backend Developer 1 <be-dev-1@agents.roboco.dev>

* [d2647edf] Frontend: Build CompanyScorecard card on Goals tab (#211)

* [12569f37] Extend CockpitSummary type and build CompanyScorecardCard component (#208)

* [12569f37] feat(cockpit): extend CockpitSummary type with completed_30d and median_lead_time_hours

Add optional delivery.completed_30d (number) and top-level
median_lead_time_hours (number | null, optional) to CockpitSummary
interface in panel/src/lib/api/cockpit.ts so the API shape captures
the new backend fields without breaking existing consumers.

* [12569f37] feat(business): add CompanyScorecardCard component

Create panel/src/components/business/company-scorecard-card.tsx
exporting CompanyScorecardCard. The card fetches /cockpit/summary
via useQuery and renders five always-visible sections:

- Delivery: in_flight, blocked, awaiting_ceo, completed_30d tiles
  (all from API response; no hardcoded numbers)
- Spend: 30d spend + projected monthly; muted 'No budget cap set'
  when cap is null; red/destructive styling only when cap is a
  non-null number AND over_budget is true
- Speed: 'X.Xh median — target: < 24h' when value present;
  'No data yet' when null/undefined; '0h' never rendered
- Two stub Objectives with 'Not tracked yet' label, muted text,
  and dashed-border styling — no fabricated numeric values
- Loading: three grouped Skeleton blocks
- Error: OfflineState with title 'Could not load scorecard data'

---------

Co-authored-by: Frontend Developer 1 <fe-dev-1@agents.roboco.dev>

* [f1f5cded] Integrate CompanyScorecardCard into GoalsTab and pass quality gate (#209)

* [f1f5cded] feat(cockpit): extend CockpitSummary with completed_30d and median_lead_time_hours

Add optional delivery.completed_30d (number) and top-level
median_lead_time_hours (number | null, optional) to CockpitSummary
interface in panel/src/lib/api/cockpit.ts. Backward compatible.

* [f1f5cded] feat(business): add CompanyScorecardCard component

Create panel/src/components/business/company-scorecard-card.tsx
exporting CompanyScorecardCard. Fetches /cockpit/summary via
useQuery and renders five always-visible sections: Delivery (no
hardcoded numbers), Spend (muted 'No budget cap set' when null;
red only when cap set AND over_budget true), Speed (X.Xh median
or 'No data yet'), two stub Objectives with dashed border and
'Not tracked yet' label. Loading: three skeleton groups. Error:
OfflineState 'Could not load scorecard data'.

* [f1f5cded] feat(goals-tab): integrate CompanyScorecardCard into GoalsTab

Import and render CompanyScorecardCard below the charter form in
goals-tab.tsx. The scorecard fetches its own data independently
so all loading/error states are handled per-card. Both cards are
always rendered in the Goals tab.

* [f1f5cded] fix(scorecard-tests): add vitest framework and CompanyScorecardCard test suite

Install vitest + @testing-library/react + @testing-library/jest-dom +
jsdom + @vitest/coverage-v8 as devDependencies in panel/.

Add panel/vitest.config.ts (jsdom env, @/* alias, coverage on
company-scorecard-card.tsx with 80% threshold).

Add panel/src/test/setup.ts (jest-dom matchers).

Update panel/package.json: add test, test:watch, typecheck scripts.

Update panel/eslint.config.mjs: ignore coverage/ directory to keep
lint clean of generated files.

Write panel/src/components/business/__tests__/company-scorecard-card.test.tsx
with 8 tests covering all 7 AC2 scenarios:
- loading skeleton rendered
- OfflineState on error
- OfflineState when data undefined
- delivery counts from mock data
- spend 'No budget cap set' when cap null
- spend destructive styling when cap non-null and over_budget true
- speed 'No data yet' when lead time null
- speed formatted value when lead time present

pnpm lint: 0 errors  pnpm typecheck: 0 errors
pnpm test: 8/8 pass  coverage: stmts 95% branches 90% fns 91% lines 95%

---------

Co-authored-by: Frontend Developer 1 <fe-dev-1@agents.roboco.dev>

---------

Co-authored-by: Frontend Developer 1 <fe-dev-1@agents.roboco.dev>

---------

Co-authored-by: Backend Developer 1 <be-dev-1@agents.roboco.dev>
Co-authored-by: Frontend Developer 1 <fe-dev-1@agents.roboco.dev>
2026-06-18 03:33:47 +02:00
Renn F 32b6d72933 fix(test): get-or-create seed agents in self-heal DB tests (unbreak CI)
The self-heal origination tests blind-inserted the fixed-uuid foundation
system agent (and a main-pm agent). In the full CI suite the app lifespan
seeds + commits those agents first, so the insert hit a duplicate-key on
pk_agents — green in isolation, red in CI. Get-or-create both (by id / by
slug) so the tests pass whether or not the agents already exist. Verified
against the real ordering (an app-lifespan integration test before this
file): 187 passed.
2026-06-17 22:57:33 +02:00
Renn F 12b8625462 chore(deploy): sync self-heal env into the registry compose too
All three compose files must carry the same app env; the registry one had
external-PR but was missing the self-heal block. Add it (identical to the build
compose) so docker-compose.yml / .yaml / .registry.yml agree — the registry
file still differs only in image source (registry pulls) and host-env paths.
2026-06-17 22:39:01 +02:00
Renn F e1cbf00474 chore(deploy): commit docker-compose.yaml so the file the NAS uses matches
Both compose files are git-TRACKED, and Docker picks docker-compose.yaml over
.yml — so the NAS reads the .yaml. The self-heal env (and prior changes) had
only been committed to the .yml, leaving the tracked .yaml stale in the repo.
Commit the .yaml so both carry the same config. Going forward: commit BOTH.
2026-06-17 22:32:42 +02:00
Renn F 3960e98da5 chore(deploy): default self-heal target to the roboco-api project
It's a monorepo registered as three cell-projects (BE/FE/UXUI) sharing one
git_url, so the CI signal is identical whichever is named; point self-heal at
roboco-api. Overridable via ROBOCO_SELF_HEAL_PROJECT_SLUG. Mirrored into the
untracked .yaml.
2026-06-17 22:15:29 +02:00
Renn F c330e5514c chore(deploy): wire self-heal env into the compose (default-off, ci.yml-scoped)
Plumb the self-healing knobs into the orchestrator service so a NAS deploy is
configured without hand-editing env: both toggles default OFF (armed from
Settings -> Feature Flags), the CI signal is scoped to ci.yml (RoboCo has
several workflows, so the unscoped "latest completed run" would be
unreliable), and PROJECT_SLUG is left for the deployment to set (which
registered project IS RoboCo). Mirrored byte-identically into the untracked
docker-compose.yaml the NAS uses.
2026-06-17 22:08:20 +02:00
Renn F 33fa21d00a feat(self-heal): scope CI signal to a workflow + warn on missing target
Two hardening fixes from the gap review:
- Optional self_heal_ci_workflow scopes the CI signal to one workflow file
  (the workflow-scoped Actions endpoint). Without it, "latest completed run
  across all workflows" could miss a red CI run masked by a later passing
  workflow, or false-trigger on a non-CI workflow — unreliable on a
  multi-workflow repo.
- The loop logs a warning when self-heal is armed but self_heal_project_slug
  is unset, so a misconfiguration isn't mistaken for "all green".

Tests cover the workflow-scoped endpoint.
2026-06-17 21:58:48 +02:00
Renn F 7ef7d8414e fix(self-heal): hold an unconfirmed fix task out of dispatch until CEO approval
Adversarial review found the load-bearing invariant broken at the dispatch
layer: _dispatch_pm_work skipped only PR_REVIEW_SOURCES, so a PENDING
team=main_pm self_heal task (assigned_to=None, confirmed_by_human=False) was
routed to Main PM and spawned BEFORE the CEO approved it — the "never start
until you approve" promise didn't hold.

Fix: the PM dispatcher now also skips source='self_heal' while
confirmed_by_human is False (before the assigned/unassigned split, so it holds
either way); the task still shows in the panel so the CEO can see and approve
it. approve_and_start flips confirmed_by_human=True (the CEO's start IS the
human confirmation), so it dispatches normally afterward. Other sources are
unaffected.

Tests: a unit test that the dispatcher holds an unconfirmed self_heal task but
routes a confirmed one and ordinary tasks, plus a DB test that approve_and_start
flips the gate. (The readiness gate was deliberately not used — a blocker there
marks the task `blocked`; the dispatch skip leaves it cleanly PENDING.)
2026-06-17 21:55:33 +02:00
Renn F 49c7b3c42a test(self-heal): httpx-mock coverage for get_latest_ci_conclusion
The CI telemetry call is the feature's only real-world I/O and was previously
exercised only through a fake source. Cover the GitHub Actions request shape
(/actions/runs, branch/status/per_page, auth) and response parsing, plus the
safe-None paths (missing token, GitHub error, no runs).
2026-06-17 21:33:26 +02:00
Renn F fc1d6b2cc6 feat(self-heal): expose the self-heal toggles in the Feature Flags panel
Add self_heal_enabled (detect + notify) and self_heal_originate_enabled (also
open fix tasks) to the panel feature-flags registry + card, so the loop can be
armed/disarmed from Settings instead of editing env. Effect is on the next
restart, like the other flags; the self_heal_project_slug target (which repo is
RoboCo) stays a deployment env setting.
2026-06-17 21:17:13 +02:00
Renn F a9064decb7 feat(self-heal): wire the dormant orchestrator loop
Register _self_heal_loop alongside the other background loops (created in
start, cancelled in stop). It returns immediately unless self_heal_enabled, so
a standard deployment adds zero behaviour and makes no CI call; when on it runs
one engine cycle per interval and commits any opened fix task. A test pins the
default-off dormancy (no sleep / CI / DB when disabled).
2026-06-17 21:03:27 +02:00
Renn F bd1fb84198 feat(self-heal): open a PENDING fix task on regression, then stop
Behind the second opt-in (self_heal_originate_enabled), a detected regression
also opens a fix task into RoboCo's own delivery lifecycle and STOPS: PENDING,
unassigned, confirmed_by_human=False, team=main_pm, source=self_heal, with
synthesized acceptance criteria and a self_heal_fp= dedupe marker. It rides the
normal dispatchers only once the CEO Approve-&-Starts it; the loop itself never
calls start / approve / merge / deploy.

- TaskService: SELF_HEAL_SOURCE, extract_self_heal_fingerprint, and
  list_open_self_heal_tasks (the dedupe + open-cap basis)
- SelfHealEngine._originate: per-signal fingerprint dedupe, per-cycle and
  rolling open-task caps, repo resolved to RoboCo's own project (notify-only
  when it can't be resolved)
- 7 DB-backed tests including the never-start / never-approve invariant
2026-06-17 21:00:45 +02:00
Renn F 606ccc3327 feat(self-heal): regression engine — detect + notify the CEO (dormant)
The detect side of the self-healing loop, modeled on the strategy engine: a
pure assess() turns breaching telemetry samples into RegressionObservations
(with a stable per-signal fingerprint for later dedupe), and run_cycle() is a
no-op unless self_heal_enabled and otherwise only sends the CEO one
ack-notification per regression. Detect + notify only — it never originates,
starts, merges, or deploys; the telemetry source is injectable for testing.
6 unit tests.
2026-06-17 20:52:47 +02:00
Renn F 7e02713cc6 feat(self-heal): CI telemetry source for RoboCo's own repo (dormant)
First slice of the production self-healing loop: a read-only telemetry source
that watches RoboCo's OWN repo CI and normalizes the latest GitHub Actions run
conclusion into breach / no-breach samples for the regression detector. It
targets only the single project named by self_heal_project_slug — RoboCo
healing itself, never other/client repos; the org's repo-agnostic delivery flow
is untouched.

- config: self_heal_enabled / self_heal_project_slug / self_heal_originate_enabled
  plus interval and open-task / per-cycle caps, all default-off
- GitService.get_latest_ci_conclusion: per-project Actions-run lookup (graceful
  None on missing token / no runs / error; never raises into the loop)
- TelemetrySample + TelemetrySource contract + GitHubCITelemetrySource
- 5 unit tests
2026-06-17 20:50:07 +02:00
Renn F 47fda2e3fa docs(changelog): flesh out the v0.6.0 entry
The first pass under-represented the release. Make the inbound-PR-review
scope explicit (real GitHub review, author allowlist, head-SHA re-review,
repo-aware polling, 22nd agent with its own image, migration 037); add the
PR-reviewer respawn-loop fix and supersede close-on-land hardening to
Fixed; and note the richer agent-facing RAG docs (new Prompter / Secretary
/ PR-reviewer role docs + the company layer) under Changed.
v0.6.0
2026-06-17 17:58:13 +02:00
Renn F 1d835ff50f chore(release): prepare v0.6.0
Bump the version to 0.6.0 across pyproject, the package, the config, and
the panel, and add the 0.6.0 CHANGELOG entry: inbound external/internal PR
review with a CEO decision queue and supersede, the panel feature-flags
card, the required-cells decomposition gate, the CEO-rejected
coordination-root deadlock fix, the panel UI pass, and registry-image
deploy.

Also refresh the locked dependencies, update the release-tag examples in
the README and deployment docs, and correct the package docstring's agent
count to 22.
2026-06-17 17:53:08 +02:00
Renn F b278c55c64 docs(rag): fix PM delegate signature + cross-link cell-pm ↔ main-pm
The cell-pm role doc showed delegate with a nonexistent nested body={...}
and omitted covers_parent_criteria, so an agent following it would make a
malformed call and burn turns rediscovering the real shape. Both PM docs now
match the actual flow_server.delegate signature (flat keywords, with
covers_parent_criteria; the subtask inherits the parent's project — resolved
from the product cell→project map for coordination roots, never passed).

Also document reassign (cell PM could call it but it was undocumented) and
cross-link the two roles: cell-pm explains submit_up hands finished work to
Main PM; main-pm explains the receiving side — the integration-branch chain,
that its complete on the root opens the master PR, and that only Main PM and
the CEO act on master.
2026-06-17 17:21:50 +02:00
Renn F ef609085e4 docs(how-to): research & strategy now toggle from Settings → Feature Flags
The business-workflow chapter said neither capability has a panel switch —
no longer true once the Feature Flags card shipped. Reframe both as
panel-toggleable (effect on next restart), with the env vars kept as the
same toggles at the source plus the bits the panel doesn't surface (the
research provider and its server-side API key).
2026-06-17 17:10:18 +02:00
Renn F c826b03ac2 test(task): DB-backed coverage for the PR-review lifecycle methods
Real-Postgres round-trips for the external/internal PR-review TaskService
helpers that the existing mock tests can't prove actually persist:

- ingest_external_pr — create-once, head-SHA dedup (unchanged head skips),
  and re-review on a new head; internal_pr source wording;
- pr_review_claim / complete_review — the planless, branchless
  pending -> in_progress -> completed lifecycle, with re-claim / re-complete
  no-ops and the "complete requires in_progress" guard;
- create_supersede_umbrella / find_supersede_umbrella — created on the same
  repo (not parented), idempotent lookup, non-review rejection, and the
  pr=5-vs-pr=50 exact-marker disambiguation;
- list_external_pr_reviews — source isolation, the data-layer half of the
  dispatcher contract (regular tasks never leak into the review queue).

Writes use flush (not commit) so the rollback-per-test fixture keeps each
case isolated from the others against the shared session-scoped test DB.
2026-06-17 16:50:11 +02:00
Renn F f27a9f9447 feat(settings): panel-tunable feature flags
Add a Feature Flags card to the Settings page that toggles env-gated
subsystems (external/internal PR review, web research, strategy engine,
pitch provisioning, RAG auto-update, transcript pruning) directly from
the panel instead of hand-editing environment variables.

Flags persist in system_settings as 'true'/'false' and are overlaid onto
the live config singleton at startup; an unset flag keeps its
environment/config default. A toggle takes effect on the next backend
restart — no per-consumer re-routing.

Backend: FEATURE_FLAGS registry + bool validator + get_bool accessor on
SettingsService; feature_flag_effective_values and
apply_persisted_feature_flags; GET /settings/feature-flags; best-effort
startup overlay in the app lifespan.

Frontend: settingsApi.getFeatureFlags / setFeatureFlag and a
FeatureFlagsCard rendered full-width below the settings grid.
2026-06-17 16:50:11 +02:00
Renn F eec35c0357 fix(orchestrator): un-deadlock a CEO-rejected coordination root
A coordination root (team=main_pm, product-linked, no repo) the CEO sends back
lands in needs_revision, but the dev dispatcher skips it (not a cell team) and
the closure path only handles paused parents — so it sat in needs_revision
forever. (NOT a foundation-spec gap: the spec already allows needs_revision ->
claimed for any role.)

- _dispatch_revision_coordination_roots: re-spawn the owning PM for a
  needs_revision coordination root so it re-coordinates the revision (registered
  in the dispatch loop after PM closure)
- _readiness_check_role_for_status: widen the dev-owned states (needs_revision,
  verifying) to also accept cell_pm/main_pm for coordination roots — a pure
  widening; normal code tasks stay dev/doc-only
- 16 unit tests (dispatcher decision + readiness widening)
2026-06-17 16:50:10 +02:00
Renn F 748ff7813e test(git): httpx-mock coverage for list_open_prs + get_pr_diff
Covers the inbound-PR read surface: list_open_prs normalization + fork/internal
classification (and the recent _fetch_open_prs/_normalize_open_pr refactor),
plus get_pr_diff's diff-media-type request — both with their safe-empty paths
on missing token / GitHub error. The DB-backed paths
(ingest/complete_review/pr_review_claim/supersede umbrella) are covered
separately.
2026-06-17 16:50:10 +02:00
Renn F 66a8ad40eb feat(pr-review): internal-PR safety reviewer — review off-task-flow org PRs
Extend the inbound-PR reviewer beyond external/fork PRs to internal org-repo
PRs that bypassed the agent task-flow (a human-pushed branch). The org's own
in-flight integration PRs are skipped — a live task owns their branch and they
already pass QA + PM review — so the reviewer only flags off-process PRs.

- config: internal_pr_enabled (default OFF, like external_pr_enabled)
- PR_REVIEW_SOURCES = (external_pr, internal_pr); generalize dispatch, dedup,
  the decision queue, the git-gate exemption, and supersede to both sources
- TaskService.active_task_owns_branch (skip lifecycle PRs) + ingest source param
  with source-aware wording
- poll loop runs when EITHER flag is on; _ingest_pr_if_reviewable picks the
  source per PR (external: flag+author-allow; internal: flag+not-task-owned)
- 11 unit tests (decision logic + branch-ownership)
2026-06-17 16:50:09 +02:00
Renn F 34de96397f fix(panel): kanban card no longer overflows; PR-review queue shows an empty state
- Kanban card: show the short 8-char task id (full id on hover) instead of the
  full UUID, which was one unbreakable token that ran off the card edge
- PR Review queue: render an empty-state card instead of returning null when
  empty, so the surface is always visible on the Command Center (matching the
  CEO Approval Queue) rather than vanishing when there's nothing to decide
2026-06-17 08:01:49 +02:00
Renn F 94395d408d feat(gateway): structured required_cells gate — reject i_am_idle on a dropped named cell
The companion to the prompt rule (60de3499): when the brief explicitly names
cells, the Main PM must create a subtask for each and not silently collapse one
into a neighbour. Records the named cells as a 'required_cells:' marker on the
parent's quick_context (no migration — same pattern as the other markers), and
adds a _pm_uncovered_required_cells_guard at i_am_idle that refuses to idle
while a named cell has no subtask. Inert when no parent carries the marker, so
legacy decompositions are never blocked (mirrors the AC-coverage guard).
TaskService.uncovered_required_cells + extract_required_cells + 7 unit tests.
2026-06-17 08:01:49 +02:00
Renn F a175b65b0f refactor(git): cut cyclomatic complexity to clear the xenon gate
xenon flagged three rank-C blocks. Extracted helpers, no behavior change:
- list_open_prs -> _fetch_open_prs + _normalize_open_pr
- create_pull_request -> _resolve_new_pr_context + _existing_pr_tuple
- validate_git_requirements -> per-transition gate helpers
All three now rank <= B; ruff + mypy + git/lifecycle tests green.
2026-06-17 07:44:01 +02:00
Renn F cd89ba0ad5 fix(panel): UI revamp — settings layout, journals/kanban scroll, agent item, projects
- Settings: cards reordered to User Info / Appearance / Data & Refresh /
  Transcript Retention / Notifications / Connection Info (the 3x2 grid)
- Journals: the page now fills the viewport; the agent list and the entry
  detail each scroll inside their own panel — removes the page + list +
  fixed-500px triple scrollbar (real layout, not bolted-on magic heights)
- Agent item: distinct per-team avatar with initials, clear selected/hover
  states, truncation, focus ring (was a generic icon repeated on every row)
- Kanban: the board fills the viewport and each column's card list scrolls
  inside it, so a full Done column no longer overflows down the page
- Projects: drop the misleading Workspace column — it read the legacy
  per-project workspace_path (never set in the per-agent workspace model), so
  it always showed 'No workspace'
2026-06-17 07:44:00 +02:00
Renn F 6ce3cc1225 fix(panel): bottom-align and size the Secretary chat composer buttons
The Start/Send buttons used items-stretch with a fixed-height button, pinning
a cramped button to the top of a tall textarea. Bottom-align the row, give
Start a real primary size and Send a proper square icon button, and cap the
textarea height so the composer reads as a deliberate input.
2026-06-17 07:16:50 +02:00
Renn F 9cc63125d2 feat(external-pr): surface in-flight reviews in the panel, not just completed
The PR-review queue only listed COMPLETED reviews and hid when empty, so while
a review was in_progress the panel showed nothing — no sign a review was
happening or where its findings go (the reviewer posts its change-request on
the PR itself). Add TaskService.list_external_pr_reviews (active reviews +
awaiting-decision, minus cancelled/decided/dismissed); the route uses it. The
panel card now shows active reviews with a 'Reviewing' badge and a link to the
PR where the change-request lands, and the Supersede/Dismiss actions only once
the review completes.
2026-06-17 07:16:49 +02:00