Close the doc gaps the audit found in the agent knowledge base and the human
walkthrough:
- config-reference: add the Grok provider env table (host ~/.grok subscription
mount, grok-build, idle-kill, cost cap) and the Self-Healing CI loop toggles.
- agent-model: provider-aware Model Configuration (ANTHROPIC default / GROK) +
add the pr_reviewer / prompter / secretary roles to the Roles table.
- tool-permissions: 'three' -> five MCP servers (roboco-optimal, roboco-docs) +
PR Reviewer / Prompter / Secretary tool sections.
- new roles/pr-reviewer.md (the 22nd agent had no role doc); permissions +
agent-uuids + task-tools 'PR Reviewer flow' all gain the role.
- api-endpoints: drop the removed USAGE_UPDATE event (only USAGE_SNAPSHOT exists).
- how-to: self-healing CI loop + Company Scorecard (ch.5), inbound external-PR
review + CEO Supersede/Dismiss queue (ch.4).
Every claim verified against current code by the audit (grok model grok-build,
auth ~/.grok, no metered API; opencode fully removed).
* [e7349d84] feat(dashboard): WS usage store, hook extension, status badge, and smooth animations (#111) (#113)
- Add src/store/usage-store.ts with typed UsageData interface, useUsageStore
Zustand store, setUsageData, clearUsageData, and setWsState actions
- Export useUsageStore and UsageData from store/index.ts
- Extend use-rate-limit-websocket.ts: rename msg type to SystemWsMessage,
add key_metrics field; add useEffect syncing wsState into useUsageStore;
add USAGE_UPDATE/USAGE_SNAPSHOT handler dispatching to useUsageStore
(RATE_LIMIT_HIT/LIFTED handling and onReconnect unchanged)
- Update CommandCenter to read key_metrics from useUsageStore when
wsState === 'connected' and usageData non-null; falls back to
useCeoOverview() (refetchInterval: 60000) when WS disconnected
- Update KeyMetricsPanel: add wsState prop, render connection status Badge
matching AgentStreamViewer pattern (bg-green-500+Wifi / bg-yellow-500+
Loader2 spin / bg-gray-500+WifiOff); add transition-all duration-300
ease-in-out to metric value spans for smooth animated updates
Co-authored-by: Frontend Developer 1 <fe-dev-1@agents.roboco.dev>
* [c9745ee8] feat(events): add USAGE_UPDATE/SNAPSHOT event types, throttled publisher, /ws/system usage bridge (#112) (#114)
- Add EventType.USAGE_UPDATE='usage.update' and EventType.USAGE_SNAPSHOT='usage.snapshot'
to the EventType StrEnum in roboco/models/events.py
- Create roboco/services/usage_events.py with _UsageThrottle class (5-second per-agent
window using time.monotonic()) and publish_usage_update() / publish_usage_snapshot()
helpers; lazy imports prevent circular dependency with roboco.events
- Extend orchestrator._sweep_token_snapshots() to publish USAGE_UPDATE per active agent
(throttled) and a USAGE_SNAPSHOT aggregate after each sweep cycle; wrapped in
contextlib.suppress so event errors never abort DB snapshot operations
- Add _handle_usage_event() to websocket_bridge.py following _handle_rate_limit_event
pattern; register USAGE_UPDATE and USAGE_SNAPSHOT subscriptions in
register_websocket_bridge_handlers() forwarding both to /ws/system via broadcast_system()
- Add unit tests: test_usage_events.py (throttle suppression, publish helpers) and
test_websocket_bridge.py extended with _handle_usage_event coverage and updated
registration assertion to include USAGE_UPDATE/USAGE_SNAPSHOT
Co-authored-by: Backend Developer 1 <be-dev-1@agents.roboco.dev>
* fix(usage-ws): reconcile the realtime token/cost contract end-to-end
The backend and frontend halves shipped mismatched contracts, so the usage
dashboard never received live data:
- The bridge forwarded the dotted event value ("usage.update") while the panel
switched on "USAGE_UPDATE"; map both to the UPPER_SNAKE type string the same
way the rate-limit handler does.
- The backend emitted token/cost telemetry but the frontend read a key_metrics
field and fed the org-metrics panel. Rewire the frontend to consume the
USAGE_SNAPSHOT token/cost payload into the "Token Usage & Cost" panel —
WS-first with polling fallback and a connection-status badge — and revert the
unrelated KeyMetricsPanel / CommandCenter wiring.
Backend cleanups in the same path:
- Replace the multi-argument publish helpers with typed UsageUpdate /
UsageSnapshot payloads, removing the too-many-arguments lint suppressions.
- Extract _fetch_agent_tokens and _persist_token_snapshot from the token sweep,
removing the too-many-statements suppression; label the live snapshot "live".
Hardening uncovered while fixing the above:
- _finalize_spawn_session pulled the full RAG stack into the
session-finalization path through a transcript-parse import; move the pure
parser into a dependency-light roboco.agent_sdk.transcript_usage module so
finalization never imports the agent SDK server.
- Reduce _finalize_spawn_session complexity by extracting
_resolve_final_token_usage, and widen the transcript-fallback guard so a read
error can never abort finalization.
Also align KeyMetricsPanel with the metrics /dashboard/ceo actually returns: it
read velocity_24h / avg_time_to_done / active_agents, none of which
get_key_metrics() emits, so four of five rows rendered "—". Render
velocity_weekly, completion_rate, documentation_coverage and active_blockers.
* docs: note live usage push over /ws/system on the usage dashboard
* fix(usage): finalize on self-exit and de-duplicate transcript token counts
Two bugs left token capture broken even after the transcript-read fallback
landed — surfaced by a live agent run:
- Agents that self-exit (the normal i_am_idle -> container shutdown, exit 0)
were never finalized. _finalize_spawn_session is only called from
stop_agent(), but a graceful self-exit goes through _handle_stopped_container,
which set the instance OFFLINE and returned without finalizing — leaving the
spawn-session row open with zero tokens. Finalize there for both graceful
(exit_reason="completed") and crash (exit_reason="crashed") exits.
- sum_transcript_usage double-counted. Claude Code logs one assistant message
as several JSONL lines (one per content block — thinking / text / tool_use),
each repeating the same message.usage, so summing every line roughly doubled
the totals. De-duplicate by message.id.
Verified against a live agent transcript: the raw sum (12, 1068, 62502, 115828)
vs the de-duped (6, 516, 62502, 63336), which matches the session's
authoritative result.usage exactly.
* feat(usage): fall back to the transcript in the live token sweep
The 60s token sweep read only the agent SDK's /usage/status, which races
container teardown and reports zero mid-run — so live usage (and the
USAGE_SNAPSHOT pushed to /ws/system) stayed at zero for active agents.
Extract _resolve_active_tokens: try the SDK, then fall back to the durable
transcript (the same source finalize uses) so running agents report live.
* feat(usage): add GET /usage/sessions for the dashboard's Recent Sessions
The panel's Recent Sessions table was mock-only — the backend had no sessions
endpoint, so production always showed 'No sessions recorded yet'. Add
UsageService.get_recent_sessions + a /usage/sessions route returning the most
recent spawn-session rows (token totals + cost), and point the panel client
at it.
---------
Co-authored-by: Frontend Developer 1 <fe-dev-1@agents.roboco.dev>
Co-authored-by: Backend Developer 1 <be-dev-1@agents.roboco.dev>
Co-authored-by: Renn F <rennf93@users.noreply.github.com>
Record the features that landed this cycle:
- CHANGELOG: provider rate-limit handling, token usage & cost analytics, the
/ws/system operator stream; plus the fixes (agent gate toolchain, usage
capture, panel endpoint shape + WS path, /public 500, provider pricing).
- CLAUDE.md: a WebSocket-streams section (incl. /ws/system + the
websocket_bridge pattern), a Rate-limiting & usage subsystem note, and the
'uv sync --extra dev' workspace-toolchain requirement.
- agent API reference: a System & realtime section (/api/system/rate-limits,
/ws/system, per-resource WS streams).