feat(grok-cli): GrokCliProvider — subscription auth mount, mirrors ClaudeCodeProvider

Replace the opencode GrokProvider with GrokCliProvider: reuses the orchestrator's
shared mount/auth/git assembly (gateway + identity) exactly like the Claude path,
mounts the host ~/.grok/auth.json read-only (SuperGrok subscription) instead of
injecting an xAI key, and sets the slim env the grok-cli entrypoint + renderer
read (ROBOCO_AGENT_ID for per-role flags, model, mcp-config, prompt). Provider
routing fields are blanked before the shared step so the grok endpoint is never
mislabelled ANTHROPIC_*. Per-role permission logic now lives in grok_cli_config,
so the provider is slim. Registry/orchestrator/exports updated; provider tests
rewritten for the CLI behavior (no XAI key, auth mount present/absent).
This commit is contained in:
Renn F
2026-06-19 03:01:46 +02:00
parent 4a1ea9d38e
commit aa4dc08d58
7 changed files with 137 additions and 350 deletions
+4 -3
View File
@@ -8,18 +8,19 @@ and falls back to its built-in Claude Code spawn when none is registered.
Backends: Backends:
- :class:`ClaudeCodeProvider` — Anthropic-protocol Claude Code container (default; - :class:`ClaudeCodeProvider` — Anthropic-protocol Claude Code container (default;
also serves Ollama Cloud / self-hosted via ``ANTHROPIC_BASE_URL`` injection). also serves Ollama Cloud / self-hosted via ``ANTHROPIC_BASE_URL`` injection).
- :class:`GrokProvider` — xAI ``grok-build-0.1`` over the OpenAI protocol. - :class:`GrokCliProvider` — xAI Grok Build via the official ``grok`` CLI on the
SuperGrok subscription (mounted ``~/.grok`` auth, parity with the Claude path).
""" """
from roboco.llm.providers.base import AgentProvider, ProviderError, SpawnResult from roboco.llm.providers.base import AgentProvider, ProviderError, SpawnResult
from roboco.llm.providers.claude_code import ClaudeCodeProvider from roboco.llm.providers.claude_code import ClaudeCodeProvider
from roboco.llm.providers.grok import GrokProvider from roboco.llm.providers.grok import GrokCliProvider
from roboco.llm.providers.registry import ProviderNotRegisteredError, ProviderRegistry from roboco.llm.providers.registry import ProviderNotRegisteredError, ProviderRegistry
__all__ = [ __all__ = [
"AgentProvider", "AgentProvider",
"ClaudeCodeProvider", "ClaudeCodeProvider",
"GrokProvider", "GrokCliProvider",
"ProviderError", "ProviderError",
"ProviderNotRegisteredError", "ProviderNotRegisteredError",
"ProviderRegistry", "ProviderRegistry",
+58 -186
View File
@@ -1,37 +1,26 @@
"""Grok provider — xAI ``grok-build-0.1`` as a native OpenAI-protocol agent. """Grok CLI provider — xAI Grok Build via the official ``grok`` CLI.
xAI's API is OpenAI-compatible *only* (``https://api.x.ai/v1``) — there is no xAI ships an official terminal coding agent (the ``grok`` CLI, "Grok Build")
native Anthropic-Messages endpoint — so a Grok agent cannot run through the authenticated by the SuperGrok subscription. RoboCo runs Grok agents on it the
Claude Code path (which speaks the Anthropic Messages API via same way it runs Claude agents on ``claude``: the orchestrator's shared container
``ANTHROPIC_BASE_URL`` injection). Instead it runs an OpenAI-protocol agent CLI assembly mounts the RoboCo MCP gateway (``mcp-config.json``), the agent HMAC
(the OpenCode pattern), pointed at ``api.x.ai/v1`` with the operator's xAI key. identity, and the git context; this provider adds the subscription auth mount
(``~/.grok``) and the runtime env the grok-cli entrypoint reads, then launches
the ``roboco-agent-grok`` image — whose entrypoint renders ``~/.grok/config.toml``
+ per-role flags (see :mod:`roboco.llm.providers.grok_cli_config`) and runs
``grok -p`` headless.
Design — this provider deliberately **reuses the orchestrator's proven container Two things differ from the Claude Code spawn:
assembly** (``_build_mount_args`` / ``_append_agent_auth_env`` / 1. **Auth** — the host's ``~/.grok`` (subscription credential from ``grok
``_append_git_context_env``). That means a Grok agent gets the *same* RoboCo MCP login``) is mounted instead of relying on a provider key; the xAI API key is
gateway wiring as every other agent **by construction**: never used. The provider routing fields are blanked before the shared mount
step so the shared builder never injects them as ``ANTHROPIC_*`` (the wrong
* ``/app/mcp-config.json`` (the ``roboco-flow`` / ``roboco-do`` gateway) and runtime) — grok authenticates from the mounted ``~/.grok``.
``/app/system-prompt.md`` are mounted by ``_build_mount_args``; 2. **Runtime** — the ``roboco-agent-grok`` image (grok CLI) instead of
* the spawn manifest + ``ROBOCO_GATEWAY_ENABLED`` are set there too; ``claude``.
* the agent HMAC identity is injected by ``_append_agent_auth_env``.
An OpenAI-protocol provider that built its own bespoke spawn would have to
re-wire the MCP config itself, and skipping it leaves agents with zero gateway
verbs. Reusing the shared assembly makes the gateway wiring non-optional here.
Only two things differ from the Claude Code spawn:
1. **LLM env** — ``OPENAI_BASE_URL`` / ``OPENAI_API_KEY`` (xAI) instead of the
``ANTHROPIC_*`` injection. The provider routing fields are blanked before
the shared mount step so the xAI endpoint is never mislabelled as Anthropic.
2. **Runtime** — the ``roboco-agent-grok`` image, whose entrypoint launches the
OpenAI-protocol CLI from the env below (model, MCP config, system prompt,
prompt). The image + the exact CLI invocation are the one remaining piece to
finalise together with xAI ("review + help implement + test").
The initial prompt is passed via an **env var, not a positional CLI arg**, which The initial prompt is passed via an **env var, not a positional CLI arg**, which
structurally avoids a flag-injection vector: a prompt starting with ``--`` passed structurally avoids a flag-injection vector.
as a positional argument could otherwise be parsed as CLI options.
""" """
from __future__ import annotations from __future__ import annotations
@@ -39,109 +28,34 @@ from __future__ import annotations
import asyncio import asyncio
import dataclasses import dataclasses
import os import os
from pathlib import Path
from typing import TYPE_CHECKING, Protocol from typing import TYPE_CHECKING, Protocol
from roboco.agents_config import get_agent_role
from roboco.llm.providers._docker import container_running, stop_container from roboco.llm.providers._docker import container_running, stop_container
from roboco.llm.providers.base import AgentProvider, ProviderError, SpawnResult from roboco.llm.providers.base import AgentProvider, ProviderError, SpawnResult
from roboco.services.gateway.role_config import get_role_config
if TYPE_CHECKING: if TYPE_CHECKING:
from pathlib import Path
from roboco.models.runtime import OrchestratorAgentConfig as AgentConfig from roboco.models.runtime import OrchestratorAgentConfig as AgentConfig
# The Grok agent image (own image, like every other agent role). Built as the # The Grok agent image (own image, like every other agent role). Overridable for
# infra follow-up — bundles the OpenAI-protocol agent CLI + an entrypoint that # tests / staged rollout.
# reads the env contract below. Overridable for tests / staged rollout.
_DEFAULT_GROK_IMAGE = os.environ.get( _DEFAULT_GROK_IMAGE = os.environ.get(
"ROBOCO_GROK_AGENT_IMAGE", "roboco-agent-grok:latest" "ROBOCO_GROK_AGENT_IMAGE", "roboco-agent-grok:latest"
) )
# Default xAI endpoint when the seeded provider row carries no base_url. # The grok CLI model id (the CLI uses ``grok-build``, verified live).
_DEFAULT_XAI_BASE_URL = "https://api.x.ai/v1" _GROK_CLI_MODEL = os.environ.get("ROBOCO_GROK_CLI_MODEL", "grok-build")
# In-container paths mounted by the orchestrator's `_build_mount_args`. # Host directory holding the SuperGrok auth (from ``grok login``). Mounted into
_MCP_CONFIG_IN_CONTAINER = "/app/mcp-config.json" # the agent's ``~/.grok`` like the Claude path mounts ``~/.claude``. Override for
_SYSTEM_PROMPT_IN_CONTAINER = "/app/system-prompt.md" # docker-in-docker / NAS deploys (the orchestrator's home is not the host's).
# opencode's data dir inside the agent (HOME=/home/agent); opencode.db lands GROK_AUTH_HOST_PATH = os.environ.get(
# here. Mounted to a per-agent host dir so the orchestrator can read usage back "ROBOCO_HOST_GROK_DIR", str(Path.home() / ".grok")
# (mirror of roboco.llm.providers.opencode_usage.DEFAULT_DB_PATH's parent).
_OPENCODE_DATA_DIR_IN_CONTAINER = "/home/agent/.local/share/opencode"
# Reasoning effort by role. grok-build-0.1 reasons heavily by default, and
# reasoning bills at the output rate, so it dominates cost. Code-quality roles
# (developer, qa, pr_reviewer) keep full reasoning; coordination / docs / board
# roles request "minimal". opencode receives this via its `--variant` flag (and
# the serve message `variant` field). NOTE: whether opencode actually applies a
# named reasoning variant to grok-build-0.1 without a provider-defined `variants`
# block is UNVERIFIED — passing the flag does not error, but the reasoning-cost
# reduction must be measured on the NAS; treat the saving as best-effort, not
# guaranteed. Operators can force one effort for ALL grok agents with the
# ROBOCO_GROK_REASONING_EFFORT env on the orchestrator (value "minimal" | "high"
# | "max", or "default"/"full" to use full reasoning).
_MINIMAL_REASONING_ROLES = frozenset(
{
"cell_pm",
"main_pm",
"documenter",
"product_owner",
"head_marketing",
"auditor",
"prompter",
"secretary",
}
) )
_FULL_REASONING_OVERRIDES = frozenset({"default", "full", "none", ""})
# Per-role opencode permission policy (Claude-parity with # In-container paths.
# orchestrator._get_role_permissions). Claude denies Write/Edit for the read-only _MCP_CONFIG_IN_CONTAINER = "/app/mcp-config.json"
# roles and Bash(git commit/push) for PMs; opencode's permission is coarser _GROK_AUTH_IN_CONTAINER = "/home/agent/.grok/auth.json"
# (allow/deny per tool class), so:
# * edit — allow only roles that write code (role_config.allows_write:
# developer / documenter). Everyone else edit=deny.
# * bash — allow only roles that legitimately run a shell; the read-only
# reviewers (qa / pr_reviewer / auditor) and the board never do. secret-scrub
# still guards bash (git-mutate / cred files) for the roles that keep it.
# * external_directory — only the pr_reviewer reads scratch outside its cwd
# (a diff it writes to /tmp). Delivery roles work inside their workspace, so
# external_directory=deny (the headless-ask auto-deny that blocked the
# pr-reviewer is moot once that one role is explicitly allowed).
_BASH_ROLES = frozenset({"developer", "documenter", "cell_pm", "main_pm"})
_EXTERNAL_DIR_ROLES = frozenset({"pr_reviewer"})
def _edit_permission_for(agent_id: str) -> str:
"""opencode ``edit`` permission for an agent's role (allow iff it writes code)."""
role = get_agent_role(agent_id) or ""
try:
return "allow" if get_role_config(role).allows_write else "deny"
except KeyError:
return "deny" # unknown role → safest
def _bash_permission_for(agent_id: str) -> str:
"""opencode ``bash`` permission for an agent's role."""
return "allow" if (get_agent_role(agent_id) or "") in _BASH_ROLES else "deny"
def _external_dir_permission_for(agent_id: str) -> str:
"""opencode ``external_directory`` permission for an agent's role."""
role = get_agent_role(agent_id) or ""
return "allow" if role in _EXTERNAL_DIR_ROLES else "deny"
def _reasoning_effort_for(agent_id: str) -> str | None:
"""Resolve the opencode --variant reasoning effort for an agent.
Returns ``None`` to use opencode's default (full) reasoning. A global
override env wins over the per-role default.
"""
override = os.environ.get("ROBOCO_GROK_REASONING_EFFORT", "").strip()
if override:
return None if override.lower() in _FULL_REASONING_OVERRIDES else override
role = get_agent_role(agent_id) or ""
return "minimal" if role in _MINIMAL_REASONING_ROLES else None
def _container_name(agent_id: str) -> str: def _container_name(agent_id: str) -> str:
@@ -149,16 +63,14 @@ def _container_name(agent_id: str) -> str:
class _GrokHost(Protocol): class _GrokHost(Protocol):
"""The orchestrator surface GrokProvider reuses for container assembly. """The orchestrator surface GrokCliProvider reuses for container assembly.
Typed as a Protocol so this module never imports ``AgentOrchestrator`` Typed as a Protocol so this module never imports ``AgentOrchestrator`` (no
(no import cycle) and is trivially mockable in tests. import cycle) and is trivially mockable in tests.
""" """
async def _remove_container(self, container_name: str) -> None: ... async def _remove_container(self, container_name: str) -> None: ...
def _ensure_opencode_data_dir(self, agent_id: str) -> None: ...
def _resolve_host_paths( def _resolve_host_paths(
self, config: AgentConfig, agent_settings_path: Path | None self, config: AgentConfig, agent_settings_path: Path | None
) -> dict[str, str | None]: ... ) -> dict[str, str | None]: ...
@@ -175,8 +87,8 @@ class _GrokHost(Protocol):
def _append_git_context_env(self, cmd: list[str], config: AgentConfig) -> None: ... def _append_git_context_env(self, cmd: list[str], config: AgentConfig) -> None: ...
class GrokProvider(AgentProvider): class GrokCliProvider(AgentProvider):
"""Spawn a Grok (xAI, OpenAI-protocol) agent as a gateway-wired container.""" """Spawn a Grok (xAI, official CLI) agent as a gateway-wired container."""
def __init__(self, host: _GrokHost, image: str | None = None) -> None: def __init__(self, host: _GrokHost, image: str | None = None) -> None:
self._host = host self._host = host
@@ -188,12 +100,6 @@ class GrokProvider(AgentProvider):
initial_prompt: str | None = None, initial_prompt: str | None = None,
agent_settings_path: Path | None = None, agent_settings_path: Path | None = None,
) -> SpawnResult: ) -> SpawnResult:
if not config.provider_auth_token:
raise ProviderError(
"GROK spawn requires an xAI API key — set the Grok provider key "
"in Settings (PUT /api/providers/grok/key).",
agent_id=config.agent_id,
)
if not config.mcp_config_path: if not config.mcp_config_path:
raise ProviderError( raise ProviderError(
"GROK spawn requires an MCP config (gateway access).", "GROK spawn requires an MCP config (gateway access).",
@@ -202,14 +108,12 @@ class GrokProvider(AgentProvider):
container_name = _container_name(config.agent_id) container_name = _container_name(config.agent_id)
await self._host._remove_container(container_name) await self._host._remove_container(container_name)
# Pre-create the opencode store dir (world-writable) before the bind mount
# so the non-root agent user can write opencode.db / repos (else EACCES).
self._host._ensure_opencode_data_dir(config.agent_id)
# Reuse the orchestrator's mount/auth/git assembly so the agent gets the # Reuse the orchestrator's mount/auth/git assembly so the agent gets the
# full MCP gateway + identity wiring. Blank the provider routing fields # full MCP gateway + identity wiring. Blank the provider routing fields
# first: otherwise the shared builder would inject the xAI endpoint as # first: otherwise the shared builder would inject the provider endpoint
# ANTHROPIC_BASE_URL/ANTHROPIC_AUTH_TOKEN (wrong protocol). # as ANTHROPIC_BASE_URL/AUTH_TOKEN — grok authenticates from the mounted
# ~/.grok, not a provider key.
mount_config = dataclasses.replace( mount_config = dataclasses.replace(
config, provider_base_url=None, provider_auth_token=None config, provider_base_url=None, provider_auth_token=None
) )
@@ -217,7 +121,7 @@ class GrokProvider(AgentProvider):
cmd = self._host._build_mount_args(container_name, mount_config, hosts) cmd = self._host._build_mount_args(container_name, mount_config, hosts)
self._host._append_agent_auth_env(cmd, config) self._host._append_agent_auth_env(cmd, config)
self._host._append_git_context_env(cmd, config) self._host._append_git_context_env(cmd, config)
self._append_opencode_data_mount(cmd, hosts) self._append_grok_auth_mount(cmd)
self._append_grok_env(cmd, config, initial_prompt) self._append_grok_env(cmd, config, initial_prompt)
cmd.append(self._image) cmd.append(self._image)
@@ -234,75 +138,43 @@ class GrokProvider(AgentProvider):
) )
return SpawnResult( return SpawnResult(
instance_id=container_name, instance_id=container_name,
extra={"container_id": stdout.decode().strip(), "model": config.model}, extra={"container_id": stdout.decode().strip(), "model": _GROK_CLI_MODEL},
) )
def _append_opencode_data_mount( @staticmethod
self, cmd: list[str], hosts: dict[str, str | None] def _append_grok_auth_mount(cmd: list[str]) -> None:
) -> None: """Mount the host's SuperGrok ``auth.json`` (read-only) into ~/.grok.
"""Mount the per-agent opencode data dir so the orchestrator can read it.
opencode persists token usage to ``opencode.db`` under its data dir Read-only so concurrent containers can't corrupt the shared subscription
(``$HOME/.local/share/opencode``). Binding a per-agent host dir there credential; grok writes its per-run state (the rendered ``config.toml``,
lets the finalizer read the store back over the shared data volume — ``sessions/``) into the image's own ``~/.grok``. One-shot delivery runs
the opencode analogue of the mounted Claude transcript. Without this a are short, so the token needs no mid-run refresh.
Grok agent finalizes at 0 tokens / $0.
""" """
opencode_host = hosts.get("opencode") auth_json = Path(GROK_AUTH_HOST_PATH) / "auth.json"
if opencode_host: if auth_json.exists():
cmd.extend(["-v", f"{opencode_host}:{_OPENCODE_DATA_DIR_IN_CONTAINER}"]) cmd.extend(["-v", f"{auth_json}:{_GROK_AUTH_IN_CONTAINER}:ro"])
def _append_grok_env( def _append_grok_env(
self, cmd: list[str], config: AgentConfig, initial_prompt: str | None self, cmd: list[str], config: AgentConfig, initial_prompt: str | None
) -> None: ) -> None:
"""Append the OpenAI-protocol (xAI) env contract the grok image consumes. """Append the runtime env the grok-cli entrypoint + renderer read.
The prompt travels as an env var, never an argv positional, so a prompt ``ROBOCO_AGENT_ID`` lets the renderer compute the per-role flags;
beginning with ``--`` cannot be parsed as a CLI flag. ``ROBOCO_MCP_CONFIG`` points it at the mounted gateway config; the prompt
travels as an env var (never an argv positional).
""" """
base_url = config.provider_base_url or _DEFAULT_XAI_BASE_URL
cmd.extend( cmd.extend(
[ [
# opencode's BUILT-IN xai provider authenticates from XAI_API_KEY
# and reads XAI_BASE_URL for the endpoint — opencode_config emits
# no provider block, so these envs are the only LLM wiring needed.
"-e", "-e",
f"XAI_API_KEY={config.provider_auth_token}", f"ROBOCO_AGENT_ID={config.agent_id}",
"-e", "-e",
f"XAI_BASE_URL={base_url}", f"ROBOCO_AGENT_MODEL={_GROK_CLI_MODEL}",
# Operational inputs for the grok image entrypoint.
"-e",
f"ROBOCO_AGENT_MODEL={config.model}",
"-e", "-e",
f"ROBOCO_MCP_CONFIG={_MCP_CONFIG_IN_CONTAINER}", f"ROBOCO_MCP_CONFIG={_MCP_CONFIG_IN_CONTAINER}",
"-e", "-e",
f"ROBOCO_SYSTEM_PROMPT={_SYSTEM_PROMPT_IN_CONTAINER}",
"-e",
f"ROBOCO_INITIAL_PROMPT={initial_prompt or ''}", f"ROBOCO_INITIAL_PROMPT={initial_prompt or ''}",
] ]
) )
if config.claude_session_id:
# Reused as the generic agent session id so the transcript stays
# locatable at finalize, exactly as on the Claude Code path.
cmd.extend(["-e", f"ROBOCO_AGENT_SESSION_ID={config.claude_session_id}"])
# Per-role opencode permissions (Claude-parity): read-only roles get
# edit=deny, only delivery roles get bash, only the pr-reviewer gets
# external-directory reads. opencode_config.main() reads these.
cmd.extend(
[
"-e",
f"ROBOCO_GROK_EDIT_PERMISSION={_edit_permission_for(config.agent_id)}",
"-e",
f"ROBOCO_GROK_BASH_PERMISSION={_bash_permission_for(config.agent_id)}",
"-e",
"ROBOCO_GROK_EXTERNAL_DIR_PERMISSION="
f"{_external_dir_permission_for(config.agent_id)}",
]
)
# Reasoning effort (opencode --variant) by role; omitted = full reasoning.
variant = _reasoning_effort_for(config.agent_id)
if variant:
cmd.extend(["-e", f"ROBOCO_GROK_VARIANT={variant}"])
async def stop(self, instance_id: str, graceful: bool = True) -> None: async def stop(self, instance_id: str, graceful: bool = True) -> None:
await stop_container(instance_id, graceful) await stop_container(instance_id, graceful)
+1 -1
View File
@@ -3,7 +3,7 @@
Usage:: Usage::
registry = ProviderRegistry() registry = ProviderRegistry()
registry.register(ModelProvider.GROK, GrokProvider(...)) registry.register(ModelProvider.GROK, GrokCliProvider(...))
provider = registry.get(ModelProvider.GROK) provider = registry.get(ModelProvider.GROK)
result = await provider.spawn(config, initial_prompt) result = await provider.spawn(config, initial_prompt)
+1 -1
View File
@@ -73,7 +73,7 @@ MODEL_CATALOG: tuple[CatalogEntry, ...] = (
CatalogEntry("kimi-k2.6:cloud", ModelProvider.OLLAMA_CLOUD, "Kimi K2.6"), CatalogEntry("kimi-k2.6:cloud", ModelProvider.OLLAMA_CLOUD, "Kimi K2.6"),
CatalogEntry("minimax-m3:cloud", ModelProvider.OLLAMA_CLOUD, "Minimax M3"), CatalogEntry("minimax-m3:cloud", ModelProvider.OLLAMA_CLOUD, "Minimax M3"),
# --- Grok (xAI, OpenAI protocol) --- # --- Grok (xAI, OpenAI protocol) ---
# Routes to the GROK provider → GrokProvider spawn (api.x.ai/v1). The xAI # Routes to the GROK provider → GrokCliProvider spawn (api.x.ai/v1). The xAI
# key is set via PUT /api/providers/grok/key. # key is set via PUT /api/providers/grok/key.
CatalogEntry("grok-build-0.1", ModelProvider.GROK, "Grok Build 0.1"), CatalogEntry("grok-build-0.1", ModelProvider.GROK, "Grok Build 0.1"),
) )
+2 -2
View File
@@ -2076,7 +2076,7 @@ class AgentOrchestrator:
container are registered. Today that is GROK (xAI, OpenAI protocol). container are registered. Today that is GROK (xAI, OpenAI protocol).
""" """
if self._provider_registry is None: if self._provider_registry is None:
from roboco.llm.providers import GrokProvider, ProviderRegistry from roboco.llm.providers import GrokCliProvider, ProviderRegistry
from roboco.models.base import ModelProvider from roboco.models.base import ModelProvider
registry = ProviderRegistry() registry = ProviderRegistry()
@@ -2085,7 +2085,7 @@ class AgentOrchestrator:
# get_agent_image for the Claude path). # get_agent_image for the Claude path).
registry.register( registry.register(
ModelProvider.GROK, ModelProvider.GROK,
GrokProvider(self, image=_qualify_agent_image("roboco-agent-grok")), GrokCliProvider(self, image=_qualify_agent_image("roboco-agent-grok")),
) )
self._provider_registry = registry self._provider_registry = registry
return self._provider_registry return self._provider_registry
+68 -154
View File
@@ -1,11 +1,12 @@
"""Tests for the LLM agent provider seam. """Tests for the LLM agent provider seam.
Covers the ProviderRegistry, the ClaudeCodeProvider adapter, and the Covers the ProviderRegistry, the ClaudeCodeProvider adapter, and the
GrokProvider (xAI / OpenAI protocol) — especially the safety properties an GrokCliProvider (xAI Grok Build via the official ``grok`` CLI) — especially the
OpenAI-protocol agent provider must hold: safety properties the Grok provider must hold:
* the agent gets the MCP gateway wiring (reuses the orchestrator mount path); * the agent gets the MCP gateway wiring (reuses the orchestrator mount path);
* the xAI endpoint is injected as XAI_* and never mislabelled ANTHROPIC_*; * the subscription auth (~/.grok) is mounted, and the provider routing fields
are blanked so the grok endpoint is never mislabelled ANTHROPIC_*;
* the prompt travels via env, so a leading ``--`` cannot become a CLI flag. * the prompt travels via env, so a leading ``--`` cannot become a CLI flag.
""" """
@@ -17,22 +18,28 @@ from unittest.mock import AsyncMock, MagicMock, patch
import pytest import pytest
from roboco.llm.providers import ( from roboco.llm.providers import (
ClaudeCodeProvider, ClaudeCodeProvider,
GrokProvider, GrokCliProvider,
ProviderError, ProviderError,
ProviderNotRegisteredError, ProviderNotRegisteredError,
ProviderRegistry, ProviderRegistry,
SpawnResult, SpawnResult,
) )
from roboco.llm.providers.grok import (
_bash_permission_for,
_edit_permission_for,
_external_dir_permission_for,
_reasoning_effort_for,
)
from roboco.models.base import ModelProvider from roboco.models.base import ModelProvider
from roboco.models.runtime import OrchestratorAgentConfig from roboco.models.runtime import OrchestratorAgentConfig
@pytest.fixture(autouse=True)
def _isolate_grok_auth(
tmp_path: Path, monkeypatch: pytest.MonkeyPatch
) -> Path:
"""Point GROK_AUTH_HOST_PATH at a fresh tmp dir so tests never mount the real
~/.grok. Tests that exercise the auth mount create ``auth.json`` themselves."""
monkeypatch.setattr(
"roboco.llm.providers.grok.GROK_AUTH_HOST_PATH", str(tmp_path)
)
return tmp_path
def _config( def _config(
*, *,
agent_id: str = "be-dev-1", agent_id: str = "be-dev-1",
@@ -60,7 +67,6 @@ class _FakeHost:
self.removed: list[str] = [] self.removed: list[str] = []
self.spawn_args: tuple[object, ...] | None = None self.spawn_args: tuple[object, ...] | None = None
self.mount_config: OrchestratorAgentConfig | None = None self.mount_config: OrchestratorAgentConfig | None = None
self.opencode_dirs_ensured: list[str] = []
async def _spawn_container( async def _spawn_container(
self, self,
@@ -74,9 +80,6 @@ class _FakeHost:
async def _remove_container(self, container_name: str) -> None: async def _remove_container(self, container_name: str) -> None:
self.removed.append(container_name) self.removed.append(container_name)
def _ensure_opencode_data_dir(self, agent_id: str) -> None:
self.opencode_dirs_ensured.append(agent_id)
def _resolve_host_paths( def _resolve_host_paths(
self, config: OrchestratorAgentConfig, agent_settings_path: Path | None self, config: OrchestratorAgentConfig, agent_settings_path: Path | None
) -> dict[str, str | None]: ) -> dict[str, str | None]:
@@ -85,7 +88,6 @@ class _FakeHost:
if config.mcp_config_path if config.mcp_config_path
else None, else None,
"settings": str(agent_settings_path) if agent_settings_path else None, "settings": str(agent_settings_path) if agent_settings_path else None,
"opencode": f"/host/opencode/{config.agent_id}",
} }
def _build_mount_args( def _build_mount_args(
@@ -134,7 +136,7 @@ def _proc(
def test_registry_register_and_get() -> None: def test_registry_register_and_get() -> None:
registry = ProviderRegistry() registry = ProviderRegistry()
provider = GrokProvider(_FakeHost()) provider = GrokCliProvider(_FakeHost())
registry.register(ModelProvider.GROK, provider) registry.register(ModelProvider.GROK, provider)
assert registry.get(ModelProvider.GROK) is provider assert registry.get(ModelProvider.GROK) is provider
assert registry.is_registered(ModelProvider.GROK) assert registry.is_registered(ModelProvider.GROK)
@@ -154,42 +156,43 @@ def test_registry_get_or_none_returns_none_when_absent() -> None:
def test_registry_unregister() -> None: def test_registry_unregister() -> None:
registry = ProviderRegistry() registry = ProviderRegistry()
registry.register(ModelProvider.GROK, GrokProvider(_FakeHost())) registry.register(ModelProvider.GROK, GrokCliProvider(_FakeHost()))
registry.unregister(ModelProvider.GROK) registry.unregister(ModelProvider.GROK)
assert not registry.is_registered(ModelProvider.GROK) assert not registry.is_registered(ModelProvider.GROK)
registry.unregister(ModelProvider.GROK) # idempotent registry.unregister(ModelProvider.GROK) # idempotent
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
# GrokProvider # GrokCliProvider
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
async def test_grok_spawn_requires_api_key() -> None:
provider = GrokProvider(_FakeHost())
with pytest.raises(ProviderError, match="xAI API key"):
await provider.spawn(_config(provider_auth_token=None))
async def test_grok_spawn_requires_mcp_config() -> None: async def test_grok_spawn_requires_mcp_config() -> None:
provider = GrokProvider(_FakeHost()) provider = GrokCliProvider(_FakeHost())
with pytest.raises(ProviderError, match="MCP config"): with pytest.raises(ProviderError, match="MCP config"):
await provider.spawn(_config(mcp_config_path=None)) await provider.spawn(_config(mcp_config_path=None))
async def test_grok_spawn_injects_xai_env_and_no_anthropic_leak() -> None: async def test_grok_spawn_does_not_require_api_key() -> None:
# Subscription auth (mounted ~/.grok) — a missing provider key is fine.
host = _FakeHost() host = _FakeHost()
provider = GrokProvider(host, image="roboco-agent-grok:test") provider = GrokCliProvider(host)
with patch("asyncio.create_subprocess_exec", AsyncMock(return_value=_proc())):
result = await provider.spawn(_config(provider_auth_token=None))
assert result.instance_id == "roboco-agent-be-dev-1"
async def test_grok_spawn_no_xai_key_and_no_anthropic_leak() -> None:
host = _FakeHost()
provider = GrokCliProvider(host, image="roboco-agent-grok:test")
with patch( with patch(
"asyncio.create_subprocess_exec", AsyncMock(return_value=_proc()) "asyncio.create_subprocess_exec", AsyncMock(return_value=_proc())
) as exec_mock: ) as exec_mock:
await provider.spawn(_config(), initial_prompt="do the work") await provider.spawn(_config(), initial_prompt="do the work")
cmd = list(exec_mock.call_args.args) cmd = list(exec_mock.call_args.args)
# opencode's built-in xai provider reads XAI_API_KEY / XAI_BASE_URL (no # The CLI authenticates from the mounted ~/.grok — the xAI key is never used.
# provider block in the rendered config — that breaks plugin-tool reg). assert not any(c.startswith("XAI_API_KEY=") for c in cmd)
assert "XAI_API_KEY=xai-secret-key" in cmd # The provider endpoint must NOT be injected as an Anthropic var.
assert "XAI_BASE_URL=https://api.x.ai/v1" in cmd
# The xAI endpoint must NOT be injected as an Anthropic var.
assert not any(c.startswith("ANTHROPIC_BASE_URL=") for c in cmd) assert not any(c.startswith("ANTHROPIC_BASE_URL=") for c in cmd)
assert not any(c.startswith("ANTHROPIC_AUTH_TOKEN=") for c in cmd) assert not any(c.startswith("ANTHROPIC_AUTH_TOKEN=") for c in cmd)
# Provider fields were blanked before the shared mount step. # Provider fields were blanked before the shared mount step.
@@ -198,23 +201,18 @@ async def test_grok_spawn_injects_xai_env_and_no_anthropic_leak() -> None:
assert host.mount_config.provider_auth_token is None assert host.mount_config.provider_auth_token is None
async def test_grok_spawn_wires_gateway_and_image_last() -> None: async def test_grok_spawn_wires_gateway_env_and_image_last() -> None:
host = _FakeHost() host = _FakeHost()
provider = GrokProvider(host, image="roboco-agent-grok:test") provider = GrokCliProvider(host, image="roboco-agent-grok:test")
with patch( with patch(
"asyncio.create_subprocess_exec", AsyncMock(return_value=_proc()) "asyncio.create_subprocess_exec", AsyncMock(return_value=_proc())
) as exec_mock: ) as exec_mock:
result = await provider.spawn(_config()) result = await provider.spawn(_config())
cmd = list(exec_mock.call_args.args) cmd = list(exec_mock.call_args.args)
# Gateway + operational env the grok image entrypoint consumes. # Gateway + operational env the grok-cli entrypoint + renderer consume.
assert "ROBOCO_MCP_CONFIG=/app/mcp-config.json" in cmd assert "ROBOCO_MCP_CONFIG=/app/mcp-config.json" in cmd
assert "ROBOCO_SYSTEM_PROMPT=/app/system-prompt.md" in cmd assert "ROBOCO_AGENT_ID=be-dev-1" in cmd # renderer computes per-role flags
# Tool restriction lives in the rendered opencode.json (opencode `tools`), assert "ROBOCO_AGENT_MODEL=grok-build" in cmd
# not a spawn env var — no ROBOCO_AGENT_TOOLS is injected.
assert not any(c.startswith("ROBOCO_AGENT_TOOLS=") for c in cmd)
# The opencode store is mounted so the orchestrator can read usage/cost
# back at finalize.
assert "/host/opencode/be-dev-1:/home/agent/.local/share/opencode" in cmd
# Identity wiring from the shared host helpers is present. # Identity wiring from the shared host helpers is present.
assert "ROBOCO_AGENT_TOKEN=hmac-be-dev-1" in cmd assert "ROBOCO_AGENT_TOKEN=hmac-be-dev-1" in cmd
# The image is the final docker-run argument. # The image is the final docker-run argument.
@@ -222,13 +220,38 @@ async def test_grok_spawn_wires_gateway_and_image_last() -> None:
assert host.removed == ["roboco-agent-be-dev-1"] assert host.removed == ["roboco-agent-be-dev-1"]
assert result == SpawnResult( assert result == SpawnResult(
instance_id="roboco-agent-be-dev-1", instance_id="roboco-agent-be-dev-1",
extra={"container_id": "cid", "model": "grok-build-0.1"}, extra={"container_id": "cid", "model": "grok-build"},
) )
async def test_grok_spawn_mounts_auth_when_present(_isolate_grok_auth: Path) -> None:
(_isolate_grok_auth / "auth.json").write_text("{}", encoding="utf-8")
host = _FakeHost()
provider = GrokCliProvider(host)
with patch(
"asyncio.create_subprocess_exec", AsyncMock(return_value=_proc())
) as exec_mock:
await provider.spawn(_config())
cmd = list(exec_mock.call_args.args)
expected = f"{_isolate_grok_auth / 'auth.json'}:/home/agent/.grok/auth.json:ro"
assert expected in cmd
async def test_grok_spawn_omits_auth_mount_when_absent() -> None:
# No auth.json in the (tmp) GROK_AUTH_HOST_PATH → no mount, no crash.
host = _FakeHost()
provider = GrokCliProvider(host)
with patch(
"asyncio.create_subprocess_exec", AsyncMock(return_value=_proc())
) as exec_mock:
await provider.spawn(_config())
cmd = list(exec_mock.call_args.args)
assert not any("/home/agent/.grok/auth.json" in c for c in cmd)
async def test_grok_spawn_prompt_is_injection_safe() -> None: async def test_grok_spawn_prompt_is_injection_safe() -> None:
host = _FakeHost() host = _FakeHost()
provider = GrokProvider(host) provider = GrokCliProvider(host)
nasty = "--model evil --session-id pwned" nasty = "--model evil --session-id pwned"
with patch( with patch(
"asyncio.create_subprocess_exec", AsyncMock(return_value=_proc()) "asyncio.create_subprocess_exec", AsyncMock(return_value=_proc())
@@ -240,18 +263,8 @@ async def test_grok_spawn_prompt_is_injection_safe() -> None:
assert nasty not in cmd assert nasty not in cmd
async def test_grok_spawn_defaults_base_url_when_route_blank() -> None:
provider = GrokProvider(_FakeHost())
with patch(
"asyncio.create_subprocess_exec", AsyncMock(return_value=_proc())
) as exec_mock:
await provider.spawn(_config(provider_base_url=None))
cmd = list(exec_mock.call_args.args)
assert "XAI_BASE_URL=https://api.x.ai/v1" in cmd
async def test_grok_spawn_raises_on_docker_failure() -> None: async def test_grok_spawn_raises_on_docker_failure() -> None:
provider = GrokProvider(_FakeHost()) provider = GrokCliProvider(_FakeHost())
with ( with (
patch( patch(
"asyncio.create_subprocess_exec", "asyncio.create_subprocess_exec",
@@ -262,105 +275,6 @@ async def test_grok_spawn_raises_on_docker_failure() -> None:
await provider.spawn(_config()) await provider.spawn(_config())
# ---------------------------------------------------------------------------
# Reasoning effort by role
# ---------------------------------------------------------------------------
def test_reasoning_effort_full_for_code_roles() -> None:
# developer / qa / pr_reviewer keep full reasoning (no variant).
assert _reasoning_effort_for("be-dev-1") is None
assert _reasoning_effort_for("be-qa") is None
assert _reasoning_effort_for("pr-reviewer-1") is None
def test_reasoning_effort_minimal_for_coordination_roles() -> None:
for slug in ("be-pm", "main-pm", "be-doc", "auditor", "product-owner"):
assert _reasoning_effort_for(slug) == "minimal", slug
def test_reasoning_effort_override(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setenv("ROBOCO_GROK_REASONING_EFFORT", "max")
assert _reasoning_effort_for("be-dev-1") == "max" # override wins over role
monkeypatch.setenv("ROBOCO_GROK_REASONING_EFFORT", "default")
assert _reasoning_effort_for("be-pm") is None # "default" => full reasoning
# ---------------------------------------------------------------------------
# Per-role opencode permissions (Claude-parity)
# ---------------------------------------------------------------------------
def test_edit_permission_allows_only_writer_roles() -> None:
assert _edit_permission_for("be-dev-1") == "allow"
assert _edit_permission_for("be-doc") == "allow"
for slug in ("be-qa", "pr-reviewer-1", "be-pm", "main-pm", "auditor"):
assert _edit_permission_for(slug) == "deny", slug
def test_bash_permission_allows_only_shell_roles() -> None:
for slug in ("be-dev-1", "be-doc", "be-pm", "main-pm"):
assert _bash_permission_for(slug) == "allow", slug
for slug in ("be-qa", "pr-reviewer-1", "auditor", "product-owner"):
assert _bash_permission_for(slug) == "deny", slug
def test_external_dir_permission_only_pr_reviewer() -> None:
assert _external_dir_permission_for("pr-reviewer-1") == "allow"
for slug in ("be-dev-1", "be-qa", "be-pm", "auditor"):
assert _external_dir_permission_for(slug) == "deny", slug
async def test_grok_spawn_sets_readonly_permissions_for_reviewer() -> None:
# A read-only reviewer (qa) gets edit=deny + bash=deny + external_dir=deny.
host = _FakeHost()
provider = GrokProvider(host)
with patch(
"asyncio.create_subprocess_exec", AsyncMock(return_value=_proc())
) as exec_mock:
await provider.spawn(_config(agent_id="be-qa"))
cmd = list(exec_mock.call_args.args)
assert "ROBOCO_GROK_EDIT_PERMISSION=deny" in cmd
assert "ROBOCO_GROK_BASH_PERMISSION=deny" in cmd
assert "ROBOCO_GROK_EXTERNAL_DIR_PERMISSION=deny" in cmd
async def test_grok_spawn_pr_reviewer_is_read_only_but_reads_scratch() -> None:
# The pr-reviewer never writes code (edit=deny) but reads its /tmp diff
# (external_directory=allow) — the one role that needs it.
host = _FakeHost()
provider = GrokProvider(host)
with patch(
"asyncio.create_subprocess_exec", AsyncMock(return_value=_proc())
) as exec_mock:
await provider.spawn(_config(agent_id="pr-reviewer-1"))
cmd = list(exec_mock.call_args.args)
assert "ROBOCO_GROK_EDIT_PERMISSION=deny" in cmd
assert "ROBOCO_GROK_EXTERNAL_DIR_PERMISSION=allow" in cmd
async def test_grok_spawn_sets_variant_for_minimal_role() -> None:
host = _FakeHost()
provider = GrokProvider(host)
with patch(
"asyncio.create_subprocess_exec", AsyncMock(return_value=_proc())
) as exec_mock:
await provider.spawn(_config(agent_id="be-pm")) # cell_pm -> minimal
cmd = list(exec_mock.call_args.args)
assert "ROBOCO_GROK_VARIANT=minimal" in cmd
async def test_grok_spawn_no_variant_for_dev_role() -> None:
host = _FakeHost()
provider = GrokProvider(host)
with patch(
"asyncio.create_subprocess_exec", AsyncMock(return_value=_proc())
) as exec_mock:
await provider.spawn(_config(agent_id="be-dev-1")) # developer -> full
cmd = list(exec_mock.call_args.args)
assert not any(c.startswith("ROBOCO_GROK_VARIANT=") for c in cmd)
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
# ClaudeCodeProvider # ClaudeCodeProvider
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
+3 -3
View File
@@ -1,6 +1,6 @@
"""The orchestrator routes only dedicated-backend providers through the registry. """The orchestrator routes only dedicated-backend providers through the registry.
GROK gets the GrokProvider; Anthropic / Ollama Cloud / self-hosted (and any GROK gets the GrokCliProvider; Anthropic / Ollama Cloud / self-hosted (and any
unknown value) return None so ``_spawn_container`` runs its built-in Claude Code unknown value) return None so ``_spawn_container`` runs its built-in Claude Code
path unchanged. This keeps the GROK addition purely additive. path unchanged. This keeps the GROK addition purely additive.
""" """
@@ -9,7 +9,7 @@ from __future__ import annotations
from unittest.mock import patch from unittest.mock import patch
from roboco.llm.providers import GrokProvider from roboco.llm.providers import GrokCliProvider
from roboco.runtime.orchestrator import AgentOrchestrator from roboco.runtime.orchestrator import AgentOrchestrator
@@ -21,7 +21,7 @@ def _make_orch() -> AgentOrchestrator:
def test_provider_for_grok_returns_grok_provider() -> None: def test_provider_for_grok_returns_grok_provider() -> None:
assert isinstance(_make_orch()._provider_for("grok"), GrokProvider) assert isinstance(_make_orch()._provider_for("grok"), GrokCliProvider)
def test_provider_for_anthropic_returns_none() -> None: def test_provider_for_anthropic_returns_none() -> None: