fix(agent): launch agent uv-run subprocesses with --no-sync

Agents with a write workspace (developer/product_owner/head_marketing/documenter)
run with cwd = their git workspace clone. Claude Code launches each MCP server
(flow/do/git-readonly/optimal/docs/search) and the SDK server as
`uv run python -m ...` from that cwd. When the clone's uv.lock drifts from the
baked image, `uv run` re-resolves and re-syncs /app/.venv against the clone's
lock — a multi-minute stall on a cold wheel cache — so the servers never reach
"connected": they sit at status="pending" and the agent gets ZERO gateway
verbs. It then can't claim/commit/idle (all MCP verbs), its Stop is rejected,
and it respawns in a loop redoing work it can't submit.

UV_PROJECT_ENVIRONMENT pins the venv location but does NOT stop the cwd-relative
resolve/sync (confirmed empirically on uv 0.11.1); `--no-sync` does, so the
servers reuse the baked /app/.venv as-is and start instantly. The /app-cwd roles
(qa/cell_pm/main_pm/auditor) were unaffected because their env already matches.

- orchestrator.py: --no-sync on all 6 generated MCP servers
- docker/scripts/sdk-startup-hook.sh: --no-sync on the agent_sdk.server launch
- test_spawn_strict_mcp.py: assert every server's args start with run,--no-sync
This commit is contained in:
Renn F
2026-06-15 22:43:05 +02:00
parent f443a60d29
commit 25aed51c04
3 changed files with 29 additions and 9 deletions
+7 -4
View File
@@ -18,15 +18,18 @@ PRECOMPACT_FILE="/tmp/roboco-precompact-${AGENT_ID}.md"
# set (torch/lancedb/pyarrow/scipy, ~350MB) into a fresh venv. On a cold
# uv wheel cache (first spawn after an image rebuild) that download takes
# minutes and the SDK/MCP layer never comes up before the agent reaps.
# Pin uv to the pre-baked image venv so it starts instantly regardless
# of cwd. (The orchestrator sets the same var in every MCP server's env
# in the generated mcp-config.json — keep both in sync.)
# Pin uv to the pre-baked image venv AND pass `--no-sync` on the run below.
# Pinning the env location alone is NOT sufficient: `uv run` still discovers
# the cwd project and re-syncs the pinned venv against the clone's (drifted)
# lock — that resync is the actual multi-minute stall. `--no-sync` skips it.
# (The orchestrator passes the same var + --no-sync to every MCP server in
# the generated mcp-config.json — keep both in sync.)
export UV_PROJECT_ENVIRONMENT=/app/.venv
# --- SDK bring-up ---------------------------------------------------------
if ! curl -sf "http://localhost:${SDK_PORT}/health" >/dev/null 2>&1; then
echo "[SDK] Starting for agent ${AGENT_ID} on port ${SDK_PORT}..."
nohup uv run python -m roboco.agent_sdk.server > "$LOG_FILE" 2>&1 &
nohup uv run --no-sync python -m roboco.agent_sdk.server > "$LOG_FILE" 2>&1 &
SDK_PID=$!
sleep 2
if curl -sf "http://localhost:${SDK_PORT}/health" >/dev/null 2>&1; then