mirror of
https://github.com/rennf93/roboco.git
synced 2026-08-03 07:23:24 +02:00
Agents with a write workspace (developer/product_owner/head_marketing/documenter) run with cwd = their git workspace clone. Claude Code launches each MCP server (flow/do/git-readonly/optimal/docs/search) and the SDK server as `uv run python -m ...` from that cwd. When the clone's uv.lock drifts from the baked image, `uv run` re-resolves and re-syncs /app/.venv against the clone's lock — a multi-minute stall on a cold wheel cache — so the servers never reach "connected": they sit at status="pending" and the agent gets ZERO gateway verbs. It then can't claim/commit/idle (all MCP verbs), its Stop is rejected, and it respawns in a loop redoing work it can't submit. UV_PROJECT_ENVIRONMENT pins the venv location but does NOT stop the cwd-relative resolve/sync (confirmed empirically on uv 0.11.1); `--no-sync` does, so the servers reuse the baked /app/.venv as-is and start instantly. The /app-cwd roles (qa/cell_pm/main_pm/auditor) were unaffected because their env already matches. - orchestrator.py: --no-sync on all 6 generated MCP servers - docker/scripts/sdk-startup-hook.sh: --no-sync on the agent_sdk.server launch - test_spawn_strict_mcp.py: assert every server's args start with run,--no-sync
58 lines
2.4 KiB
Bash
58 lines
2.4 KiB
Bash
#!/bin/bash
|
|
# SessionStart hook — starts the SDK server and prints the pre-rendered
|
|
# task briefing into the session so Claude doesn't burn its first turns
|
|
# on tool discovery (give_me_work → read files).
|
|
#
|
|
# The briefing is written by the orchestrator before the container spawns
|
|
# (_write_agent_briefing) and mounted read-only at /app/briefing.md.
|
|
|
|
SDK_PORT="${ROBOCO_SDK_PORT:-9000}"
|
|
AGENT_ID="${ROBOCO_AGENT_ID:-unknown}"
|
|
LOG_FILE="/tmp/sdk-server.log"
|
|
BRIEFING_FILE="/app/briefing.md"
|
|
PRECOMPACT_FILE="/tmp/roboco-precompact-${AGENT_ID}.md"
|
|
|
|
# #179: this hook runs with cwd = the agent's workspace, so a bare
|
|
# `uv run` resolves a cwd-relative `.venv` (≠ the baked /app/.venv),
|
|
# ignores VIRTUAL_ENV with a warning, and RE-SYNCS the full dependency
|
|
# set (torch/lancedb/pyarrow/scipy, ~350MB) into a fresh venv. On a cold
|
|
# uv wheel cache (first spawn after an image rebuild) that download takes
|
|
# minutes and the SDK/MCP layer never comes up before the agent reaps.
|
|
# Pin uv to the pre-baked image venv AND pass `--no-sync` on the run below.
|
|
# Pinning the env location alone is NOT sufficient: `uv run` still discovers
|
|
# the cwd project and re-syncs the pinned venv against the clone's (drifted)
|
|
# lock — that resync is the actual multi-minute stall. `--no-sync` skips it.
|
|
# (The orchestrator passes the same var + --no-sync to every MCP server in
|
|
# the generated mcp-config.json — keep both in sync.)
|
|
export UV_PROJECT_ENVIRONMENT=/app/.venv
|
|
|
|
# --- SDK bring-up ---------------------------------------------------------
|
|
if ! curl -sf "http://localhost:${SDK_PORT}/health" >/dev/null 2>&1; then
|
|
echo "[SDK] Starting for agent ${AGENT_ID} on port ${SDK_PORT}..."
|
|
nohup uv run --no-sync python -m roboco.agent_sdk.server > "$LOG_FILE" 2>&1 &
|
|
SDK_PID=$!
|
|
sleep 2
|
|
if curl -sf "http://localhost:${SDK_PORT}/health" >/dev/null 2>&1; then
|
|
echo "[SDK] Ready (PID: ${SDK_PID})"
|
|
else
|
|
echo "[SDK] Starting in background (PID: ${SDK_PID}, check ${LOG_FILE})"
|
|
fi
|
|
fi
|
|
|
|
# Reset budget/terminal counters at the start of every session.
|
|
curl -sf -m 2 -X POST "http://localhost:${SDK_PORT}/budget/reset" >/dev/null 2>&1 || true
|
|
|
|
# --- Briefing + PreCompact recovery --------------------------------------
|
|
# Compact restore comes FIRST so it's clear what this session is resuming.
|
|
if [[ -s "$PRECOMPACT_FILE" ]]; then
|
|
echo "### Resumed from compact"
|
|
cat "$PRECOMPACT_FILE"
|
|
echo
|
|
fi
|
|
|
|
if [[ -s "$BRIEFING_FILE" ]]; then
|
|
cat "$BRIEFING_FILE"
|
|
fi
|
|
|
|
exit 0
|