fix(prompts): hoist ToolSearch activation to top of system prompt

Smoke-7: be-dev-1 hit "Edit exists but is not enabled in this context."
Claude Code v2.1.69+ defers built-in tools (Edit, Write, Read, etc.)
behind a ToolSearch call. Weak models (minimax-m2.7) skip soft
directives buried in the briefing.

Also: 4 role prompts (developer, cell_pm, main_pm, board) claimed
"no ToolSearch needed" — a lie that compounds the problem. The
manifest registers MCP tools; built-in tools are still deferred.

Fix: compose_prompt now prepends a tool-load directive layer as the
FIRST block in the system prompt. It names the exact ToolSearch call
the role needs:
- developer/documenter: Read, Bash, Grep, Glob, Task, TodoWrite, Edit, Write
- qa/pm/board: Read, Bash, Grep, Glob, Task, TodoWrite (no Edit/Write)

The directive includes the failure mode it prevents so the model
understands what skipping the call causes.

Updated role-prompt lines that lied about ToolSearch.

7 new tests pin: directive is the first block; developer/documenter
get Edit/Write; qa/pm don't; failure-mode message is present.
This commit is contained in:
Renn F
2026-05-15 03:14:43 +02:00
parent 417b8c5f29
commit 197b1576c3
6 changed files with 164 additions and 4 deletions
+1 -1
View File
@@ -16,7 +16,7 @@ If you find yourself reaching for `Bash git`, `Edit`, or any execution tool, sto
- **Product Owner**: product vision, feature priorities, accept/reject delivered work. - **Product Owner**: product vision, feature priorities, accept/reject delivered work.
- **Head of Marketing**: positioning, announcements, user feedback. - **Head of Marketing**: positioning, announcements, user feedback.
- **Auditor**: read everything, observe quality and compliance, escalate critical issues directly to CEO. - **Auditor**: read everything, observe quality and compliance, escalate critical issues directly to CEO.
- Your verb manifest is loaded — no `ToolSearch` needed. - Your verb manifest is loaded — MCP verbs are registered. Built-in tools (`Read`, `Bash`, `Task`, etc.) are deferred; run the `ToolSearch` call in the **First Action Required** block at the top of this prompt before your first read/bash.
## Your verbs ## Your verbs
+1 -1
View File
@@ -12,7 +12,7 @@ You merge what your developers submit (leaf PRs into your cell branch via `compl
- Your `task_id` (your cell-PM task) and `agent_id` are pre-baked into the gateway session. - Your `task_id` (your cell-PM task) and `agent_id` are pre-baked into the gateway session.
- Your team: backend / frontend / ux_ui. Your dev slugs: `be-dev-1`, `be-dev-2` (backend), `fe-dev-1`, `fe-dev-2` (frontend), `ux-dev-1`, `ux-dev-2` (UX). Your QA: `be-qa`/`fe-qa`/`ux-qa`. Your documenter: `be-doc`/`fe-doc`/`ux-doc`. - Your team: backend / frontend / ux_ui. Your dev slugs: `be-dev-1`, `be-dev-2` (backend), `fe-dev-1`, `fe-dev-2` (frontend), `ux-dev-1`, `ux-dev-2` (UX). Your QA: `be-qa`/`fe-qa`/`ux-qa`. Your documenter: `be-doc`/`fe-doc`/`ux-doc`.
- Your verb manifest is loaded — no `ToolSearch` needed. - Your verb manifest is loaded — MCP verbs are registered. Built-in tools (`Read`, `Bash`, `Task`, etc.) are deferred; run the `ToolSearch` call in the **First Action Required** block at the top of this prompt before your first read/bash.
- Workspace: `/data/workspaces/{project}/{team}/{your-slug}/` — but you have no `Edit`/`Write` permission; this is just where merge operations resolve. - Workspace: `/data/workspaces/{project}/{team}/{your-slug}/` — but you have no `Edit`/`Write` permission; this is just where merge operations resolve.
## Your verbs ## Your verbs
+1 -1
View File
@@ -10,7 +10,7 @@ You write code; you do not coordinate. If you find yourself thinking "let me als
- Your `task_id` and `agent_id` are pre-baked into the gateway session — every verb knows who you are. - Your `task_id` and `agent_id` are pre-baked into the gateway session — every verb knows who you are.
- Your workspace path: `/data/workspaces/{project}/{team}/{your-slug}/`. - Your workspace path: `/data/workspaces/{project}/{team}/{your-slug}/`.
- Your verb manifest is loaded — you do **not** need a `ToolSearch` call. - Your verb manifest is loaded — MCP verbs (`mcp__roboco-flow__*`, `mcp__roboco-do__*`) are already registered. Built-in tools (`Edit`, `Write`, `Read`, `Bash`, etc.) are deferred; the **First Action Required** block at the top of this prompt names the exact `ToolSearch` call to activate them. Do it once before your first edit/commit.
- Acceptance criteria, dev notes, parent context: call `evidence(task_id)` to fetch the task body and PR diff (if any). - Acceptance criteria, dev notes, parent context: call `evidence(task_id)` to fetch the task body and PR diff (if any).
## Your verbs ## Your verbs
+1 -1
View File
@@ -12,7 +12,7 @@ You merge what your Cell PMs submit (cell PRs into your root branch via `complet
- Your `task_id` (your root coordination task) and `agent_id` are pre-baked into the gateway session. - Your `task_id` (your root coordination task) and `agent_id` are pre-baked into the gateway session.
- Your cell-PM slugs: `be-pm`, `fe-pm`, `ux-pm`. Your team: `board`. Your channel: `main-pm-board`. - Your cell-PM slugs: `be-pm`, `fe-pm`, `ux-pm`. Your team: `board`. Your channel: `main-pm-board`.
- Your verb manifest is loaded — no `ToolSearch` needed. - Your verb manifest is loaded — MCP verbs are registered. Built-in tools (`Read`, `Bash`, `Task`, etc.) are deferred; run the `ToolSearch` call in the **First Action Required** block at the top of this prompt before your first read/bash.
- Workspace: `/data/workspaces/{project}/board/main-pm/` — but you have no `Edit`/`Write` permission; this is just where merge operations resolve. - Workspace: `/data/workspaces/{project}/board/main-pm/` — but you have no `Edit`/`Write` permission; this is just where merge operations resolve.
## Your verbs ## Your verbs
+65
View File
@@ -105,6 +105,70 @@ def _autogen_verbs_layer(prompts_path: Path, role: "AgentRole") -> str | None:
return _load_layer(prompts_path / "_generated" / f"{role_value}.md") return _load_layer(prompts_path / "_generated" / f"{role_value}.md")
# Built-in Claude Code tools each role's session needs at spawn time.
# Smoke-7 surfaced: be-dev-1 hit "Edit exists but is not enabled in this
# context" because Claude Code v2.1.69+ defers built-in tools behind a
# ToolSearch activation. Weak models skip the soft directive in the
# briefing, so we hoist the exact call into a top-of-system-prompt
# layer (highest-priority instruction the model sees).
#
# Read/Bash/Grep/Glob/Task/TodoWrite are needed by every role; Edit/Write
# only by roles that author code or docs.
_BUILTIN_TOOLS_COMMON: tuple[str, ...] = (
"Read",
"Bash",
"Grep",
"Glob",
"Task",
"TodoWrite",
)
_BUILTIN_TOOLS_AUTHORS: tuple[str, ...] = (*_BUILTIN_TOOLS_COMMON, "Edit", "Write")
_ROLE_BUILTIN_TOOLS: dict[str, tuple[str, ...]] = {
"developer": _BUILTIN_TOOLS_AUTHORS,
"documenter": _BUILTIN_TOOLS_AUTHORS,
"qa": _BUILTIN_TOOLS_COMMON,
"main_pm": _BUILTIN_TOOLS_COMMON,
"cell_pm": _BUILTIN_TOOLS_COMMON,
"product_owner": _BUILTIN_TOOLS_COMMON,
"head_marketing": _BUILTIN_TOOLS_COMMON,
"auditor": _BUILTIN_TOOLS_COMMON,
}
def _tool_load_directive_layer(role: "AgentRole") -> str | None:
"""Top-of-prompt instruction to activate deferred built-in tools.
Claude Code v2.1.69+ defers built-in tools (Edit, Write, Read, etc.)
behind a ToolSearch call to save context tokens. Without ToolSearch,
`Edit` returns "Edit exists but is not enabled in this context." Weak
models skip this if it's buried in the briefing; we hoist it into the
system prompt's very first block so it's the highest-priority
instruction. The call has no side effects beyond enabling the tools.
"""
role_value = role.value if hasattr(role, "value") else str(role)
tools = _ROLE_BUILTIN_TOOLS.get(role_value)
if not tools:
return None
tool_list = ",".join(tools)
return (
"# FIRST ACTION REQUIRED (mandatory)\n"
"\n"
"Before any other tool call you MUST run this ToolSearch once.\n"
"Built-in tools (Edit, Write, Read, etc.) are deferred until you\n"
"activate them. Skipping this step results in errors like\n"
'"Edit exists but is not enabled in this context" that waste your\n'
"tool-call budget and may leave the task in a stuck state.\n"
"\n"
"Copy this verbatim as your first action:\n"
"\n"
f'```\nToolSearch(query="select:{tool_list}")\n```\n'
"\n"
"Only after that call returns successfully should you proceed to\n"
"the role-specific workflow below."
)
def _lifecycle_layer(prompts_path: Path, role: "AgentRole") -> str | None: def _lifecycle_layer(prompts_path: Path, role: "AgentRole") -> str | None:
"""Load the canonical lifecycle fragment for this role. """Load the canonical lifecycle fragment for this role.
@@ -155,6 +219,7 @@ def compose_prompt(
parts: list[str] = [] parts: list[str] = []
for layer in ( for layer in (
_tool_load_directive_layer(role),
_lifecycle_layer(prompts_path, role), _lifecycle_layer(prompts_path, role),
_load_layer(prompts_path / "base.md"), _load_layer(prompts_path / "base.md"),
_role_layer(prompts_path, role), _role_layer(prompts_path, role),
@@ -0,0 +1,95 @@
"""Smoke-7: system prompt opens with a ToolSearch directive that activates
deferred built-in tools (Edit, Write, Read, ...).
Original bug: be-dev-1 called Edit and got "Edit exists but is not enabled
in this context" because Claude Code v2.1.69+ defers built-in tools behind
a ToolSearch call. The role prompts said "no ToolSearch needed" a lie
for built-in tools so weak models skipped the activation step.
Fix: compose_prompt now prepends a tool-load directive layer that names
the exact ToolSearch call for the role. It's the highest-priority block in
the system prompt so even weak models follow it.
"""
from __future__ import annotations
from roboco.agents.factories._base import compose_prompt
from roboco.models import AgentRole, Team
def _composed_prompt_for(role: AgentRole, team: Team | None = None) -> str:
"""Compose the prompt for a role and team."""
return compose_prompt(role, team, agent_slug="test-agent")
def test_developer_prompt_starts_with_tool_load_directive() -> None:
"""Developer system prompt begins with the ToolSearch activation block."""
prompt = _composed_prompt_for(AgentRole.DEVELOPER, Team.BACKEND)
assert prompt.startswith("# FIRST ACTION REQUIRED"), (
f"Developer prompt must lead with the tool-load directive. "
f"Got first 80 chars: {prompt[:80]!r}"
)
def test_developer_directive_names_edit_and_write() -> None:
"""Developer ToolSearch call lists Edit + Write (the smoke-7 wedge)."""
prompt = _composed_prompt_for(AgentRole.DEVELOPER, Team.BACKEND)
assert 'ToolSearch(query="select:' in prompt
# Find the ToolSearch line
line = next(line for line in prompt.splitlines() if "ToolSearch(query=" in line)
assert "Edit" in line, f"Developer ToolSearch missing Edit: {line!r}"
assert "Write" in line
def test_documenter_directive_names_edit_and_write() -> None:
"""Documenter needs Edit/Write too — they author docs."""
prompt = _composed_prompt_for(AgentRole.DOCUMENTER, Team.BACKEND)
line = next(line for line in prompt.splitlines() if "ToolSearch(query=" in line)
assert "Edit" in line
assert "Write" in line
def _tool_list_from_directive(prompt: str) -> list[str]:
"""Extract the comma-separated tool list from the ToolSearch call."""
line = next(line for line in prompt.splitlines() if "ToolSearch(query=" in line)
# `ToolSearch(query="select:Read,Bash,...")` — pull out the names.
start = line.find("select:") + len("select:")
end = line.find('"', start)
return line[start:end].split(",")
def test_qa_directive_excludes_edit_and_write() -> None:
"""QA reads but doesn't author — directive must NOT activate Edit/Write."""
tools = _tool_list_from_directive(_composed_prompt_for(AgentRole.QA, Team.BACKEND))
assert "Edit" not in tools, f"QA must not activate Edit: tools={tools}"
assert "Write" not in tools
assert "Read" in tools
assert "Bash" in tools
def test_pm_directives_exclude_edit_and_write() -> None:
"""PMs coordinate; they don't author code. No Edit/Write activation."""
for role in (AgentRole.MAIN_PM, AgentRole.CELL_PM):
tools = _tool_list_from_directive(_composed_prompt_for(role))
assert "Edit" not in tools, f"{role.value} must not activate Edit: {tools}"
assert "Write" not in tools
def test_directive_explains_why_it_matters() -> None:
"""The block must mention 'Edit exists but is not enabled' so the agent
understands what skipping the call causes."""
prompt = _composed_prompt_for(AgentRole.DEVELOPER, Team.BACKEND)
assert "Edit exists but is not enabled" in prompt
def test_directive_is_first_layer_before_lifecycle() -> None:
"""First Action block precedes the lifecycle and base layers."""
prompt = _composed_prompt_for(AgentRole.DEVELOPER, Team.BACKEND)
first_idx = prompt.find("# FIRST ACTION REQUIRED")
lifecycle_idx = prompt.find("Lifecycle")
base_idx = prompt.find("RoboCo Agent — Base")
assert first_idx == 0
if lifecycle_idx > -1:
assert first_idx < lifecycle_idx
if base_idx > -1:
assert first_idx < base_idx