Files
buzz/crates
c405ad1d4b feat(agent): fix Anthropic prompt caching with Databricks (+ MCP proxy/TLS passthrough) (#3463)
> On 8 tasks matched by name across the two runs, cost fell $8.36 →
$1.77 (4.71×) and wall-clock 12,423 s → 1,085 s (11.45×).

## Summary

Two independent, self-contained fixes to `buzz-agent`/`buzz-acp`, split
out of the benchmark branch so they can land while the harness work
continues:

1. **Request and surface Anthropic prompt caching.** buzz never sent a
`cache_control` breakpoint, so on the Databricks Anthropic route
`cache_read_input_tokens` was **structurally always 0** and the ~10×
cache-read discount was never claimed. This teaches `anthropic_body()`
to mark the cacheable prefix, and plumbs the cache split end-to-end so
accounting can price it.
2. **Pass proxy + TLS-trust env into MCP tool subprocesses**, so agent
tools on a proxy-only host stop reporting a live network as offline.

## Why the caching gap matters

The Anthropic Messages API does **not** cache unless the request carries
a `cache_control` breakpoint, and the Databricks AI Gateway — a
third-party proxy in front of the model, in the same category as
Bedrock/Vertex — does **not** auto-cache (only the first-party Anthropic
API and Claude-on-AWS do zero-config caching). So every request was
billed cold.

Measured live against the Databricks gateway
(`databricks-claude-opus-5`, 2026-07-28), the same call with and without
a single `cache_control` marker:

| Run | `input_tokens` | `cache_creation` | `cache_read` | latency |
|---|---|---|---|---|
| No `cache_control`, two byte-identical calls | 121,625 | 0 | **0** |
~9.3 s |
| With one marker — cold (write) | 4 | 121,625 | 0 | 9.3 s |
| With one marker — warm (read) | 4 | 0 | **121,625** | **4.5 s** |

One marker moved 121,625 tokens from full-price input to a 0.1× cache
read and roughly halved latency (a clean, isolated ~2.07× prefill
speedup on this single-threaded microbenchmark). The gateway honours
`cache_control`; buzz simply never sent it.

At fleet scale this was a real budget item. Across matched
Terminal-Bench solo sweeps (89 tasks, `-n 20`, before the fix), the two
OpenAI-route models independently landed at ~86–87% cache reads — the
expected shape for an agentic loop, where system + tools + append-only
history repeat every turn — while the Anthropic route returned a hard 0%
on every receipt:

| Condition | Route | Input tokens | Cache reads | Cost | Cost if
uncached | Discount |
|---|---|---|---|---|---|---|
| luna (`gpt-5-6`) | OpenAI | 20,320,818 | **17.7M (87.0%)** | $6.96 |
$22.87 | **3.28×** |
| sol (`gpt-5-6`) | OpenAI | 22,312,290 | **19.2M (85.9%)** | $37.07 |
$123.35 | **3.33×** |
| opus (`claude-opus-5`) | Anthropic | 12,459,822 | **0 (0.0%)** |
$81.31 | $81.31 | **1.00×** |

Applying luna's measured 87% read rate to the opus token counts at list
prices (`input $5/M`, `cached_input $0.5/M`, `output $25/M`) puts the
opus run at **~$32.53 vs the $81.31 actually paid — a ~60% overspend on
those 49 trials (~$89 on a full sweep)**. That is an upper bound (it
prices every cached token at the 0.1× read rate and ignores the 1.25×
write premium), and the opus discount is structurally smaller than
luna/sol's because opus emits ~3.5× more uncacheable output per trial,
which sets a floor on what caching can recover.

There is also a plausible **second-order effect**: Databricks appears to
meter its per-minute rate limit on *uncached* input tokens, so the
missing cache also cost rate-limit headroom — the opus endpoint lost 63%
of its trials to fatal 429s while running alone at one-third of a GPT
endpoint's raw throughput. This is a hypothesis, not a proven mechanism
(the only zero-cache condition is also the only Anthropic endpoint), but
it is the reading that explains the throttling with one rule instead of
two.

## Post-fix results (provisional — first trials of an in-flight re-run)

On 8 tasks matched by name across the two runs, cost fell **$8.36 →
$1.77 (4.71×)** and wall-clock **12,423 s → 1,085 s (11.45×)**.

| Metric | before (`4a955a858`) | after (`3bef1f6a`) |
|---|---|---|
| Cache reads as % of input | **0.0%** | **78.7%** (still climbing
toward the ~86% steady state) |
| `cost_usd_no_cache_discount / cost_usd` | **1.00×** | **2.18×**
(tracking the projected ~2.5×) |
| Trials with a fatal 429 (same `-n 20`) | **63%** | **15–19%** |

To be clear about attribution: **~2× of that is the clean prefill saving
from caching itself**; the rest is second-order — cached requests burn
far less rate-limit budget, so they stall less and redo less destroyed
work. The 11.45× is a system-level result specific to this throttled
workspace, not a caching benchmark. Quality held (7/8 solved in each
run). A controlled low-`-n` A/B (neither arm hitting a 429), which the
`BUZZ_AGENT_PROMPT_CACHING` opt-out exists to enable, is still owed
before this becomes a published claim.

## What changed

### 1. Request caching (`llm.rs`, `config.rs`)

`anthropic_body()` emits ephemeral `cache_control` breakpoints, gated by
`BUZZ_AGENT_PROMPT_CACHING` (**default on**, `=0` to opt out):

- **Static prefix** — marker on the `system` block. Prefix order is
`tools → system → messages`, so this single marker caches **tools +
system** together. Byte-identical on every turn of a run, and survives a
context handoff (system/tools come from cfg/mcp, not `self.history`).
- **Rolling tail + leapfrog** — marker on the last block of the last
**two** messages. The append-only history re-reads the prior turn's
prefix from cache; marking two messages (not one) keeps consecutive
breakpoints inside Anthropic's **20-block lookback window** even as tool
parallelism rises, avoiding a silent full-price miss.

An empty system prompt stays a bare string (Anthropic rejects empty text
blocks), and below-threshold prefixes are silently not cached, so the
flag is safe on by default.

### 2. Surface the cache split end-to-end — the plumbing (`types.rs`,
`llm.rs`, `agent.rs`, `lib.rs`, `usage.rs`, `acp.rs`)

This is the part that makes gaps like the one above **visible** instead
of silent. A consumer that prices all of `input_tokens` at the full rate
can't tell a route that's caching from one that isn't — the total looks
right either way. So:

- `LlmResponse` gains `cached_input_tokens` (a **subset** of
`input_tokens`, never an addition); `parse_anthropic` / `parse_openai` /
`parse_responses` each populate it.
- A `usage_first()` helper reads the cache count wherever a provider
hides it — flat `cache_read_input_tokens` (Anthropic),
`prompt_tokens_details.cached_tokens` (OpenAI chat),
`input_tokens_details.cached_tokens` (Responses) — taking the **first
present value, never a sum**. Reading only flat keys is exactly why the
OpenAI route's nested `cached_tokens` had *also* been going unclaimed:
`prompt_tokens` is already inclusive, so the total looked correct while
the discount silently went unreported.
- The per-turn/per-session accumulators and the goose `usage_update`
payload now carry `accumulatedCachedInputTokens`; `buzz-acp`
deserializes it (`serde` default `0` for goose, which doesn't send it)
and logs `cached=<n>`.

### 3. Fix a Databricks MLflow-route double-count (`llm.rs`)

The Databricks MLflow route reports the flat Anthropic-spelled
`cache_read_input_tokens` *alongside* an already-inclusive
`prompt_tokens`, so the old code summed them and nearly doubled the
count — inflating both the context-budget gate and cost.
`openai_chat_input_tokens()` now reads `prompt_tokens` alone. Verified
on a live `databricks-glm-5-2` response where `prompt_tokens +
completion == total` proves inclusivity. (Anthropic's native route
genuinely *excludes* the cache fields and is still summed — the two
never collide, because `claude*` models route to the Anthropic path.)

### 4. Proxy + TLS-trust passthrough into MCP tools (`mcp.rs`) —
independent fix

`buzz-agent` `env_clear()`s each MCP child, and the allowlist carried no
proxy/TLS vars. On a proxy-only host that doesn't degrade the tools, it
**blinds** them: apt, curl, pip, git connect directly, the egress
firewall resets the socket, and the agent reports "Connection reset by
peer" — indistinguishable from a genuinely offline task. Adds both
spellings of `HTTP(S)_PROXY`/`NO_PROXY`/`ALL_PROXY` (curl/git read
lowercase; Go/Python read uppercase; libcurl ignores uppercase
`HTTP_PROXY`) plus `SSL_CERT_FILE`/`SSL_CERT_DIR` for TLS-terminating
proxies that present their own CA.

## Testing

- `cargo fmt --all -- --check`, `cargo clippy -p buzz-agent -p buzz-acp
--all-targets -- -D warnings` — clean.
- `cargo test -p buzz-agent -p buzz-acp` — **all green** (632 + 299 lib
tests plus integration suites, 0 failures). New tests cover: the three
breakpoints and the disabled/empty-system/single-message edge cases;
nested-vs-flat cache parsing for all three routes; the Databricks
inclusive-`prompt_tokens` fix; wire deserialization of
`accumulatedCachedInputTokens`; and the proxy/TLS passthrough allowlist.
- Pre-push lefthook suite green (branch-skew, rust-tests, test,
desktop-check/test/tauri).

## Relationship to the benchmark branch

These are the non-`benchmarks/` changes from
`benchmark/harness-accounting-and-solo`, lifted onto a clean base off
`main` so they can merge independently.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Atish Patel <atish@squareup.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-29 09:36:09 -04:00
..
2026-07-27 14:18:24 -04:00