Files
8eb6e3eb60 fix(agents): run live Databricks discovery instead of the fallback list (#2890)
## Problem

The Databricks model dropdown offers a handful of stale models — and
there's no way to tell that list apart from the real one. The AI Gateway
exposes **66** chat/embedding endpoints on `block-lakehouse-production`,
but the picker was showing a short list that includes models the gateway
no longer serves and embedding endpoints that can't chat at all.

Three independent defects, all on the discovery path:

**1. Live discovery never ran for agents with no saved provider.**
`get_agent_models` gates every in-process discovery attempt on the
provider (`is_openai_compatible_provider` / `is_anthropic_provider` /
`is_databricks_provider`), reading it straight from `record.provider`.
That field is `null` for every agent record created before provider
persistence — and for any agent that inherits its provider from the
build. So all three gates saw `None`, no HTTP discovery ran, and the
request fell through to the `buzz-acp models` subprocess. On the
Databricks path that subprocess returns `discovery_failure_fallback` —
the small hardcoded `DATABRICKS_V2_KNOWN_MODELS` catalog — which the
frontend renders exactly like a live catalog. An internal DMG that bakes
`BUZZ_AGENT_PROVIDER=databricks_v2` and a `DATABRICKS_HOST` still got
the fallback.

**2. The fallback list couldn't represent the running model.**
When discovery genuinely fails, the picker should at minimum be able to
show what the agent is actually configured with. For `DatabricksV2` it
couldn't: the fallback returned only the hardcoded slate, so a model
like `databricks-gpt-5-5` wasn't selectable in its own picker.

**3. Embedding endpoints were offered as chat models.**
`databricks-bge-large-en` was selectable (visible in the dialog today).
The v2 endpoints payload carries no `task` or `state` field, so there is
nothing to filter on but the name.

## Changes

- **`effective_discovery_provider`** (new,
`desktop/src-tauri/src/commands/agent_models_env.rs`) — an explicit
provider (saved record value, or the create/edit dialog's current form
value) still always wins; when there is none, discovery falls back to
the runtime's own provider env var (`GOOSE_PROVIDER`,
`BUZZ_AGENT_PROVIDER`, …) read off the merged env, which by that point
already carries the baked build floor and the process env. Wired into
both `get_agent_models` and `discover_agent_models`.
`SavedAgentModelDiscoveryConfig` now carries `provider_env_var` from
`known_acp_runtime`, so each runtime reads *its own* key rather than a
shared guess.
- The relay-mesh branches in `discover_agent_models` deliberately keep
using `input.provider`: those key off a deliberate provider selection,
never a baked default.
- **Asserted vs inferred matters for missing credentials.** The OpenAI
and Anthropic gates error on a missing API key, while the Databricks
gate falls through; an inferred provider hitting the first two would
have replaced a working subprocess catalog with `config:
ANTHROPIC_API_KEY required` (`export GOOSE_PROVIDER=anthropic` is
goose's documented way to pick a provider, and it keeps the key in its
own keyring). So `effective_discovery_provider` returns a
`DiscoveryProvider` that remembers how the value was resolved, and
`required_env` only reports a missing credential for an asserted
provider. A wrong guess declines and lets the subprocess answer.
- **`is_chat_capable_endpoint`** (new,
`crates/buzz-agent/src/catalog.rs`) — applied in
`parse_v2_endpoints_page`. Drops `*embedding*` and segment-matched `bge`
/ `gte` endpoints; keeps everything unrecognised (fail-open, so a new
model family is never hidden). Segment matching is why it's `split('-')`
and not `contains`: a substring check would swallow legitimate names.
- **`discovery_failure_fallback`** for `Provider::DatabricksV2` now
leads with the configured model (deduped against the known slate,
blank-tolerant), so a failed discovery still yields a picker that can
show the running model. The configured model is trimmed once up front —
`resolve_model` doesn't trim, so a padded `DATABRICKS_MODEL` used to
slip past the dedupe and appear twice.
- **`sort_v2_endpoints_newest_first`** (new, second commit) — the
catalog is now ordered newest-first on each endpoint's
`created_timestamp`, ties broken by name. Previously Buzz sorted
nothing, so the gateway's own order reached the picker: it pages in two
phases (Databricks-managed, then workspace-created — the page token
decodes to `{"phase":"user"}`), each alphabetical, which buried
`databricks-claude-opus-5` 8th behind five older Claude endpoints and
`goose-claude-opus-5` — the newest endpoint in the catalog — 55th of 63.
Sorting in `fetch_v2_models` means both discovery paths inherit it with
no wire or type changes, and the combobox filter preserves incoming
order. Endpoints with an absent or unparseable timestamp sort last
rather than first, so a wire-shape change degrades to "unordered at the
bottom" instead of "shuffled to the top".
- The name tiebreak is load-bearing: eleven managed endpoints share one
placeholder timestamp (`1699610000000`), so without it their relative
order would vary between runs. That placeholder is also not always
accurate — a few genuinely recent endpoints
(`databricks-kimi-k2-7-code`, `databricks-llama-4-maverick`) land at the
bottom with the 2023 batch. The gateway offers nothing better to sort
on.
- Env/provider lookup helpers moved out of `agent_models.rs` into
`agent_models_env.rs`. This keeps the command module under the file-size
limit **without ratcheting the override up** — the existing 1079 entry
is untouched (file is now 1066 lines).

## Verification

Live against `block-lakehouse-production`, release build:

```
BUZZ_ACP_AGENT_COMMAND=$PWD/target/release/buzz-agent \
BUZZ_AGENT_PROVIDER=databricks_v2 \
DATABRICKS_HOST=https://block-lakehouse-production.cloud.databricks.com \
DATABRICKS_MODEL=databricks-gpt-5-5 \
./target/release/buzz-acp models --json
```

- before: 66 endpoints, including `databricks-bge-large-en`,
`databricks-gte-large-en`, `databricks-qwen3-embedding-0-6b`
- after: **63** endpoints, `[.models[] | select(.id |
test("embedding|-bge-|-gte-"))]` → `[]`

Top of the list after the sort commit:

```
goose-claude-opus-5                2026-07-24
databricks-claude-opus-5           2026-07-23
databricks-gemini-3-6-flash        2026-07-20
databricks-gemini-3-5-flash-lite   2026-07-20
databricks-inkling                 2026-07-14
```

Tests: 15 new (8 in `catalog.rs` — including the two-wire-shape
timestamp parse, the sort's tiebreak/no-timestamp cases, and the
padded-model dedupe — and 7 plus one assertion in
`agent_models_tests.rs`, 3 of them covering the asserted/inferred
credential split), two existing tests updated. `just check`, `just
test-unit`, and `just desktop-tauri-test` all pass (1636 desktop-tauri
tests, 274 buzz-agent lib tests).

Not run locally: the Docker-backed integration suite (`just test`) —
this diff touches neither `buzz-relay`, `buzz-db`, nor `buzz-auth`.

## Follow-ups (deliberately out of scope)

Two inference-path defects found while investigating, both reproduced
live against the gateway and both independent of discovery:

1. **Gemini thought signatures are dropped.** The gateway returns a bare
`thoughtSignature` on tool calls; the external-model serving endpoints
return it nested as `extra_content.google.thought_signature`. Neither
shape is round-tripped, so multi-turn tool use on `databricks-gemini-*`
fails with a 400 on the second turn.
2. **Array-shaped `content` is silently discarded.** Some models return
OpenAI `content` as a block array rather than a string; `parse_openai`'s
`str_field` returns `None` and the text is dropped.

The legacy `serving-endpoints` path does not work around either one, and
costs reasoning support on the GPT-5 family.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 00:56:18 +00:00
..

buzz-agent

Minimal, unbreakable ACP-compliant LLM agent. Stdio in, tool calls out. Non-streaming. No persistence. No cleverness.

ACP is the Agent Client Protocol — JSON-RPC 2.0 over stdio between a client (Zed, JetBrains, buzz-acp, …) and an agent. MCP is how the agent talks to its tools.

buzz-agent is the agent.

What It Is

        +--------+   stdio (JSON-RPC 2.0)   +---------------+
        | client | <----------------------> |  buzz-agent |
        +--------+        ACP frames        +---------------+
                                              │            │
                                              │            │ rmcp (stdio)
                                              │            ▼
                                              │       MCP servers
                                              │       (your tools)
                                              ▼
                                            HTTPS
                                              │
                                              ▼
                                  Anthropic Messages API
                                   or any OpenAI-compat
                                  (vLLM, llama.cpp, OpenRouter,
                                   Block Gateway, Ollama, …)

A client sends session/prompt. The agent loops: call the LLM → get tool calls → run them via MCP → feed results back → repeat. The loop terminates when the LLM stops asking for tools, the round cap is hit, or the client cancels.

The agent's output is its tool calls. Generated text is forwarded to the client as agent_message_chunk updates, but the real work happens in the tools. The LLM call is non-streaming — one HTTP POST, one response.

Quick Start

# Build
cargo build --release -p buzz-agent

# Run against Anthropic
BUZZ_AGENT_PROVIDER=anthropic \
ANTHROPIC_API_KEY=sk-ant-... \
ANTHROPIC_MODEL=claude-sonnet-4-5 \
  ./target/release/buzz-agent

# Or any OpenAI-compatible endpoint
BUZZ_AGENT_PROVIDER=openai \
OPENAI_COMPAT_API_KEY=sk-... \
OPENAI_COMPAT_MODEL=gpt-5 \
OPENAI_COMPAT_BASE_URL=https://api.openai.com/v1 \
  ./target/release/buzz-agent

# Or Databricks model serving via OAuth 2.0 PKCE
BUZZ_AGENT_PROVIDER=databricks \
DATABRICKS_HOST=https://dbc-...cloud.databricks.com \
DATABRICKS_MODEL=goose-claude-4-6-sonnet \
  ./target/release/buzz-agent

That's the whole setup. The agent reads JSON-RPC frames from stdin, writes them to stdout, and logs to stderr.

ACP Transcript

A complete round-trip. Lines starting with → are client→agent (stdin); ← are agent→client (stdout). Each line is one newline-terminated JSON value. Comments are not part of the wire.

// 1. Handshake.
→ {"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":1,"clientCapabilities":{}}}
← {"jsonrpc":"2.0","id":1,"result":{
    "protocolVersion":1,
    "agentCapabilities":{
      "loadSession":false,
      "promptCapabilities":{"image":false,"audio":false,"embeddedContext":false},
      "mcpCapabilities":{"http":false,"sse":false}
    },
    "agentInfo":{"name":"buzz-agent","version":"0.1.0"}
  }}

// 2. Open a session. The client passes the MCP servers to spawn.
→ {"jsonrpc":"2.0","id":2,"method":"session/new","params":{
    "cwd":"/tmp",
    "mcpServers":[{"name":"echo","command":"/usr/local/bin/echo-mcp","args":[],"env":[]}]
  }}
← {"jsonrpc":"2.0","id":2,"result":{"sessionId":"ses_a1b2c3d4e5f6a7b8"}}

// 3. Prompt. The agent loops until the LLM stops calling tools.
→ {"jsonrpc":"2.0","id":3,"method":"session/prompt","params":{
    "sessionId":"ses_a1b2c3d4e5f6a7b8",
    "prompt":[{"type":"text","text":"echo hello"}]
  }}

// 4. Agent emits tool_call (status: pending) — visible to the UI.
← {"jsonrpc":"2.0","method":"session/update","params":{
    "sessionId":"ses_a1b2c3d4e5f6a7b8",
    "update":{
      "sessionUpdate":"tool_call",
      "toolCallId":"toolu_01XYZ",
      "title":"echo__say",
      "kind":"other",
      "status":"pending",
      "rawInput":{"text":"hello"}
    }
  }}

// 5. Agent moves the call to in_progress, runs the MCP tool, then completed.
← {"jsonrpc":"2.0","method":"session/update","params":{
    "sessionId":"ses_a1b2c3d4e5f6a7b8",
    "update":{"sessionUpdate":"tool_call_update","toolCallId":"toolu_01XYZ","status":"in_progress"}
  }}
← {"jsonrpc":"2.0","method":"session/update","params":{
    "sessionId":"ses_a1b2c3d4e5f6a7b8",
    "update":{
      "sessionUpdate":"tool_call_update",
      "toolCallId":"toolu_01XYZ",
      "status":"completed",
      "content":[{"type":"content","content":{"type":"text","text":"hello"}}]
    }
  }}

// 8. The model sees the result, decides it's done, and the prompt resolves.
← {"jsonrpc":"2.0","id":3,"result":{"stopReason":"end_turn"}}

That's ACP. Three request methods (initialize, session/new, session/prompt), one inbound notification (session/cancel), and three outbound update variants (agent_message_chunk, tool_call, tool_call_update). The full server is hand-rolled in main.rs.

Configuration

Everything is environment variables. No flags, no config files. (We are a subprocess; subprocess config is environment.)

Variable Default Notes
BUZZ_AGENT_PROVIDER — Required. anthropic, openai, databricks, or databricks_v2. No implicit fallback — the agent errors at startup when this is unset.
ANTHROPIC_API_KEY — Required when provider=anthropic.
ANTHROPIC_MODEL — Required when provider=anthropic.
ANTHROPIC_BASE_URL https://api.anthropic.com
ANTHROPIC_API_VERSION 2023-06-01
OPENAI_COMPAT_API_KEY — Required when provider=openai.
OPENAI_COMPAT_MODEL — Required when provider=openai.
OPENAI_COMPAT_BASE_URL https://api.openai.com/v1 Point at vLLM, llama.cpp, OpenRouter, Ollama, etc.
OPENAI_COMPAT_API auto auto | chat | responses. auto picks Responses for *.openai.com, Chat Completions everywhere else.
DATABRICKS_HOST — Required when provider=databricks or provider=databricks_v2.
DATABRICKS_MODEL — Required when provider=databricks or provider=databricks_v2.
DATABRICKS_TOKEN — Optional static bearer escape hatch. If unset, Databricks uses browser OAuth + refresh cache.
BUZZ_AGENT_SYSTEM_PROMPT built-in Inline system prompt.
BUZZ_AGENT_SYSTEM_PROMPT_FILE — File path. Mutually exclusive with the above.
BUZZ_AGENT_MAX_ROUNDS 0 Tool-loop iteration cap. 0 = unlimited.
BUZZ_AGENT_MAX_OUTPUT_TOKENS 32768 Per LLM call. Headroom for large tool-call inputs (e.g. file writes via heredoc); Sonnet 4 / Opus 4 cap at 64K.
BUZZ_AGENT_MAX_CONTEXT_TOKENS 200000 Provider context window used by the handoff gate.
BUZZ_AGENT_MAX_HANDOFFS 10 Max context handoffs per session before falling back to truncation.
BUZZ_AGENT_LLM_TIMEOUT_SECS 240 Max seconds with no response bytes before abandoning an LLM call (per-read inactivity, not wall-clock).
BUZZ_AGENT_TOOL_TIMEOUT_SECS 660 Per-tool call timeout in seconds
BUZZ_AGENT_MAX_PARALLEL_TOOLS 8 Max concurrent tool calls per turn (1 = sequential)
BUZZ_AGENT_MAX_SESSIONS unlimited Max concurrent ACP sessions. Sessions are cheap; default has no cap.
BUZZ_AGENT_MAX_LINE_BYTES 4194304 4 MiB. Hard cap on inbound JSON-RPC frames.
BUZZ_AGENT_MAX_HISTORY_BYTES 1048576 1 MiB. Old turns are evicted past this.
BUZZ_AGENT_MAX_TOOL_RESULT_TEXT_BYTES 51200 50 KiB. Per-result cap on tool-output text; oversize is middle-elided (head + tail kept) with an inline marker. Images are exempt.

Providers

buzz-agent speaks a few HTTP dialects. Pick with BUZZ_AGENT_PROVIDER.

Provider BUZZ_AGENT_PROVIDER Endpoint (auto) Tested with
Anthropic anthropic POST {base}/v1/messages claude-sonnet-4-5, claude-opus-4
OpenAI openai POST {base}/responses gpt-5, gpt-5-mini, o4-mini, gpt-4o
vLLM openai POST {base}/chat/completions any tool-calling model
llama.cpp openai POST {base}/chat/completions any tool-calling GGUF
Ollama openai POST {base}/chat/completions llama3.1, qwen2.5-coder
OpenRouter openai POST {base}/chat/completions anything they route
Block Gateway openai POST {base}/chat/completions gpt-5, claude
Databricks databricks POST {host}/serving-endpoints/{model}/invocations goose-claude-4-6-sonnet
Databricks AI Gateway v2 databricks_v2 POST {host}/ai-gateway/{provider}/v1/... databricks-gpt-5-5, databricks-claude-opus-4-7

If BUZZ_AGENT_PROVIDER=anthropic is selected without ANTHROPIC_API_KEY, or BUZZ_AGENT_PROVIDER=openai is selected without OPENAI_COMPAT_API_KEY, the agent returns an error — there is no implicit fallback to another provider.

provider=openai speaks two HTTP dialects: the Responses API (/v1/responses, required for GPT-5 / o-series tool-calling on OpenAI's own service) and the Chat Completions API (/chat/completions, the broadly-supported OpenAI-compatible wire format).

By default (OPENAI_COMPAT_API=auto) the agent picks Responses when OPENAI_COMPAT_BASE_URL points at an *.openai.com host and Chat Completions everywhere else. Pin the choice explicitly with OPENAI_COMPAT_API=chat or OPENAI_COMPAT_API=responses for providers that diverge from the default (e.g. a Responses-compatible self-hosted gateway).

Provider is a Rust enum with one match in Llm::complete. There is no trait, no Box<dyn>, no async-trait. Adding a provider is a match arm and one body/parse pair in llm.rs.

MCP Servers

The client passes MCP server specs in session/new. The agent spawns each one as a stdio subprocess, calls tools/list, and merges everything into a single tool catalog the LLM sees. Tool names are namespaced as server__tool (double underscore separator). Bare tool names containing __ are rejected at registration.

Example: a single echo MCP server.

{
  "jsonrpc": "2.0",
  "id": 2,
  "method": "session/new",
  "params": {
    "cwd": "/work",
    "mcpServers": [
      {
        "name": "echo",
        "command": "/usr/local/bin/echo-mcp",
        "args": ["--mode", "stdio"],
        "env": [
          { "name": "ECHO_VERBOSE", "value": "1" }
        ]
      }
    ]
  }
}

Multiple servers: just add more entries. Tool calls fan out to the right server by namespace prefix.

Transport: stdio only. No HTTP, no SSE. We advertise this in agentCapabilities (mcpCapabilities.http: false, mcpCapabilities.sse: false); spec-compliant clients won't ask for what we don't have.

Security Model

The trust boundary is the operator who launched the agent. The harness, MCP server binaries, and API keys are all trusted. Untrusted input — model output, tool results, prompts — is bounded.

Boundary Mechanism
Stdout discipline Single-consumer mpsc channel feeding stdout. No two tasks can interleave bytes. All logs go to stderr.
MCP child env Whitelist (PATH, HOME, TERM, LANG, LC_ALL, TMPDIR) plus what the client explicitly passes. Your ANTHROPIC_API_KEY does not leak into MCP children.
MCP child lifetime Process group via setpgid(0,0) in pre_exec. On transport break or shutdown: killpg(SIGKILL). Grandchildren die too.
Server poisoning After a timeout or transport break, the offending server is marked dead. Future calls trigger a lazy restart with exponential backoff. Other servers keep working.
Frame size BUZZ_AGENT_MAX_LINE_BYTES (default 4 MiB). Oversize → connection killed.
LLM response size 16 MiB hard cap. Both Content-Length precheck and streaming-buffer cap.
Cancellation tokio::select! { biased; _ = cancel.changed() => ... } at every loop boundary. Cancel always wins the race.
Session isolation Unlimited concurrent sessions by default (configurable via BUZZ_AGENT_MAX_SESSIONS). One prompt per session at a time. Each session gets its own MCP servers.
tool_use ↔ tool_result pairing Encoded in the type system. Every ToolCall and ToolResult carries a provider_id: String (not Option).

Bounded Everything

Limit Default Where
Inbound JSON-RPC frame 4 MiB BUZZ_AGENT_MAX_LINE_BYTES
Single prompt 1 MiB MAX_PROMPT_BYTES
History window 1 MiB BUZZ_AGENT_MAX_HISTORY_BYTES
LLM response body 16 MiB MAX_LLM_RESPONSE_BYTES
LLM error body 4 KiB MAX_LLM_ERROR_BODY_BYTES
Tool result body (total, incl. images) 8 MiB MAX_TOOL_RESULT_BYTES
Tool result text 50 KiB BUZZ_AGENT_MAX_TOOL_RESULT_TEXT_BYTES
MCP servers / session 16 MAX_MCP_SERVERS
Tools / session 128 MAX_TOOLS_PER_SESSION
Tool description bytes 1 KiB MAX_DESCRIPTION_BYTES
Tool schema bytes 4 KiB MAX_SCHEMA_BYTES (oversize → replaced with {})
Tool calls per turn 64 MAX_TOOL_CALLS_PER_TURN
Loop rounds 0 (unlimited) BUZZ_AGENT_MAX_ROUNDS
LLM read inactivity timeout 240 s BUZZ_AGENT_LLM_TIMEOUT_SECS
Tool call timeout 660 s BUZZ_AGENT_TOOL_TIMEOUT_SECS

What This Is NOT

A short list, because the answer is mostly "no":

  • Not a framework. No plugins, no recipes, no slash commands, no modes. MCP servers can participate in agent lifecycle via hook tools (_Stop, _PostCompact), but these are advisory, fail-open, and budget-bounded — not a plugin system.
  • Not streaming. One non-streaming HTTP POST per round. The LLM's generated text is forwarded to the client as agent_message_chunk, but there is no token-level streaming.
  • Not persistent. Everything is in-memory, per-process. No SQLite. When context fills, the agent summarizes its own history and continues (context handoff). No external persistence.
  • Not an SDK. This is a binary. The protocol seam is stdin/stdout. Use it from any language.
  • Not a UI. No TUI, no web, no notifications. The client renders.
  • Not authenticated. API keys come from env. Use systemd, Docker secrets, or a wrapper.
  • Not networked MCP. Stdio transport only. No HTTP/SSE MCP transport.
  • Not load-able. No session/load. We advertise loadSession: false.
  • Not a router. No agent-to-agent, no fan-out, no orchestration. One model. One loop.

Concurrency model:

                  ┌──── reader task ──────────┐
                  │  (stdin → JSON-RPC → ...) │
                  │                           │
   stdin ─────────┤   dispatch                │
                  │     │                     │
                  │     ├── initialize        │  (sync reply)
                  │     ├── session/new       │  (sync reply)
                  │     ├── session/prompt ───┼─── spawn ──> prompt task
                  │     │                     │              │
                  │     ├── session/cancel ───┼─> watch::send│ (biased select wins)
                  │     │                     │              │
                  └───────────────────────────┘              │
                                                             │
                  ┌── writer task ────────────────┐          │
   stdout ────────┤  mpsc<WireMsg> consumer       │<─────────┘
                  │  (the only stdout writer)     │
                  └───────────────────────────────┘

One reader, one writer, up to 8 concurrent prompt tasks (one per session).

Building

cargo build --release -p buzz-agent

Testing

cargo test -p buzz-agent

Test strategy is real subprocess, no mocks:

  • Fake LLM — tests/fake_llm.rs and the helpers in tests/regressions.rs spin up a real tokio::net::TcpListener on port 0, parse Content-Length, and return scripted JSON. No HTTP mocking library.
  • Fake MCP server — tests/bin/fake_mcp.rs is a separate binary controlled by env vars: FAKE_MCP_HANG_INIT, FAKE_MCP_TOOL_DELAY, FAKE_MCP_SPAWN_GRANDCHILD, etc. Each fault path is a real process being abused.
  • Regression tests are the changelog. Each #[test] in regressions.rs is named for the bug it locks down: assistant_text_preserved_across_prompts, cancel_leaves_history_valid_for_next_prompt, mcp_init_timeout_kills_child, oversize_line_kills_connection. Read them in order to learn the protocol's failure modes.