Agent commits were authored by a raw 63-character npub, which makes `git
log`, `git blame`, and GitHub's author column effectively unreadable.
This uses the agent's display name for `user.name` instead, while
leaving the pubkey where it does real work.
## What changes
`build_git_env` in `crates/buzz-dev-mcp/src/shim.rs` now reads
`BUZZ_ACP_DISPLAY_NAME`, sanitizes it, and uses the result as
`user.name`. When the variable is absent or unusable it falls back to
`info.npub` — byte-identical to today's behavior.
`user.email`, `user.signingkey`, and the whole credential/signing block
are untouched. The pubkey is what NIP-98 auth, NIP-GS signing, and
contributor matching key on, and it stays in the email verbatim.
`crates/buzz-acp/src/lib.rs` forwards the variable into the dev-mcp
server's declared env, mirroring the existing `BUZZ_AUTH_TAG` block. It
reads `std::env::var` directly rather than going through `Config`, so
the variable is picked up whenever the process has it.
`crates/buzz-agent/src/mcp.rs` adds one `PASSTHROUGH_ENV` entry so ACP
clients that spawn `buzz-agent` without declaring the variable on the
wire still propagate it.
## Why a dedicated variable
`BUZZ_ACP_DISPLAY_NAME` is its own contract rather than a reuse of the
ACP session title. Commits outlive sessions: a session title is
per-session UI chrome and may be composed downstream into `Agent ·
#channel`, and if that composed form ever reached the env var, git
attribution would change silently with no test able to catch it. Git
identity gets a variable whose contract is "bare agent display name,
never channel-qualified."
Nothing writes it yet — a one-line Desktop write lands as a follow-up.
Until then `std::env::var` returns `Err`, the npub fallback fires, and
behavior is byte-for-byte current `main`.
## Sanitizing
Strip control characters, Unicode format characters, and angle brackets;
collapse whitespace runs, trim, cap at 80 characters (by `chars()`, so a
multi-byte name is never split mid-UTF-8).
Angle brackets go because git drops them silently rather than erroring:
`Duncan <evil@x.com>` renders as `Duncan evil@x.com <hex@relay>`. It
forges nothing, but it reads as though it might.
The empty result also has to cover more than literal emptiness. git's
`ident.c` treats a set of characters as "crud" — stripped from both
ends, and fatal when a name is *nothing but* those characters:
```
$ git -c user.name=';;' commit -m t
fatal: name consists only of disallowed characters: ;;
```
Verified against git 2.54.0 by committing with each ASCII byte 32..=126
as the entire `user.name`: exactly space, `"`, `'`, `,`, `:`, `;`, `<`,
`>`, `\` abort, plus all control characters (the predicate is `c <=
32`). `.` is not crud in this version, despite older lore. Names that
merely *contain* crud are fine — `O'Brien` and `Smith, Jr.` both commit
cleanly — so the check is "at least one non-crud character survives,"
not "no crud present." Without it, a display name of `;;` or `""` would
abort every commit that agent makes.
## Unicode format characters
`char::is_control` covers only category `Cc`. Category `Cf` — zero-width
spaces and joiners, bidi embedding and override marks, invisible math
operators, tag characters — is neither control, nor whitespace, nor git
crud, so those characters survived every one of the checks above. A
display name of nothing but U+200B ZERO WIDTH SPACE therefore satisfied
"at least one non-crud character survives" and git accepted the commit
with a visually blank author:
```
# pre-fix, BUZZ_ACP_DISPLAY_NAME set to two U+200B
$ git log -1 --format='%an' | xxd -p
e2808be2808b0a
```
Embedded marks were the other half: a trailing U+202E RIGHT-TO-LEFT
OVERRIDE reorders everything after it, so a stored author line renders
as something other than what it stores — the same confusion class the
angle-bracket filtering exists to prevent.
`is_unicode_format` rejects the whole `Cf` category rather than the
known-bad marks, because the boundary that matters is "invisible or
reorders text", not "the codepoint someone thought of". The 21 ranges
come from the UCD's `DerivedGeneralCategory.txt` (17.0.0), cross-checked
against Python's `unicodedata` (16.0.0); both yield exactly the same
set. They are inlined as a `matches!` rather than pulling in a
Unicode-tables crate for one predicate, and a test asserts both
endpoints of every range plus the codepoints immediately outside them —
including U+2065, which sits inside the U+2060 block but is unassigned
rather than `Cf`.
Filtering happens inside the existing per-word filter, so a format-only
name collapses to empty and falls out through the same `None` → npub
path as a crud-only name. No new fallback logic. And because filtering
precedes truncation, invisible padding cannot eat the 80-character
budget.
## NUL is handled one layer up
An interior NUL is a sibling constraint that cannot be fixed here: it
makes `Command::env` fail the entire spawn before this code runs, so it
has to die at the writer. #3028 establishes that pattern for the session
title in `resolve_session_title` via `filter(|c| !c.is_control())`, and
the Desktop follow-up that writes `BUZZ_ACP_DISPLAY_NAME` inherits it.
The shim sanitizer is a second line of defense for values that arrive
from somewhere other than Desktop.
## Verified end to end
Driving the real `buzz-dev-mcp` binary over stdio MCP and committing
inside its shimmed environment:
```
# BUZZ_ACP_DISPLAY_NAME="Duncan Idaho"
Duncan Idaho <dcfd242e...0f95@buzz.block.builderlab.xyz>
verify_exit=0
# BUZZ_ACP_DISPLAY_NAME unset
npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e...0f95@buzz.block.builderlab.xyz>
verify_exit=0
# BUZZ_ACP_DISPLAY_NAME=";;" (crud-only; would otherwise be fatal)
npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e...0f95@buzz.block.builderlab.xyz>
verify_exit=0
# BUZZ_ACP_DISPLAY_NAME=U+200B U+200B (format-only; would otherwise be blank)
npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e...0f95@buzz.block.builderlab.xyz>
verify_exit=0
# BUZZ_ACP_DISPLAY_NAME="Duncan" + U+202E (bidi override stripped)
Duncan <dcfd242e...0f95@buzz.block.builderlab.xyz>
verify_exit=0
# BUZZ_ACP_DISPLAY_NAME="Dun" + U+200B + "can" (zero-width removed, word not split)
Duncan <dcfd242e...0f95@buzz.block.builderlab.xyz>
verify_exit=0
```
Signature verification passes in every case — the signing identity is
unchanged.
`Related: #3028` — it establishes the Desktop-side env plumbing this
builds beside; the one-line Desktop follow-up that writes
`BUZZ_ACP_DISPLAY_NAME` alongside the session title ships after it
merges. Not a dependency: with the variable absent, `std::env::var`
returns `Err` and the npub fallback keeps current behavior exactly.
---------
Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
Signed-off-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
Co-authored-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
Uses MeshLLM built-in `mesh` collective intelligence / Mixture of Agents
when Buzz Auto sees two or more distinct physical models. With zero or
one model, Auto remains ordinary `auto`. This is an alternative way to
improve tool responses, accuracy, and resistance to hallucination when a
high-latency distributed mesh contains diverse models.
Models and members may come and go: collective routing enables only
after stable capacity, drops on confirmed contraction, and can recover
later. Mesh-specific failures retry once through ordinary Auto.
This update also pins MeshLLM to a v0.73.1-compatible backport of
[MeshLLM #1074](https://github.com/Mesh-LLM/mesh-llm/pull/1074), so
client-only Buzz nodes cannot enter model election or download a remote
provider model. Buzz preserves the selected local sharing model and
switches an existing client to sharing across a controlled app restart,
retaining one runtime and one `:9337` / `:3131` pair per machine.
Validation:
- Full local `just ci` passes on the cleaned branch.
- MeshLLM host-runtime suite: 1,568 passed, 0 failed; strict Clippy
passes.
- Buzz desktop Tauri suite with `mesh-llm`: 1,721 passed, 0 failed;
strict feature Clippy passes.
- Playwright covers client-to-share using the saved local model and no
destructive stop.
- Packaged two-machine testing proved single-model routing, dual-model
collective routing, tool-markup fallback, and runtime reuse.
- Packaged client-only recheck routed a real Mini Buzz turn through M5
while Mini stayed `is_client=true`, `is_host=false`, hosted no models,
and created no Gemma cache.
Builds on the recovery work merged in #2823; this PR does not duplicate
it.
---------
Signed-off-by: Michael Neale <michael.neale@gmail.com>
Signed-off-by: Tyler Longwell <tlongwell@block.xyz>
Co-authored-by: Michael Neale <michael.neale@gmail.com>
Co-authored-by: npub1qyvc0c5kl4gqv2fd97fsk46tu378sqgy35vc83rvgfwne90sel7s0ed67d <011987e296fd5006292d2f930b574be47c7801048d1983c46c425d3c95f0cffd@buzz.block.builderlab.xyz>
Co-authored-by: Tyler Longwell <tlongwell@block.xyz>
## Problem
The Databricks model dropdown offers a handful of stale models — and
there's no way to tell that list apart from the real one. The AI Gateway
exposes **66** chat/embedding endpoints on `block-lakehouse-production`,
but the picker was showing a short list that includes models the gateway
no longer serves and embedding endpoints that can't chat at all.
Three independent defects, all on the discovery path:
**1. Live discovery never ran for agents with no saved provider.**
`get_agent_models` gates every in-process discovery attempt on the
provider (`is_openai_compatible_provider` / `is_anthropic_provider` /
`is_databricks_provider`), reading it straight from `record.provider`.
That field is `null` for every agent record created before provider
persistence — and for any agent that inherits its provider from the
build. So all three gates saw `None`, no HTTP discovery ran, and the
request fell through to the `buzz-acp models` subprocess. On the
Databricks path that subprocess returns `discovery_failure_fallback` —
the small hardcoded `DATABRICKS_V2_KNOWN_MODELS` catalog — which the
frontend renders exactly like a live catalog. An internal DMG that bakes
`BUZZ_AGENT_PROVIDER=databricks_v2` and a `DATABRICKS_HOST` still got
the fallback.
**2. The fallback list couldn't represent the running model.**
When discovery genuinely fails, the picker should at minimum be able to
show what the agent is actually configured with. For `DatabricksV2` it
couldn't: the fallback returned only the hardcoded slate, so a model
like `databricks-gpt-5-5` wasn't selectable in its own picker.
**3. Embedding endpoints were offered as chat models.**
`databricks-bge-large-en` was selectable (visible in the dialog today).
The v2 endpoints payload carries no `task` or `state` field, so there is
nothing to filter on but the name.
## Changes
- **`effective_discovery_provider`** (new,
`desktop/src-tauri/src/commands/agent_models_env.rs`) — an explicit
provider (saved record value, or the create/edit dialog's current form
value) still always wins; when there is none, discovery falls back to
the runtime's own provider env var (`GOOSE_PROVIDER`,
`BUZZ_AGENT_PROVIDER`, …) read off the merged env, which by that point
already carries the baked build floor and the process env. Wired into
both `get_agent_models` and `discover_agent_models`.
`SavedAgentModelDiscoveryConfig` now carries `provider_env_var` from
`known_acp_runtime`, so each runtime reads *its own* key rather than a
shared guess.
- The relay-mesh branches in `discover_agent_models` deliberately keep
using `input.provider`: those key off a deliberate provider selection,
never a baked default.
- **Asserted vs inferred matters for missing credentials.** The OpenAI
and Anthropic gates error on a missing API key, while the Databricks
gate falls through; an inferred provider hitting the first two would
have replaced a working subprocess catalog with `config:
ANTHROPIC_API_KEY required` (`export GOOSE_PROVIDER=anthropic` is
goose's documented way to pick a provider, and it keeps the key in its
own keyring). So `effective_discovery_provider` returns a
`DiscoveryProvider` that remembers how the value was resolved, and
`required_env` only reports a missing credential for an asserted
provider. A wrong guess declines and lets the subprocess answer.
- **`is_chat_capable_endpoint`** (new,
`crates/buzz-agent/src/catalog.rs`) — applied in
`parse_v2_endpoints_page`. Drops `*embedding*` and segment-matched `bge`
/ `gte` endpoints; keeps everything unrecognised (fail-open, so a new
model family is never hidden). Segment matching is why it's `split('-')`
and not `contains`: a substring check would swallow legitimate names.
- **`discovery_failure_fallback`** for `Provider::DatabricksV2` now
leads with the configured model (deduped against the known slate,
blank-tolerant), so a failed discovery still yields a picker that can
show the running model. The configured model is trimmed once up front —
`resolve_model` doesn't trim, so a padded `DATABRICKS_MODEL` used to
slip past the dedupe and appear twice.
- **`sort_v2_endpoints_newest_first`** (new, second commit) — the
catalog is now ordered newest-first on each endpoint's
`created_timestamp`, ties broken by name. Previously Buzz sorted
nothing, so the gateway's own order reached the picker: it pages in two
phases (Databricks-managed, then workspace-created — the page token
decodes to `{"phase":"user"}`), each alphabetical, which buried
`databricks-claude-opus-5` 8th behind five older Claude endpoints and
`goose-claude-opus-5` — the newest endpoint in the catalog — 55th of 63.
Sorting in `fetch_v2_models` means both discovery paths inherit it with
no wire or type changes, and the combobox filter preserves incoming
order. Endpoints with an absent or unparseable timestamp sort last
rather than first, so a wire-shape change degrades to "unordered at the
bottom" instead of "shuffled to the top".
- The name tiebreak is load-bearing: eleven managed endpoints share one
placeholder timestamp (`1699610000000`), so without it their relative
order would vary between runs. That placeholder is also not always
accurate — a few genuinely recent endpoints
(`databricks-kimi-k2-7-code`, `databricks-llama-4-maverick`) land at the
bottom with the 2023 batch. The gateway offers nothing better to sort
on.
- Env/provider lookup helpers moved out of `agent_models.rs` into
`agent_models_env.rs`. This keeps the command module under the file-size
limit **without ratcheting the override up** — the existing 1079 entry
is untouched (file is now 1066 lines).
## Verification
Live against `block-lakehouse-production`, release build:
```
BUZZ_ACP_AGENT_COMMAND=$PWD/target/release/buzz-agent \
BUZZ_AGENT_PROVIDER=databricks_v2 \
DATABRICKS_HOST=https://block-lakehouse-production.cloud.databricks.com \
DATABRICKS_MODEL=databricks-gpt-5-5 \
./target/release/buzz-acp models --json
```
- before: 66 endpoints, including `databricks-bge-large-en`,
`databricks-gte-large-en`, `databricks-qwen3-embedding-0-6b`
- after: **63** endpoints, `[.models[] | select(.id |
test("embedding|-bge-|-gte-"))]` → `[]`
Top of the list after the sort commit:
```
goose-claude-opus-5 2026-07-24
databricks-claude-opus-5 2026-07-23
databricks-gemini-3-6-flash 2026-07-20
databricks-gemini-3-5-flash-lite 2026-07-20
databricks-inkling 2026-07-14
```
Tests: 15 new (8 in `catalog.rs` — including the two-wire-shape
timestamp parse, the sort's tiebreak/no-timestamp cases, and the
padded-model dedupe — and 7 plus one assertion in
`agent_models_tests.rs`, 3 of them covering the asserted/inferred
credential split), two existing tests updated. `just check`, `just
test-unit`, and `just desktop-tauri-test` all pass (1636 desktop-tauri
tests, 274 buzz-agent lib tests).
Not run locally: the Docker-backed integration suite (`just test`) —
this diff touches neither `buzz-relay`, `buzz-db`, nor `buzz-auth`.
## Follow-ups (deliberately out of scope)
Two inference-path defects found while investigating, both reproduced
live against the gateway and both independent of discovery:
1. **Gemini thought signatures are dropped.** The gateway returns a bare
`thoughtSignature` on tool calls; the external-model serving endpoints
return it nested as `extra_content.google.thought_signature`. Neither
shape is round-tripped, so multi-turn tool use on `databricks-gemini-*`
fails with a 400 on the second turn.
2. **Array-shaped `content` is silently discarded.** Some models return
OpenAI `content` as a block array rather than a string; `parse_openai`'s
`str_field` returns `None` and the text is dropped.
The legacy `serving-endpoints` path does not work around either one, and
costs reasoning support on the GPT-5 family.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>