Commit Graph
2003 Commits
Author SHA1 Message Date
npub1223z34hd7vtwc6qj4s7flsxkj644nlre2nthu7lrrmkumhu3xddsrx9r6wandTaylor Ho 7b87864d55 test(desktop): align onboarding backup integration coverage
Reconcile focused backup assertions with Tyler’s post-main behavior: default generator surfaces, optional verification copy, hidden saved paths, and fresh re-entry state.

Co-authored-by: Taylor Ho <taylorkmho@gmail.com>
Signed-off-by: Taylor Ho <taylorkmho@gmail.com>
2026-07-30 00:01:41 -07:00
Taylor Hoandnpub1223z34hd7vtwc6qj4s7flsxkj644nlre2nthu7lrrmkumhu3xddsrx9r6w 6261595c04 feat(desktop): polish backup flow completion
- Focus the backup password input on entry and submit downloads with Enter, including queued encryption.
- Remove the pre-download discard confirmation so Back immediately resets and exits the optional flow.
- Give successful verification dedicated page copy and a primary Finish action.
- Add an explicit visibility toggle for the unlocked private key while keeping key material masked until requested.
- Preserve the existing verification success message with a minimal, lightly blurred key treatment.

Co-authored-by: Taylor Ho <taylorkmho@gmail.com>
Signed-off-by: Taylor Ho <taylorkmho@gmail.com>
2026-07-29 23:52:33 -07:00
Taylor Hoandnpub1223z34hd7vtwc6qj4s7flsxkj644nlre2nthu7lrrmkumhu3xddsrx9r6w 9a7faa9425 feat(desktop): refine optional backup verification
- Present backup verification as an optional post-download step with contextual headings and a Skip for now action.
- Match the verification password field to the backup password input and move its primary action into the onboarding footer.
- Add an upward-pulsing file-to-password illustration with reduced-motion support and a stronger selected-file affordance.
- Remove saved filesystem paths from the verification view to avoid exposing local details.
- Reset encrypted-backup state whenever users re-enter the creation flow.
- Confirm before leaving only when an unfinished pre-download password or encryption would be discarded.

Co-authored-by: Taylor Ho <taylorkmho@gmail.com>
Signed-off-by: Taylor Ho <taylorkmho@gmail.com>
2026-07-29 23:52:32 -07:00
Taylor Hoandnpub1223z34hd7vtwc6qj4s7flsxkj644nlre2nthu7lrrmkumhu3xddsrx9r6w 09bc460196 fix(desktop): wait for explicit backup download
- Keep eager NIP-49 encryption silent without changing the password form or advancing the backup flow.
- Commit the encrypted payload for saving only after the user clicks the download button.
- Preserve queued downloads during encryption and show the rotating progress ticker until the payload is ready.
- Update reducer coverage for silent background encryption while leaving end-to-end tests unchanged.

Co-authored-by: Taylor Ho <taylorkmho@gmail.com>
Signed-off-by: Taylor Ho <taylorkmho@gmail.com>
2026-07-29 23:52:31 -07:00
Taylor Hoandnpub1223z34hd7vtwc6qj4s7flsxkj644nlre2nthu7lrrmkumhu3xddsrx9r6w bfc79dfb59 fix(desktop): scope inverted buttons to backup flows
- Restore the standard primary CTA treatment for general onboarding steps.
- Add a dedicated inverted CTA style for dark backup-security surfaces.
- Apply the white button treatment only to backup creation, file selection, verification, and password-reset actions.

Co-authored-by: Taylor Ho <taylorkmho@gmail.com>
Signed-off-by: Taylor Ho <taylorkmho@gmail.com>
2026-07-29 23:52:31 -07:00
Taylor Hoandnpub1223z34hd7vtwc6qj4s7flsxkj644nlre2nthu7lrrmkumhu3xddsrx9r6w 4a8aec83ff fix(desktop): invert onboarding primary buttons
- Give shared in-step primary actions a white background with dark labels.
- Add coordinated hover colors that preserve contrast across onboarding surfaces.
- Update the shared CTA documentation to describe the inverted treatment.

Co-authored-by: Taylor Ho <taylorkmho@gmail.com>
Signed-off-by: Taylor Ho <taylorkmho@gmail.com>
2026-07-29 23:52:30 -07:00
Taylor Hoandnpub1223z34hd7vtwc6qj4s7flsxkj644nlre2nthu7lrrmkumhu3xddsrx9r6w 6b624e9771 feat(desktop): illustrate encrypted backup flow
- Add a key-to-password-to-lock timeline to the onboarding backup form.
- Animate four connector dots in a downward pulse with reduced-motion support.
- Preserve the compact settings variant and hide decoration in short windows.
- Keep balanced spacing and semantic security-theme colors across the illustration.

Co-authored-by: Taylor Ho <taylorkmho@gmail.com>
Signed-off-by: Taylor Ho <taylorkmho@gmail.com>
2026-07-29 23:52:30 -07:00
Taylor Hoandnpub1223z34hd7vtwc6qj4s7flsxkj644nlre2nthu7lrrmkumhu3xddsrx9r6w 8aaa80daee fix(desktop): darken backup password field controls
- Increase contrast for the spotlight password field border and focus ring against its white surface.
- Apply dark default and hover colors to the password generator and visibility controls.
- Remove the obsolete dark-surface icon treatment from the light password field.

Co-authored-by: Taylor Ho <taylorkmho@gmail.com>
Signed-off-by: Taylor Ho <taylorkmho@gmail.com>
2026-07-29 23:52:29 -07:00
Taylor Hoandnpub1223z34hd7vtwc6qj4s7flsxkj644nlre2nthu7lrrmkumhu3xddsrx9r6w 1432a32644 fix(desktop): improve backup password control contrast
- Give the spotlight password field a white background, subtle border, rounded corners, and dark password and placeholder text.
- Keep the saved-password mask legible against the light input surface.
- Render password generator controls in the standard popover surface instead of the dark textured security treatment.

Co-authored-by: Taylor Ho <taylorkmho@gmail.com>
Signed-off-by: Taylor Ho <taylorkmho@gmail.com>
2026-07-29 23:52:29 -07:00
Taylor Hoandnpub1223z34hd7vtwc6qj4s7flsxkj644nlre2nthu7lrrmkumhu3xddsrx9r6w 72e4e319a4 fix(desktop): clarify keychain onboarding copy
- Use the resolved identity storage backend to name the system keychain in the initial backup guidance.
- Keep local-file and unknown storage states accurate without exposing boot-time identity creation.
- Explain in the backup options card that the computer may request the user's password when Buzz needs to read the key.

Co-authored-by: Taylor Ho <taylorkmho@gmail.com>
Signed-off-by: Taylor Ho <taylorkmho@gmail.com>
2026-07-29 23:52:28 -07:00
Taylor Hoandnpub1223z34hd7vtwc6qj4s7flsxkj644nlre2nthu7lrrmkumhu3xddsrx9r6w e90a2062bb feat(desktop): refine backup onboarding security flow
- Add generated compact light and dark noisy surfaces with reusable tone and size options for cards, popovers, and confirmation dialogs.
- Keep the dark backup option and password panels unboxed, reduce their padding, and top-align parallel card titles for a clearer visual hierarchy.
- Explain identity keys and password-protected backup files in plain language while preserving distinct device, password-manager, and portable-backup choices.
- Normalize onboarding primary, secondary, and icon controls and add reduced-motion-safe transitions between backup creation and verification stages.
- Update focused onboarding E2E coverage for current navigation, compact texture contracts, title alignment, and narrow viewport geometry.

Co-authored-by: Taylor Ho <taylorkmho@gmail.com>
Signed-off-by: Taylor Ho <taylorkmho@gmail.com>
2026-07-29 23:52:28 -07:00
Taylor Hoandnpub1223z34hd7vtwc6qj4s7flsxkj644nlre2nthu7lrrmkumhu3xddsrx9r6w 38edd28336 feat(desktop): explain system keychain protection during onboarding
- Track whether the active identity is stored in the system keyring, local fallback file, environment, or ephemeral memory.
- Expose storage metadata through the Tauri identity response and shared TypeScript identity model.
- Add a responsive first information card beside the two backup actions without revealing boot-time identity creation.
- Show keychain password education only when secure storage succeeded and accurate neutral copy for fallback states.
- Extend identity persistence tests and split storage types into focused modules to preserve file-size limits.

Co-authored-by: Taylor Ho <taylorkmho@gmail.com>
Signed-off-by: Taylor Ho <taylorkmho@gmail.com>
2026-07-29 23:52:27 -07:00
Taylor Hoandnpub1223z34hd7vtwc6qj4s7flsxkj644nlre2nthu7lrrmkumhu3xddsrx9r6w ca754a5daa fix(desktop): decouple backup surface navigation
- Replace the password-backup surface's onboarding Skip and Next actions with a dedicated Back control.
- Keep Backup key and Back in the existing bottom-docked CTA location while preserving inline settings behavior.
- Route Back to Backup options with the correct backward transition instead of advancing onboarding.

Co-authored-by: Taylor Ho <taylorkmho@gmail.com>
Signed-off-by: Taylor Ho <taylorkmho@gmail.com>
2026-07-29 23:52:27 -07:00
Taylor Hoandnpub1223z34hd7vtwc6qj4s7flsxkj644nlre2nthu7lrrmkumhu3xddsrx9r6w e8c7f15aee feat(desktop): integrate backup options into identity onboarding
- Replace the standalone password-backup step with copy and encrypted-file options inside the identity-key stage.
- Add a midnight security theme, return-to-onboarding control, and downward transition back into the core yellow flow.
- Restore seven-page onboarding pagination and keep backup progress available when users revisit the security view.
- Update onboarding E2E coverage and screenshots for the inline options, direct setup path, and password-backup return behavior.

Co-authored-by: Taylor Ho <taylorkmho@gmail.com>
Signed-off-by: Taylor Ho <taylorkmho@gmail.com>
2026-07-29 23:52:20 -07:00
npub1qyvc0c5kl4gqv2fd97fsk46tu378sqgy35vc83rvgfwne90sel7s0ed67dandTyler Longwell 6a24905a2e Merge remote-tracking branch 'origin/main' into eva/nip49-local-backup
Conflict resolutions:
- relay.rs: took main's submit refactor (submit_signed_event_at_with_keys
  funnel replaces the removed submit_signed_event); the persona flush loop
  on main already publishes through the guarded funnel.
- egress_guard boundary 3 retargeted to the pre-signed funnel entry, and
  the /events inventory rows updated for main's relay.rs / sharing.rs
  layout (sharing.rs site is its in-file mock-relay test fixture; its
  production publish goes through the guarded boundary-1 funnel).
- personas/snapshot/import.rs + e2eBridge.ts: kept both sides' additions.

Co-authored-by: Tyler Longwell <tlongwell@block.xyz>
Signed-off-by: Tyler Longwell <tlongwell@block.xyz>
2026-07-29 18:27:39 -04:00
f95fdc1a10 feat(agent,acp): wire provider total_tokens through NIP-AM publish chain (#3593)
## What

Wires genuine provider-reported `total_tokens` through the full
buzz-agent → buzz-acp publish chain so kind-44200 events carry real
per-turn and cumulative totals for OpenAI-backed models, while
preserving all existing behaviour for Anthropic and external harnesses
(goose, claude-code).

## Why

Live prod data showed 0 of 1,934 archived reports carry `totalTokens`.
Both hardcoded `total_tokens: None` in `pool.rs` and the absent field in
`buzz-agent`'s parser are root causes. This is the backend half of a
two-track fix; the display-fallback half lands in
[#2035](https://github.com/block/buzz/pull/2035).

## Changes

**`crates/buzz-agent/src/types.rs`**
- Added `total_tokens: Option<u64>` to `LlmResponse` with an explicit
doc comment that NIP-AM forbids deriving it.
- Added `TurnTotalState` enum (`Unseen | Exact(u64) | Unknown`) with
`fold()` and `exact_value()` — the tri-state accumulator that
distinguishes not-yet-observed from permanently poisoned.

**`crates/buzz-agent/src/llm.rs`**
- `parse_responses` and `parse_openai`: read `usage.total_tokens` from
OpenAI Chat Completions (including Databricks routes) and the Responses
API via `sum_usage`.
- Anthropic: explicit `total_tokens: None` — no genuine total available;
NIP-AM forbids summing categories.

**`crates/buzz-agent/src/agent.rs`**
- Added `turn_total_state: &'a mut TurnTotalState` to `RunCtx`.
- Fold `response.total_tokens` into the accumulator after each
usage-bearing response; non-usage-bearing responses (keepalive/stream
frames) do not poison.

**`crates/buzz-agent/src/lib.rs`**
- Added `accumulated_total_state: TurnTotalState` to `Session` (default
`Unseen`).
- Per-turn state passed to `RunCtx`, folded into session cumulative
after each turn.
- Emits `accumulatedTotalTokens` in `usage_update` only when cumulative
is `Exact(n)`.

**`crates/buzz-acp/src/usage.rs`**
- Added `accumulated_total_tokens: Option<u64>` (serde default) to
`UsageUpdatePayload` — optional for goose compat.
- Added `last_total: Option<u64>` to `SessionState`.
- Added `turn_total_tokens` and `cumulative_total_tokens` to `TurnUsage`
(field-local — never affect `delta_reliable`).
- Derive turn-total delta only when prev and current are both `Some` and
monotonic; absence, decrease, or no baseline leaves only the total delta
null without touching input/output reliability.

**`crates/buzz-acp/src/pool.rs`**
- Replaced both hardcoded `total_tokens: None` in
`publish_agent_turn_metric` with `usage.turn_total_tokens` and
`usage.cumulative_total_tokens`.

## Tests

20 new tests across the four touched files:

| File | Tests |
|------|-------|
| `types.rs` | `TurnTotalState` fold, accumulation, exact_value, default
(7 tests) |
| `llm.rs` | Chat present/absent, Responses present/absent, Anthropic
always-None (5 tests) |
| `usage.rs` | First turn no baseline, second-turn delta, cumulative
decrease (field-local), current absent, goose-shaped deserialization,
baseline absent (6 tests) |
| `pool.rs` | Exact turn+cumulative mapping, null totals never derived
(2 tests) |

`cargo test -p buzz-acp -p buzz-agent` — all passing, 0 failures.

## Scope

Boundary: `crates/buzz-agent/**` + `crates/buzz-acp/**` only. Desktop
unchanged.
`costUsd` explicitly out of scope.

Related: [#2035](https://github.com/block/buzz/pull/2035)

---------

Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
Co-authored-by: npub1g8493u0xfsjrvflg4n08ezd7vec99mnwzlv0qgwpr9d7gvjwhuzqx59rhw <41ea58f1e64c243627e8acde7c89be667052ee6e17d8f021c1195be4324ebf04@buzz.block.builderlab.xyz>
2026-07-29 18:15:30 -04:00
WesandGitHub 3e48f1b236 chore(release): release Buzz Desktop version 0.5.2 (#3624)
## Buzz Desktop release v0.5.2

### Changes since v0.5.1:

- feat(cli): mirror Desktop mention delivery
([#3330](https://github.com/block/buzz/pull/3330))
([`7adc46268`](https://github.com/block/buzz/commit/7adc46268d5e93f0b1d4dc8e700af22815dcac1b))
- fix(desktop): deduplicate relay outage notification
([#3579](https://github.com/block/buzz/pull/3579))
([`66e705492`](https://github.com/block/buzz/commit/66e7054928cc29395f828467c3e8c81b7408dd29))
- fix(desktop): reconcile thread arrivals at bottom
([#3585](https://github.com/block/buzz/pull/3585))
([`b42a8d447`](https://github.com/block/buzz/commit/b42a8d447e3a2b85b2313dc4fdd123731fd8bba3))
- Improve emoji autocomplete matching
([#3571](https://github.com/block/buzz/pull/3571))
([`259de6afb`](https://github.com/block/buzz/commit/259de6afbe0cc0d106e57ebdb2323064990e4122))
- Fix shared agent avatar import profiles
([#3578](https://github.com/block/buzz/pull/3578))
([`324bd6b46`](https://github.com/block/buzz/commit/324bd6b464de5751e12abbd155376046ce3d2afc))
- Fix inline raster avatars in agent catalog
([#3581](https://github.com/block/buzz/pull/3581))
([`7e9b77f72`](https://github.com/block/buzz/commit/7e9b77f72d82e019a99f074f1c9829be30c57ae1))
- feat(agent): make Gemini and MLflow-route models usable through
databricks_v2 ([#3569](https://github.com/block/buzz/pull/3569))
([`4a1ebf25c`](https://github.com/block/buzz/commit/4a1ebf25c782fc6a68f0a69e6f866f793a259a1f))

**To release:** merge this PR. The tag and build will happen
automatically.

Signed-off-by: Wes <wesbillman@users.noreply.github.com>
v0.5.2
2026-07-29 14:57:35 -07:00
b18e559ae2 docs: add Linux rendering troubleshooting guide (#3573)
## What

Adds `docs/linux-rendering-troubleshooting.md` — the user-facing
troubleshooting page for Linux rendering failures.

## What's in the doc

**Crash: `colrv1_configure_skpaint` assertion abort (AppImage, Fedora
40+)**

Root cause: the AppImage bundles WebKitGTK compiled against FreeType
2.11.1, but `libfreetype.so.6` is not bundled — WebKit loads the host's
FreeType at runtime. FreeType 2.13.0 added a field to
`FT_ColorStopIterator` (16 → 20 bytes); on hosts with FreeType ≥ 2.13
the struct-layout mismatch corrupts Skia's COLRv1 color-stop arithmetic,
causing the assertion abort. Fix: upgrade to v0.5.2+ (build container
bumped to `ubuntu:24.04` in
[#3602](https://github.com/block/buzz/pull/3602)). Includes the glibc
floor table (2.35 → 2.39) and `.deb`/`.rpm` guidance for Ubuntu 22.04 /
Debian 12 users. A manual fontconfig workaround is preserved for users
stuck on older AppImages.

**Blank window / dmabuf renderer (NVIDIA, AppImage)**

Covers the auto-fix shipped in v0.5.1
([#3271](https://github.com/block/buzz/pull/3271)) and the
`--safe-rendering` flag for cases where auto-detection misses.

**AMD RDNA4 / transparent window
([#2643](https://github.com/block/buzz/issues/2643))**

Documents the three-variable workaround verified by the reporter
(`GDK_BACKEND=x11`, `WEBKIT_DISABLE_DMABUF_RENDERER=1`,
`WEBKIT_SKIA_ENABLE_CPU_RENDERING=1`).

Also includes a crash-log capture recipe and issue-filing checklist.

Context: [#2548](https://github.com/block/buzz/issues/2548),
[#2982](https://github.com/block/buzz/issues/2982),
[#2643](https://github.com/block/buzz/issues/2643),
[#2338](https://github.com/block/buzz/issues/2338).

Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
Co-authored-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
2026-07-29 17:54:33 -04:00
Xule LinandGitHub 5aeed7c7a2 fix(desktop): discover bun-installed agent CLIs in ~/.bun/bin (#3343)
## Problem

`common_binary_paths()` probes mise shims, `~/.local/bin`, volta, asdf,
and (further down `resolve_command_uncached`) nvm's default bin dir —
but not bun's global bin directory, `~/.bun/bin`.

bun's installer appends its bin dir to `~/.zshrc` / `~/.bashrc`, which
are **interactive**-only. A login shell never sources them, so
`find_via_login_shell()` can't recover the path either. That's the same
failure mode already called out in this file for nvm:

```rust
// Check nvm's default Node.js bin directory — nvm initializes via
// ~/.zshrc (interactive) which is not loaded by a login shell, so
// `node`, `npm`, and npm-global shims installed there are otherwise
// invisible.
```

So for a GUI-launched desktop app, every rung of the resolution ladder
misses a bun-installed CLI:

1. workspace dev dirs — no
2. `command_looks_like_path` — no, presets use bare names
3. Buzz-managed npm/node dirs — no
4. current process PATH — launchd's minimal PATH on a Finder launch
5. `find_via_login_shell` — `.zshrc` not sourced
6. `common_binary_paths()` — **`~/.bun/bin` absent**
7. nvm default bin — no

This matters because bun is a common install route for the agent CLIs
Buzz targets. Kimi Code in particular ships as an npm package
(`@moonshot-ai/kimi-code`), so `bun add -g` puts it at `~/.bun/bin/kimi`
— exactly where discovery doesn't look.

## Reproduction

On macOS with `codex` and `kimi` installed via bun, launching Buzz from
Finder:

- Kimi Code shows **"CLI needed"**
- both CLIs run fine in an interactive terminal

Probing the way `find_via_login_shell` does, in a clean environment:

```console
$ env -i HOME=$HOME /bin/zsh -l -c 'command -v -- codex; command -v -- kimi'
(nothing)
```

Launching the app with the bun dir on PATH resolves both immediately:

```console
$ env PATH="$HOME/.bun/bin:$PATH" /Applications/Buzz.app/Contents/MacOS/buzz-desktop
```

## Change

One entry appended to the home-relative list in `common_binary_paths()`.
It goes **last** so it cannot shadow a directory that already resolves —
the change can only add resolutions, never alter existing ones.

## Testing

`cargo fmt --check` passes.

I was not able to run the full `just ci` gate locally: `ring 0.17.14`
fails to build in this environment against the macOS 26.2 SDK (`cc`
error compiling `p256-nistz.c`), which is unrelated to this change.
Relying on CI for the rest — the diff adds one `PathBuf` to an existing
`Vec<PathBuf>` and introduces no new API.

## Notes

- Related to #3084, which adds `~/.kimi-code/bin` for the same class of
GUI-launch discovery failure. That covers Kimi's standalone installer;
this covers the bun/npm-global install route. They're complementary —
I've left a note on that PR.
- Only `~/.bun/bin` is added. bun's global packages live under
`~/.bun/install/global/node_modules` but are symlinked into
`~/.bun/bin`, so the single directory is sufficient.
- Worth noting `~/.bun/bin` contains no `node`/`npm`/`npx`, so appending
it can't shadow a system Node toolchain.

Signed-off-by: Xule Lin <43122877+linxule@users.noreply.github.com>
2026-07-29 21:50:55 +00:00
88d08e426c feat(desktop): guide and verify password-protected backups (#3617)
**Category:** improvement
**User Impact:** Users can create a password-protected key backup, save
it locally, and prove that it works before finishing onboarding or
signing out.

**Problem:** Tyler's NIP-49 foundation made encrypted local backup
possible, but the end-to-end ceremony still needed clearer progression,
trustworthy verification, and a Settings treatment consistent with the
rest of Identity. **Solution:** This stack adds a guided
create/download/test flow, verifies the actual saved backup through Rust
without returning secret key material, minimizes password lifetime in
the webview, and restores warm, plain-language presentation across
onboarding and Settings.

> [!NOTE]
> This PR is intentionally stacked on #2937 (`eva/nip49-local-backup`).
Review only the 19 commits in this stack.

<details>
<summary>File changes</summary>

**desktop/src-tauri/src/commands/identity.rs**  
Adds off-thread verification of NIP-49 backups and returns public
identity details only.

**desktop/src-tauri/src/commands/identity_key_backup_tests.rs**  
Covers verification, identity matching, and failure behavior at the
command boundary.

**desktop/src-tauri/src/egress_guard_tests.rs**  
Keeps the NIP-49 source inventory aligned with the new verification
path.

**desktop/src-tauri/src/key_backup.rs**  
Supports real backup decryption for verification while zeroizing
submitted passwords.

**desktop/src-tauri/src/key_backup_tests.rs**  
Exercises valid, wrong-password, damaged, and different-identity
backups.

**desktop/src-tauri/src/lib.rs**  
Registers the verification command with the desktop runtime.

**desktop/src/features/communities/ui/WelcomeSetup.tsx**  
Routes community onboarding through the revised backup ceremony.

**desktop/src/features/onboarding/lib/encryptedBackup.test.mjs**  
Covers the hardened backup state model and password-clearing behavior.

**desktop/src/features/onboarding/lib/encryptedBackup.ts**  
Models creation and verification with opaque request correlation and
short-lived passwords.

**desktop/src/features/onboarding/ui/BackupStep.tsx**  
Presents the backup choice clearly and advances into the dedicated
download step.

**desktop/src/features/onboarding/ui/BackupTestFlow.tsx**  
Adds the polished select-file, enter-password, verification, error, and
success experience.

**desktop/src/features/onboarding/ui/CommunityOnboardingFlow.tsx**  
Connects community onboarding to the updated backup steps.

**desktop/src/features/onboarding/ui/DownloadKeyStep.tsx**  
Adds the dedicated encrypted-backup download step and pre-creation skip
escape hatch.

**desktop/src/features/onboarding/ui/EncryptedBackupCreator.tsx**  
Guides password generation, backup creation, native saving, safe retry,
and re-download.

**desktop/src/features/onboarding/ui/MachineOnboardingFlow.tsx**  
Sequences chooser, download, exact-file test, and setup progression.

**desktop/src/features/onboarding/ui/NostrKeyImportForm.tsx**  
Aligns encrypted-key import behavior with the backup flow.

**desktop/src/features/onboarding/ui/NsecMaskedDisplay.tsx**  
Improves masked/revealed key presentation and safely wraps long private
keys.

**desktop/src/features/onboarding/ui/OnboardingChrome.tsx**  
Supports the revised onboarding layout and transitions.

**desktop/src/features/onboarding/ui/SetupStep.tsx**  
Integrates the completed backup ceremony with final setup.

**desktop/src/features/onboarding/ui/onboardingFlowSteps.test.mjs**  
Updates flow-level assertions for the new step sequence.

**desktop/src/features/settings/ui/EncryptedBackupRow.tsx**  
Adds sibling Create and Test rows that match the surrounding Identity
settings rhythm.

**desktop/src/features/settings/ui/ProfileSettingsCard.tsx**  
Places password-backup controls in the Identity card.

**desktop/src/features/settings/ui/SignOutSection.tsx**  
Uses explicit backup self-attestation plus the existing wipe phrase
without requiring a raw-key reveal first.

**desktop/src/shared/api/tauriIdentity.ts**  
Exposes typed backup verification to the frontend.

**desktop/src/shared/lib/ncryptsecSourceScan.test.mjs**  
Updates the frontend source allowlist for the verification path.

**desktop/src/testing/e2eBridge.ts**  
Models creation, saving, and verification for browser coverage.

**desktop/tests/e2e/onboarding-backup.spec.ts**  
Covers backup creation, exact-file testing, errors, retries, and
completion.

**desktop/tests/e2e/onboarding-docked-cta-screenshots.spec.ts**  
Updates docked CTA visual coverage for the revised progression.

**desktop/tests/e2e/onboarding.spec.ts**  
Keeps broader onboarding coverage aligned with the backup ceremony.

**desktop/tests/e2e/profile-nsec-reveal.spec.ts**  
Covers Settings create/test behavior and updated success copy.

**desktop/tests/e2e/signout-confirmation.spec.ts**  
Covers backup self-attestation, wipe phrase confirmation, and reset
behavior.

**desktop/tests/helpers/fileDrag.ts**  
Adds reusable file-drop support for backup test coverage.

**desktop/tests/helpers/onboarding.ts**  
Routes tests through chooser, download, and skip behavior consistently.

</details>

## Reproduction steps

1. Start fresh onboarding and choose the password-protected backup path.
2. Continue to **Backup your key**, create a password, save the
`.ncryptsec` file, and confirm that **Next** remains gated until testing
succeeds.
3. Select that exact file, enter its password, and verify the **Your
backup works!** success state.
4. Open **Settings → Identity** and confirm that **Create password
backup** and **Test password backup** appear as sibling rows.
5. Test a valid backup, then test wrong-password and different-identity
cases to confirm clear, non-secret-bearing outcomes.
6. Open sign out and confirm that backup self-attestation plus typing
`wipe all my data` remain required.

## Screenshots

### Onboarding flow

| 1. Creating your key | 2. Key created | 3. Create backup password |
|---|---|---|
| <img width="1280" height="720" alt="image"
src="https://github.com/user-attachments/assets/fe5ebf60-02ad-4584-8e64-ff137bafac02"
/> | <img width="1280" height="720" alt="image"
src="https://github.com/user-attachments/assets/0b8c7c6c-2ae5-494e-bb81-0a7b47ea1b66"
/> | <img width="1280" height="720" alt="image"
src="https://github.com/user-attachments/assets/9fa8f6c0-832a-46ec-9af8-55dcae0adbb8"
/> |
| **4. Select the saved backup** | **5. Enter its password** | **6.
Verification succeeds** |
| <img width="1280" height="720" alt="image"
src="https://github.com/user-attachments/assets/7ed1528c-6a1f-4411-bd3e-c57413320cdf"
/> | <img width="1280" height="720" alt="image"
src="https://github.com/user-attachments/assets/088ec3ed-2735-4b9a-a0c5-6d176b3f06c5"
/> | <img width="1280" height="720" alt="image"
src="https://github.com/user-attachments/assets/7081ce32-9371-46b2-9692-23ff77c39d35"
/> |

### Settings flow

| Create and Test tools | Tested backup success |
|---|---|
| <img width="1280" height="720" alt="image"
src="https://github.com/user-attachments/assets/978ab672-5510-46d8-9116-6df4f1bda34c"
/> | <img width="1280" height="720" alt="image"
src="https://github.com/user-attachments/assets/56a1764d-3c3e-40e2-9f4b-962c5e96a36c"
/> |

## Verification

- Desktop typecheck
- Full desktop JS: 3,733/3,733
- `pnpm check`
- Rust desktop lib: 1,862 passed, 14 ignored; diagnostic: 3/3
- Focused Playwright: 16/16 with one worker

Native save/cancel remains the residual smoke-test risk: browser E2E
mocks the KDF/native save-dialog boundary, while real decrypt and
identity matching are covered in Rust.

---------

Signed-off-by: Taylor Ho <taylorkmho@gmail.com>
Co-authored-by: npub1223z34hd7vtwc6qj4s7flsxkj644nlre2nthu7lrrmkumhu3xddsrx9r6w <52a228d6edf316ec6812ac3c9fc0d696ab59fc7954d77e7be31eedcddf91335b@buzz.block.builderlab.xyz>
2026-07-29 17:42:55 -04:00
005b5b819a feat(tracing): correlate trace IDs in relay logs (#3608)
## Summary
Correlates trace + span IDs with logs, allowing traces and logs to be
bridged seamlessly

### Related issue
none found

### Testing
Unit tests

Signed-off-by: David Grochowski <dgrochowski@squareup.com>
Co-authored-by: Amp <amp@ampcode.com>
2026-07-29 17:01:26 -04:00
581baa6254 chore(ci): bump Linux AppImage build container to ubuntu:24.04 (#3602)
## What and why

The Buzz AppImage is built on `ubuntu:22.04`, which links WebKitGTK
against FreeType 2.11.1. Because `libfreetype.so.6` is on the
linuxdeploy community excludelist, the bundled WebKit loads the
**host's** FreeType at runtime instead of the bundled one.

FreeType 2.13.0 (released 2023-02-09) added `FT_Bool read_variable` to
`FT_ColorStopIterator`, growing the struct from 16 to 20 bytes. Any host
running FreeType ≥ 2.13 (Fedora 42+, Ubuntu 24.04+) has a struct-layout
mismatch with the 22.04-compiled WebKit. The mismatched offsets corrupt
color-stop index arithmetic inside Skia's COLRv1 renderer, producing the
assertion abort in issues #2548 and #2982:

```
stl_vector.h:1123: Assertion '__n < this->size()' failed.
... colrv1_configure_skpaint(FT_Face, ...) ...
```

## Fix

Bump the build container to `ubuntu:24.04` (noble), which ships FreeType
**2.13.2**. Noble's struct layout matches every crash-affected host. The
ABI mismatch disappears and the crash is eliminated at root.

WebKitGTK also advances from **2.50.4** (jammy backport) to **2.52.3**
(noble backport).

## Glibc floor change

| Build base | glibc floor | Oldest supported AppImage distro |
|---|---|---|
| ubuntu:22.04 (before) | 2.35 | Ubuntu 22.04 LTS, Debian 12 |
| ubuntu:24.04 (after) | 2.39 | Ubuntu 24.04 LTS, Fedora 40+ |

Ubuntu 22.04 LTS and Debian 12 users lose AppImage support. Both
distributions continue to receive first-class `.deb` / `.rpm` packages,
which use the system WebKit and are unaffected. The crash-affected users
(Fedora 42/44, Ubuntu 24.04+) all have glibc ≥ 2.39.

## Changes

- `.github/workflows/linux-canary.yml:24` — container pin updated to
`ubuntu:24.04@sha256:4fbb8e6a…`
- `.github/workflows/release.yml:479` — same container pin updated
- `.github/workflows/release.yml:501` — comment version string updated
from 22.04 to 24.04

`fix-appimage.sh` and `desktop/src-tauri/**` are untouched. The #3573
fontconfig stopgap remains active; retirement is a separate follow-on PR
once this fix is verified on a shipped build.

## Sequencing

`docs/linux-rendering-troubleshooting.md` (introduced in #3573) will
receive a glibc-floor callout section once #3573 merges — adding it here
would conflict with #3573's open branch.

Context: #2548, #2982.

Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
Co-authored-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
2026-07-29 16:38:44 -04:00
7adc46268d feat(cli): mirror Desktop mention delivery (#3330)
🤖
## Summary

Agent-authored mentions currently depend on matching visible `@Name`
text to channel profiles. That makes notification delivery ambiguous
when names collide or profiles change, and it encourages an extra
post-send lookup just to confirm that the intended `p` tags were
emitted.

This change makes `buzz messages send` mirror Desktop's existing model:
the message keeps a readable name in its content while the recipient
pubkey is supplied separately.

```bash
buzz messages send \
  --channel <UUID> \
  --content '@Alice could you review this?' \
  --mention <alice-hex-or-npub>
```

`--mention` is repeatable. The CLI normalizes and deduplicates explicit
pubkeys, merges them with any names it can resolve from the channel, and
gives explicit identities priority under the existing 50-mention limit.

Before uploading attachments, signing, or publishing, the command checks
every resulting pubkey against the channel's current membership:

- Members are mentioned normally.
- Non-members stop the send and produce an actionable error.
- `--allow-non-member-mentions` deliberately sends notifying `p` tags
without adding anyone to the channel.

Sending a message never changes membership. On success,
`mention_pubkeys` is read from the exact signed event and returned with
the relay response, so callers can verify the emitted recipients without
another query.

Managed-agent guidance teaches this single-command mention flow. Desktop
mention behavior and the Nostr event schema are unchanged. Forum
guidance is intentionally handled separately in #3596.

### Related issue

None found. This replaces the earlier guidance-only approach in this PR
with the underlying CLI behavior it required.

### Testing

- `cargo test -p buzz-sdk`
- `cargo test -p buzz-cli`
- `cargo test -p buzz-acp`
- `cargo test --manifest-path desktop/src-tauri/Cargo.toml`

---------

Signed-off-by: npub1fdupjvyregj3z2tx7gx5x6py04zw89jm5usef9lyea4f3vcgh8qq9zgkdz <4b78193083ca25112966f20d4368247d44e3965ba7219497e4cf6a98b308b9c0@buzz.block.builderlab.xyz>
Signed-off-by: npub13n66s06epmqf2kc3v373ez8hj65cuzyvxzjf93vwpervxqn2u7jq2qd9je <8cf5a83f590ec0955b11647d1c88f796a98e088c30a492c58e0e46c3026ae7a4@buzz.block.builderlab.xyz>
Co-authored-by: npub1fdupjvyregj3z2tx7gx5x6py04zw89jm5usef9lyea4f3vcgh8qq9zgkdz <4b78193083ca25112966f20d4368247d44e3965ba7219497e4cf6a98b308b9c0@buzz.block.builderlab.xyz>
Co-authored-by: npub13n66s06epmqf2kc3v373ez8hj65cuzyvxzjf93vwpervxqn2u7jq2qd9je <8cf5a83f590ec0955b11647d1c88f796a98e088c30a492c58e0e46c3026ae7a4@buzz.block.builderlab.xyz>
2026-07-29 16:37:53 -04:00
66e7054928 fix(desktop): deduplicate relay outage notification (#3579)
🤖
## Summary
- Keep the relay reconnect notification dismissed across repeated
connection retries during one continuous outage.
- Re-arm the notification after recovery or relay lifecycle replacement,
including switches between communities that use the same relay URL.
- Preserve the existing dedicated path for authentication and other
application-level errors.

### How an outage is tracked
The AppShell-owned relay-card hook treats an outage as one contiguous
runtime episode rather than assigning it a persisted ID. A hook-local
`outageActiveRef` is armed by the first qualifying unreachable/degraded
observation. While it is armed, intermediate retry states (`connecting`,
`reconnecting`, `stalled`, and `disconnected`) belong to that same
episode, so retry churn cannot clear dismissal or emit another
notification.

The hook receives the same lifecycle identity used by community
initialization: community ID plus `reinitKey`. This distinguishes
multiple communities even when they share a relay URL, and it also
changes when the active community is explicitly reinitialized. The latch
and dismissal are reset when that identity changes or when the relay
singleton reports its authoritative `idle` teardown state. A successful
`connected` state also closes the episode and re-arms the next outage.

These boundaries deliberately bias toward re-notifying rather than
suppressing a later outage: recovery, community switch/reinit, or relay
teardown cannot leave the hook stuck believing an old outage is still
active. No outage state is persisted beyond the mounted hook lifecycle.

### Related issue
None found.

### Testing
- `pnpm --dir desktop typecheck`
- `pnpm --dir desktop test` — 3,769 passed
- `pnpm --dir desktop check`
- `pnpm --dir desktop build:e2e`
- `pnpm --dir desktop exec playwright test
tests/e2e/sidebar-relay-card.spec.ts --project=integration` — 11 passed

---------

Signed-off-by: Larry <8cf5a83f590ec0955b11647d1c88f796a98e088c30a492c58e0e46c3026ae7a4@buzz.block.builderlab.xyz>
Co-authored-by: npub13n66s06epmqf2kc3v373ez8hj65cuzyvxzjf93vwpervxqn2u7jq2qd9je <8cf5a83f590ec0955b11647d1c88f796a98e088c30a492c58e0e46c3026ae7a4@buzz.block.builderlab.xyz>
2026-07-29 15:48:06 -04:00
b42a8d447e fix(desktop): reconcile thread arrivals at bottom (#3585)
## Summary
- reconcile stale native-scroll anchors when a reply arrives at the
physical floor
- clear the thread new-message affordance instead of incrementing it
from stale cached state
- preserve the existing mid-history path and add direct lifecycle
regression coverage

## Why
PR #3411 fixed geometry-driven reconciliation, but the reply-arrival
branch still trusted a cached `message` anchor without checking the
rendered position. Native anchoring could return a short thread to the
floor without another scroll/resize callback, then the next reply
incremented the pill anyway.

## Verification
- Desktop checks passed
- Desktop typecheck passed
- focused lifecycle test passed (6/6)
- push hook full Desktop unit suite passed (3,770/3,770)
- `git diff --check` passed

Signed-off-by: Wes <wesbillman@users.noreply.github.com>
Co-authored-by: Carl <c7ebe626f000404285d3686e1dc74cc07cc60a9754a150041ba132e14bd3e2ec@buzz.block.builderlab.xyz>
2026-07-29 12:46:47 -07:00
Krishna CandGitHub 047533c56c fix(mobile): keep TLS on relays joined by invite (#3139)
## Summary

Communities joined via an invite link never connect: the app dials
`ws://` on port 80 instead of `wss://` on 443 and sits on
"Reconnecting…" indefinitely.

`RelayConfig.baseUrl` is documented as an HTTP origin, but the two
onboarding flows disagree on what they persist:

- **Device pairing** validates and stores `https://` —
`pairing_provider.dart:657` throws on anything else.
- **Invite join** stores the relay URL straight off the invite link, and
`deep_link.dart:165` always emits `ws://` or `wss://`.

`wsUrl` only special-cased `https://`, so a `wss://` base fell through
to the plaintext branch:

```dart
final scheme = uri.scheme == 'https' ? 'wss' : 'ws';   // 'wss' is not 'https'
```

The claim request itself succeeds, because `_claimUrlFromRelay`
(`invite_join_provider.dart:242`) maps `wss → https` explicitly. Only
the socket path is missing that conversion — which is why the community
appears, correctly named, and then never loads.

The same `baseUrl` also feeds `/query` (`relay_session.dart:136`), media
upload (`media_upload.dart:765`), Blossom auth (`media_auth.dart:128`)
and `relayClientProvider` (`relay_provider.dart:113`), so those requests
were malformed too. Where port 80 *does* answer, it is additionally a
silent TLS downgrade after `validateInviteRelayUri` insisted on
`wss://`.

This folds the websocket schemes back to their HTTP equivalents in
`baseUrl` itself, so every consumer is correct by construction rather
than needing a second getter remembered at each call site, and
communities **already persisted** with `wss://` are repaired on read
without a migration. `community_icon_provider.dart:46` already performs
this same conversion locally.

One subtlety worth flagging for review: the normalization is derived in
the getter rather than applied in the constructor, so the constructor
stays `const`. The compile-time fallback at `relay_provider.dart:77`
relies on const canonicalization for a stable identity across rebuilds,
and Riverpod's `defaultUpdateShouldNotify` is `previous != next`
(`element.dart:361`), which falls back to identity for this class. A
`factory` constructor here yields a fresh instance per rebuild, which
tears down and resubscribes every listener —
`channels_provider_test.dart` catches it as an unexpected unsubscribe
during reconnect.

### Related issue

Fixes #2662.

### Testing

`flutter test` — **705 passed, 1 skipped, 0 failed**
`flutter analyze` — No issues found
`dart format --set-exit-if-changed .` — 249 files, 0 changed

Run against the Hermit-pinned SDK (Flutter 3.41.7 / Dart 3.11.5),
matching CI.

10 new unit tests in `mobile/test/shared/relay/relay_config_test.dart`
covering both onboarding schemes, `http`/`https` passthrough,
non-default ports, and agreement between the invite and pairing paths
for the same relay.

Verified end-to-end against a self-hosted relay behind `tailscale
serve`, which terminates TLS on 443 and leaves port 80 closed. Relay
logs show the invite claim succeeding over HTTPS at the moment of
joining, while no WebSocket connection ever arrives — no `WebSocket
connection established`, no NIP-42 auth, no `kind:0` profile, no push
registration — across the relay's entire history, even though the member
row is present and correct. Port-80 refusals are not logged by
`tailscaled`'s netstack, which is why the retries leave no trace
server-side. Reproduced on both iOS and Android.

---------

Signed-off-by: Krishna C <github@kumb.uk>
2026-07-29 12:02:42 -07:00
klopez4212andGitHub 259de6afbe Improve emoji autocomplete matching (#3571)
## Summary
- Show all colon emoji autocomplete matches
- Rank exact and prefix shortcodes before weaker matches
- Add a regression test and screenshot

## Validation
- `pnpm test`
- `pnpm build`
- `pnpm exec playwright test --project=smoke
tests/e2e/custom-emoji.spec.ts --grep "exact standard shortcode"`
- `just desktop-tauri-clippy`

Native Tauri tests were attempted but could not link because the local
disk filled during compilation.

---------

Signed-off-by: kenny lopez <klopez4212@gmail.com>
2026-07-29 19:49:57 +01:00
324bd6b464 Fix shared agent avatar import profiles (#3578)
> Carl is updating this pull request on Wes's behalf.

## Summary

- upload an embedded raster avatar through the existing authenticated
media pipeline before minting or persisting an imported shared agent
- store and publish only the resulting hosted URL so agent kind:0
profiles remain within content limits

## Root cause

Snapshot import recovered raster avatar pixels as a large inline base64
data URL. That value was persisted and placed into the agent's kind:0
profile. The relay rejected the oversized profile, so other clients
could not resolve the imported agent's avatar.

## Scope

This is intentionally the forward fix only. It changes two Desktop files
and does **not** add migration or reconciliation behavior for previously
imported agents. Existing affected imports must be re-imported or fixed
manually.

## Validation

- successful pre-push Desktop suite: 1,863 passed, 14 ignored, 0 failed
- all pre-push Rust/Desktop gates green, including all-target clippy
- valid >256 KiB PNG import → production MIME detection/sanitization →
bounded signed kind:0 containing only the hosted URL
- upload failure, malformed data, and URL-only avatar cases covered
- independent fresh review by Princess Donut: clean, no blocking
findings

Signed-off-by: Wes <wesbillman@users.noreply.github.com>
Co-authored-by: Carl <c7ebe626f000404285d3686e1dc74cc07cc60a9754a150041ba132e14bd3e2ec@buzz.block.builderlab.xyz>
2026-07-29 11:03:20 -07:00
7e9b77f72d Fix inline raster avatars in agent catalog (#3581)
## Summary

- render existing shared personas whose catalog avatar is a bounded
inline PNG, JPEG, GIF, or WebP data URL
- keep rejecting arbitrary, malformed, unsupported, and oversized
`data:` URLs
- preserve hosted relay URLs as the forward format; catalog browsing
remains read-only

## Root cause

Paul's live shared kind:30175 head contains a 144,878-character
`data:image/png;base64,...` avatar. The catalog projection accepted
HTTP(S) URLs and bounded percent-encoded SVG emoji avatars only, so it
projected Paul's avatar to `null` before `ProfileAvatar` rendered it.
The owner still saw the local persona avatar, producing the reported
owner/viewer mismatch.

This patch accepts only four raster MIME types with strict base64 shape
and a 256 KiB total URL cap at the existing catalog parsing boundary. It
repairs already-signed heads such as Paul without viewer-side uploads or
publication side effects.

Hosted media remains the canonical forward path. #3578 uploads inline
raster avatars during snapshot import, preventing the known source from
creating future inline persona/profile values; existing signed catalog
heads still need this compatibility path until their owners republish.

## Agent instruction finding

The catalog publishes `AgentDefinition.system_prompt` verbatim as the
user-authored **Agent instruction**, as documented by the sharing UI and
NIP-AP. No Buzz base/core/runtime prompt is concatenated in the publish,
catalog, or import path. This PR therefore does not remove authored
instructions and accidentally strip copied agents of their behavior.

## Validation

- targeted `personaCatalogRelay.test.mjs`: 24 passed
- Desktop typecheck: passed
- pre-push Desktop frontend suite: 3,771 passed
- pre-push Desktop checks: passed
- `git diff --check`: passed
- independent review: no blockers, 9.4/10

Signed-off-by: Wes <wesbillman@users.noreply.github.com>
Co-authored-by: Carl <c7ebe626f000404285d3686e1dc74cc07cc60a9754a150041ba132e14bd3e2ec@buzz.block.builderlab.xyz>
2026-07-29 17:36:12 +00:00
Will PflegerandGitHub ddd468723a revert(acp): remove dead GOOSE_ACP_SCHEDULER_DISABLED env injection (#3576)
## Summary

[block/buzz#3144](https://github.com/block/buzz/pull/3144) injected
`GOOSE_ACP_SCHEDULER_DISABLED=true` into every `AcpClient::spawn` call
as a forward-compatible no-op, intended to suppress the cron scheduler
in goose ACP children once the matching reader landed in goose. That
reader only ever existed in
[aaif-goose/goose#10738](https://github.com/aaif-goose/goose/pull/10738),
which was closed unmerged.

[goose#10781](https://github.com/aaif-goose/goose/pull/10781) (Lifei
Zhou, merged 2026-07-29) disables the ACP scheduler by default at the
source: `goose acp` now requires `--enable-scheduler` to start a
scheduler. Buzz-spawned children therefore get no scheduler with zero
configuration — making the `GOOSE_ACP_SCHEDULER_DISABLED` injection
permanently dead code.

## What changes

Removes from `crates/buzz-acp/src/acp.rs`:

- `GOOSE_SCHEDULER_DISABLED_ENV` constant
- `cmd.env(GOOSE_SCHEDULER_DISABLED_ENV, "true")` injection in
`AcpClient::spawn`
- `spawn_injects_scheduler_disabled_env_by_default` test
- `spawn_scheduler_disabled_env_overrides_conflicting_extra_env` test
- `spawn_and_read_child_env` helper (unreferenced once the two tests
above are gone)

No other files are affected.

## Why now

Leaving dead code that references an env var no reader will ever consume
misleads future maintainers about the actual scheduler-isolation
mechanism. The isolation is now an upstream default, not a Buzz
injection.

Reverts: [block/buzz#3144](https://github.com/block/buzz/pull/3144)
Related:
[aaif-goose/goose#10781](https://github.com/aaif-goose/goose/pull/10781)

Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
2026-07-29 13:08:32 -04:00
4a1ebf25c7 feat(agent): make Gemini and MLflow-route models usable through databricks_v2 (#3569)
## Summary

Makes Gemini — and every other non-Claude, non-GPT-5 model on the
Databricks MLflow route (`databricks_v2`) — usable in an agent loop.
These are the non-`benchmarks/` changes from
`benchmark/harness-accounting-and-solo`, lifted onto a clean base off
`main` so they can land independently while the harness work continues.

Two defects made these models unusable, one fatal and one silent. Both
live only in `openai_body` / `parse_openai`, which is the
least-exercised of the three `databricks_v2` sub-routes — the
`luna`/`sol` conditions run the Responses route and the `opus`
conditions run the Anthropic route, so **this change is inert for every
model already in use** and only lights up the MLflow path.

## Why a third route at all

`databricks_v2_route_for_model` buckets by model family: `claude*` →
Anthropic Messages, `gpt-5`/code-names → OpenAI Responses, **everything
else → MLflow chat-completions**. Gemini, Qwen, gpt-oss, and friends all
fall through to that third pair — and both bugs below live only there.

## D1 — dropped thought signatures (fatal)

Gemini returns a `thoughtSignature` on every tool call and **requires it
echoed back**. `openai_body` reserialized each call as `{id, type,
function}` only, dropping the field, so the next request 400'd:

```
HTTP 400  Function call is missing a thought_signature in functionCall parts.
```

For a coding agent this fires on the **first** tool call, so the model
never completes a single turn.

**Position is load-bearing.** A four-shape replay probe against the live
gateway established that the signature must sit as a *sibling* of
`function` — nesting it inside `function{}` fails with the *same* 400 as
omitting it. A fix that "preserves the field" without preserving its
position passes a unit test and still 400s.

The fix: `ToolCall` gains `provider_extra: Map<String, Value>`.
`parse_openai` captures every top-level wire key except the three we
model (`id`, `type`, `function`); `openai_body` re-emits them beside
`function`. Keeping *whatever we did not model*, rather than naming
`thoughtSignature`, means the next provider with an opaque per-call
token needs no change here. The Responses and Anthropic replay shapes
are fully modelled, so they pass `Default::default()` and stay
**byte-identical** to before.

### D1b — duplicate tool-call ids (same root cause)

Gemini returns the **function name** as the id, so two parallel calls to
one function arrive sharing an id — and that id is what pairs a
`role:"tool"` result back to its call, making two results
indistinguishable. `dedupe_provider_ids` suffixes collisions
(`get_weather`, `get_weather-2`). Safe because both halves of the
pairing (the assistant `tool_calls[].id` and the result's
`tool_call_id`) are re-emitted from this same value; the provider never
sees its original id again.

## D2 — block-array content discarded (silent, worse than a crash)

`parse_openai` read `content` with `as_str()`, which returns `""` for
anything that isn't a JSON string. Gemini (and Qwen35, gpt-oss) send an
array of typed blocks:

```json
"content": [
  {"type": "reasoning", "summary": [{"type": "summary_text", "text": "…"}]},
  {"type": "text", "text": "391"}
]
```

So the model answered and the answer was thrown away — no error, no
warning, just a turn that looked like the model had said nothing. On a
benchmark this reads as "Gemini is bad at the task" rather than "buzz
dropped the reply."

`openai_content_parts` now accepts either shape — string as before, or a
block array where `text` blocks concatenate into text and `reasoning`
blocks into reasoning (Gemini nests the prose one level down under
`summary`). Message-level `reasoning_content` / `reasoning` still win
when present, so DeepSeek and vLLM-style hosts are unchanged; block
reasoning is the last fallback.

## Also: a turn-start log line (`buzz-acp` `pool.rs`)

Small, independent observability change that also rides in the
non-benchmark delta: `run_prompt_task` now emits a `pool::prompt` "turn
starting" line, labelled by the same `prompt_label` helper as
`log_stop_reason`, so a log reads as start/stop pairs. An unpaired start
is the only durable evidence that a turn was entered and never returned
— without it, a stalled agent and an agent nobody woke leave identical
(zero-completion) logs.

## Interaction with #3538

#3538 (already merged) rewrote `databricks_v2_route_for_model` to route
by boundary-aware model-family segments. That change and this one touch
**different functions** in `llm.rs` — routing vs. body/parse — and
compose cleanly; the family routing decides *which* pair runs, and this
fixes the MLflow pair it can now select.

## Testing

- `cargo fmt --all -- --check`, `cargo clippy -p buzz-agent -p buzz-acp
--all-targets -- -D warnings` — clean.
- `cargo test -p buzz-agent -p buzz-acp` — all green (304 + 632 lib
tests plus integration suites, 0 failures). Five new tests cover:
block-array text extraction, plain-string regression, passthrough
capture (and non-duplication of the modelled keys), replay position
(`thoughtSignature` beside `function`, not inside it), and id
de-duplication.
- Wire evidence: the four-shape replay table and the reasoning-effort
probe were run against `block-lakehouse-staging` (recorded in the design
doc).

## Relationship to the benchmark branch

The full design write-up (four-shape replay table, position-matters
analysis, effort verification, and open pricing item) lives in
`docs/08-gemini-provider-fixes.md` on
`benchmark/harness-accounting-and-solo`. The benchmark manifests and
endpoint-config entries that exercise these models are separable and
stay on that branch.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Atish Patel <atish@squareup.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-29 12:49:51 -04:00
9beb3b8c6e fix(cli): mask credential env values in --help output (#3570)
clap renders live env var values in help text by default. Three args
carrying credentials were exposed this way:

- `BUZZ_PRIVATE_KEY` in `buzz-cli` (`crates/buzz-cli/src/lib.rs`)
- `BUZZ_AUTH_TAG` in `buzz-cli`
- `BUZZ_PRIVATE_KEY` in `buzz-acp` (`crates/buzz-acp/src/config.rs`)

Add `hide_env_values = true` to each. Env var names remain visible for
discoverability; only their runtime values are withheld from `--help`
output.

Also adds a regression guard in each crate's test module that walks the
clap command tree (recursing into subcommands for `buzz-cli`) and
asserts every arg whose env var name contains `KEY`, `SECRET`, `TOKEN`,
`PASSWORD`, `CRED`, or `AUTH` has `hide_env_values` set. This prevents
future credential-bearing args from being added without the masking in
place.

Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
Co-authored-by: Duncan <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
2026-07-29 12:23:02 -04:00
9752b816a9 Serialize Tauri pre-push checks (#3567)
## Summary

- combine Desktop Tauri clippy and tests into one pre-push command
- run clippy first, then tests
- keep unrelated pre-push commands parallel

## Why

PR #3555 added clippy as a separate command while the pre-push group
uses `parallel: true`. That can start clippy and tests simultaneously
against the same Cargo target directory, leaving one command waiting on
Cargo's build lock and making pushes appear stalled.

Serializing only these two Cargo-heavy checks avoids lock contention
while retaining the CI-equivalent clippy command and existing test
coverage.

## Validation

- `lefthook validate`
- forced `desktop-tauri-checks` through Lefthook with an instrumented
`just`; observed `desktop-tauri-clippy` followed by `desktop-tauri-test`

Signed-off-by: Wes <wesbillman@users.noreply.github.com>
Co-authored-by: Carl <c7ebe626f000404285d3686e1dc74cc07cc60a9754a150041ba132e14bd3e2ec@buzz.block.builderlab.xyz>
2026-07-29 08:52:50 -07:00
WesandGitHub a13085e9ac chore(release): release Buzz Desktop version 0.5.1 (#3566)
## Buzz Desktop release v0.5.1

### Changes since v0.5.0:

- perf(desktop): move observer-feed archive and decrypt commands off
main thread ([#3415](https://github.com/block/buzz/pull/3415))
([`294c8c821`](https://github.com/block/buzz/commit/294c8c821de51442a8c384c0bdb66b1a10224ca0))
- fix(desktop): preserve shared agent fidelity
([#3553](https://github.com/block/buzz/pull/3553))
([`f7a3988ba`](https://github.com/block/buzz/commit/f7a3988ba13b590d9a55a7e8413fc3fb5ffbef18))
- feat(agent): route Claude/GPT model families to their native gateway
wire ([#3538](https://github.com/block/buzz/pull/3538))
([`6438dedf8`](https://github.com/block/buzz/commit/6438dedf83a9dbe1853e484326911bf6c7f1618c))
- Refine community invite limits
([#3529](https://github.com/block/buzz/pull/3529))
([`24d90d128`](https://github.com/block/buzz/commit/24d90d1280a9325c6cbcf8eea30ac54db5afd2cb))
- feat(agent): fix Anthropic prompt caching with Databricks (+ MCP
proxy/TLS passthrough)
([#3463](https://github.com/block/buzz/pull/3463))
([`c405ad1d4`](https://github.com/block/buzz/commit/c405ad1d4b1da061c11b3d26761252d41dcc62d3))
- feat: add explicit entry for claude-opus-5 in model config
([#2831](https://github.com/block/buzz/pull/2831))
([`90e058ebf`](https://github.com/block/buzz/commit/90e058ebf68137e048a409aec6616519379ff726))
- fix(desktop): clear stale thread new-message pill
([#3411](https://github.com/block/buzz/pull/3411))
([`55a3ed7b9`](https://github.com/block/buzz/commit/55a3ed7b9217cee5b23e0a5441947dc929b2a38c))
- fix(ci): ratchet file sizes against the base tree
([#3352](https://github.com/block/buzz/pull/3352))
([`9227bdf58`](https://github.com/block/buzz/commit/9227bdf58ad6664ae3c1078888f2181ec19c4da4))
- feat(desktop): apply WebKit rendering workarounds at startup on Linux
([#3271](https://github.com/block/buzz/pull/3271))
([`3ece4461d`](https://github.com/block/buzz/commit/3ece4461df8a7b9663a8e68327483b8377d4086d))
- fix(desktop): stabilize flaky DM expansion E2E ordering assertions
([#2004](https://github.com/block/buzz/pull/2004))
([`913d564ce`](https://github.com/block/buzz/commit/913d564ce0f35924291bf3eeab6508517a6d8d1f))
- fix(desktop): paint community rail full height
([#3382](https://github.com/block/buzz/pull/3382))
([`1d3b810ad`](https://github.com/block/buzz/commit/1d3b810ad70d6325718ed91e723f32c4a376d5e1))
- feat(desktop): add custom harness inline from agent dialogs
([#3252](https://github.com/block/buzz/pull/3252))
([`b0503d80c`](https://github.com/block/buzz/commit/b0503d80c298b1ece3b0a43b41d316829a3379e7))
- feat(desktop): refine agent catalog sharing
([#2439](https://github.com/block/buzz/pull/2439))
([`a35771fc4`](https://github.com/block/buzz/commit/a35771fc441cdc3c6f517f419037206783b502d2))
- fix(desktop): keep drafts out of the Inbox All view
([#3217](https://github.com/block/buzz/pull/3217))
([`3afa129ee`](https://github.com/block/buzz/commit/3afa129ee785cc74d921d0ba969254a8255c4cc0))
- fix(desktop): restore the inbox icon in the sidebar
([#3341](https://github.com/block/buzz/pull/3341))
([`00ede2e7a`](https://github.com/block/buzz/commit/00ede2e7aa7eb95571b7db3ebbd163adbf6cf74e))
- fix(desktop): gate codex-acp on a minimum supported version
([#3254](https://github.com/block/buzz/pull/3254))
([`4e3998f36`](https://github.com/block/buzz/commit/4e3998f36e36d68b9a93dcbd85f0864450bb8f5f))
- feat(cli): add users set-status command for NIP-38 profile status
([#3253](https://github.com/block/buzz/pull/3253))
([`60158fce3`](https://github.com/block/buzz/commit/60158fce3e670f11bb35d42627857ccaea50ff06))
- fix(composer): scope multiline block formatting
([#3246](https://github.com/block/buzz/pull/3246))
([`5457c947a`](https://github.com/block/buzz/commit/5457c947a74f5ba4b979f9c6411aa7626a858387))

**To release:** merge this PR. The tag and build will happen
automatically.

Signed-off-by: Wes <wesbillman@users.noreply.github.com>
v0.5.1
2026-07-29 15:44:30 +00:00
51bb97d2be Run Tauri clippy in pre-push (#3555)
## Summary

- run Desktop Tauri clippy from pre-push for every path that can affect
the Tauri crate
- reuse `just desktop-tauri-clippy`, keeping the local command identical
to Desktop Core CI
- leave the existing Tauri test hook unchanged

## Why

PR #3553 exposed a hook gap: `cargo test` allowed an unused-import
warning that CI's `clippy -D warnings` correctly rejected. Running the
same recipe before push catches that class of failure locally without
duplicating CI flags in Lefthook.

## Validation

- `lefthook run pre-push --command desktop-tauri-clippy --force`
- confirmed it invokes `cargo clippy --manifest-path
desktop/src-tauri/Cargo.toml --all-targets -- -D warnings`
- command passed

Signed-off-by: Wes <wesbillman@users.noreply.github.com>
Co-authored-by: Carl <c7ebe626f000404285d3686e1dc74cc07cc60a9754a150041ba132e14bd3e2ec@buzz.block.builderlab.xyz>
2026-07-29 08:20:49 -07:00
Will PflegerandGitHub 294c8c821d perf(desktop): move observer-feed archive and decrypt commands off main thread (#3415)
Opening the agent observer feed could beachball the app. In Tauri 2, a
sync (`pub fn`) command body runs on the **main thread** — only `async
fn` commands run on the runtime pool. Five commands on the observer-feed
open path were sync, so panel open ran SQLite I/O and secp256k1 work on
the macOS main thread:

| Command | Main-thread work |
|---|---|
| `decrypt_observer_event` | Schnorr ID + signature verify, then NIP-44
decrypt — once per frame |
| `read_archived_observer_events_for_channel` | Opens the archive DB,
runs the channel-index JOIN, returns up to 200 raw JSON blobs per page |
| `read_unindexed_observer_rows` | Opens the DB, returns **all**
not-yet-indexed kind-24200 rows in one shot |
| `index_observer_channel_id` | Opens the DB, loops N upserts |
| `delete_save_subscription` | Opens the DB, one delete |

Eager hydration loads up to 10 pages × 200 frames on panel open, so
that's up to 10 main-thread DB reads plus up to 2,000 sequential
verify+decrypt calls before any scrolling. The one-shot backfill makes
it worse on the first open after history accumulates: one read of every
unindexed row, a decrypt per row, then a batch upsert — all on the main
thread, and all proportional to archive size.

The four archive commands now route their DB work through the existing
`run_archive_db_task` helper (`spawn_blocking` + `open_db`), matching
`list_save_subscriptions`, `read_archived_events`, and `archive_events`
directly around them. `decrypt_observer_event` becomes `async fn` +
`tauri::async_runtime::spawn_blocking`, with `state.signing_keys()`
extracted before the spawn since `State` is not `Send` — the same
pattern `sign_event` uses from #1222.

No frontend changes: `invoke` is already promise-based, so the TS
wrappers in `tauriArchive.ts` and `tauriObserver.ts` are unchanged.

This removes the freeze, not the work. Eager hydration still takes the
same wall time — the feed shows a loading state instead of blocking the
UI. Batching the per-frame decrypt IPC (2,000 round-trips into one
command) would cut the latency itself; that's deliberately out of scope
here.

Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
2026-07-29 11:17:35 -04:00
f7a3988ba1 fix(desktop): preserve shared agent fidelity (#3553)
## Summary

Fixes two distinct fidelity failures in direct agent sharing:

- The sender now puts the same effective avatar shown on the agent card
into People-share and file-export snapshot PNGs, including
profile/kind:0 fallback avatars.
- The importer now persists the visible PNG body as the portable avatar
instead of ignoring it in favor of sender-local manifest references.
- Export materializes inherited runtime, provider, and model identifiers
verbatim, while preserving explicit definition values. It does not
translate or substitute configuration for a different recipient setup.
- Sharing waits for a profile-only fallback avatar query, preventing an
early-click race.

The PNG import path keeps the existing safety invariant: decode is
capped at 2048×2048 / 32 MiB and re-encoded avatars above the 2 MiB
inline limit fall back to the manifest reference. The exact transparent
1×1 no-avatar placeholder is ignored.

The original Tyler↔Wes screenshot demonstrates both stages: Wren's
attachment had an avatar that disappeared after **Add agent**
(receiver/import failure), while Pinky's attachment was already blank
(sender/projection failure).

### Related issue

N/A — reported and traced in the linked Buzz conversation.

### Testing

- `cargo test --manifest-path desktop/src-tauri/Cargo.toml
commands::personas::snapshot` — 57 passed
- `pnpm exec tsc --noEmit`
- Biome check on changed frontend/E2E files
- Pre-push hooks:
  - desktop check
  - desktop tests
  - desktop Tauri tests — 1853 passed, 14 ignored
  - file-size ratchet

The People-share E2E regression asserts that a profile-only avatar
reaches `avatarPngDataUrl` in the real encode command payload.

---------

Signed-off-by: Wes <wesbillman@users.noreply.github.com>
Co-authored-by: Carl <c7ebe626f000404285d3686e1dc74cc07cc60a9754a150041ba132e14bd3e2ec@buzz.block.builderlab.xyz>
2026-07-29 08:13:54 -07:00
klopez4212andGitHub 4555899ab2 Polish mobile navigation and menus (#3486)
## Summary
- Add a shared footer fade behind the floating tabs on Home, Activity,
and Search.
- Use a shared anchored popover for Activity filters and section
actions, with working section move controls.
- Polish message grouping/press states and remove the initial Search
back button.
<img width="630" height="1368" alt="Screenshot 2026-07-29 at 08 49 37"
src="https://github.com/user-attachments/assets/9e787adf-0bb3-49c6-8224-5819e8cfb1ad"
/>

### Testing
- `flutter analyze`
- `flutter test`
- Release build installed and checked on a connected iPhone

### Screenshots
A real-device Activity baseline showing the original solid footer is
attached in a PR comment. The updated review build was checked on the
connected iPhone.

---------

Signed-off-by: kenny lopez <klopez4212@gmail.com>
2026-07-29 16:10:51 +01:00
6438dedf83 feat(agent): route Claude/GPT model families to their native gateway wire (#3538)
## Summary

Databricks v2 chooses the gateway wire format — OpenAI Responses,
Anthropic Messages, or MLflow chat — purely from substrings in the
endpoint name. There is no family field on the endpoint to key off, so
the substring set *is* the routing contract. The matcher only recognised
`gpt-5`/`gpt5` and `claude`, which makes correct billing depend on every
Claude endpoint happening to be named with the literal string "claude".

## Why this matters

Getting a Claude model onto the Anthropic Messages route is exactly what
lets buzz attach the `cache_control` breakpoint (the fix in #3463). If a
Claude endpoint's catalog name omits "claude" — an alias, a bare
`opus-5`, a `goose-opus-5` — it silently falls through to the MLflow
(OpenAI-wire) path, where Anthropic prompt caching is **structurally
impossible**. The result is the same failure #3463 fixed: 0% cache
reads, the full ~10x read discount lost, and no error — a naming
convention quietly holding up a billing-correctness invariant.

## What changed

`databricks_v2_route_for_model` (`crates/buzz-agent/src/llm.rs`) now
matches broader, case-insensitive marker sets:

- **Claude → Anthropic Messages:** `claude`, `opus`, `sonnet`, `haiku`,
`mythos`, `fable` — the Claude family names and release code names, so a
Claude endpoint reaches the cache-capable route regardless of how it's
named.
- **GPT → OpenAI Responses:** the `gpt` family (now `gpt` on its own,
not just `gpt-5`) plus the GPT-5 launch code names `sol`, `luna`,
`terra`.

OpenAI markers are evaluated first, preserving the prior `gpt-5`-first
precedence for any name that could carry both. Names matching neither
set still fall through to the MLflow chat route.

## Testing

- `cargo fmt`, `cargo clippy -p buzz-agent --all-targets -- -D warnings`
— clean.
- `cargo test -p buzz-agent` — all green (299 lib + integration suites,
0 failures). The `databricks_v2_routes_by_model_family` test was
expanded to cover each new marker, the GPT-5 code names,
case-insensitivity, and the unchanged MLflow fallback (including
`gemini`).

## Relationship to #3463

#3463 taught the Anthropic path to request caching; this makes sure
Claude models actually land on that path. Follow-up still open:
surfacing `cache_creation_input_tokens` end-to-end so a persistent
`reads == 0 && writes == 0` reveals a disabled cache regardless of which
wire a model takes — happy to do that next.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Signed-off-by: Atish Patel <atish@squareup.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-29 15:02:12 +00:00
klopez4212andGitHub 24d90d1280 Refine community invite limits (#3529)
## Summary

- Simplify the community invite dialog around link sharing.
- Add matching expiry and use-limit dropdowns, with sensible preset use
caps.
- Cover the default unlimited and selected-limit invite payloads.

## Validation

- `pnpm -C desktop run build:e2e`
- `pnpm -C desktop exec playwright test
tests/e2e/invite-link-copy.spec.ts
tests/e2e/invites-settings-screenshots.spec.ts --project=smoke`

Signed-off-by: kenny lopez <klopez4212@gmail.com>
2026-07-29 15:42:17 +01:00
klopez4212andGitHub ce01e930ed Polish mobile typing indicator (#3528)
## Summary

- Present channel and thread typing status in a composer-matched
container.
- Animate the strip so the message list moves smoothly as typing begins
and ends.
- Increase typing-label contrast and avatar/padding for readability.

## Pixel 10 snapshot

![Typing indicator above the
composer](https://raw.githubusercontent.com/block/buzz/31de9f86a76fe61498bc7f2931d9e574827a9aa2/pr-3528--typing-indicator.png)

## Validation

- `flutter test test/features/channels/channel_detail_page_test.dart`
- `flutter analyze`

Signed-off-by: kenny lopez <klopez4212@gmail.com>
2026-07-29 15:35:41 +01:00
c405ad1d4b feat(agent): fix Anthropic prompt caching with Databricks (+ MCP proxy/TLS passthrough) (#3463)
> On 8 tasks matched by name across the two runs, cost fell $8.36 →
$1.77 (4.71×) and wall-clock 12,423 s → 1,085 s (11.45×).

## Summary

Two independent, self-contained fixes to `buzz-agent`/`buzz-acp`, split
out of the benchmark branch so they can land while the harness work
continues:

1. **Request and surface Anthropic prompt caching.** buzz never sent a
`cache_control` breakpoint, so on the Databricks Anthropic route
`cache_read_input_tokens` was **structurally always 0** and the ~10×
cache-read discount was never claimed. This teaches `anthropic_body()`
to mark the cacheable prefix, and plumbs the cache split end-to-end so
accounting can price it.
2. **Pass proxy + TLS-trust env into MCP tool subprocesses**, so agent
tools on a proxy-only host stop reporting a live network as offline.

## Why the caching gap matters

The Anthropic Messages API does **not** cache unless the request carries
a `cache_control` breakpoint, and the Databricks AI Gateway — a
third-party proxy in front of the model, in the same category as
Bedrock/Vertex — does **not** auto-cache (only the first-party Anthropic
API and Claude-on-AWS do zero-config caching). So every request was
billed cold.

Measured live against the Databricks gateway
(`databricks-claude-opus-5`, 2026-07-28), the same call with and without
a single `cache_control` marker:

| Run | `input_tokens` | `cache_creation` | `cache_read` | latency |
|---|---|---|---|---|
| No `cache_control`, two byte-identical calls | 121,625 | 0 | **0** |
~9.3 s |
| With one marker — cold (write) | 4 | 121,625 | 0 | 9.3 s |
| With one marker — warm (read) | 4 | 0 | **121,625** | **4.5 s** |

One marker moved 121,625 tokens from full-price input to a 0.1× cache
read and roughly halved latency (a clean, isolated ~2.07× prefill
speedup on this single-threaded microbenchmark). The gateway honours
`cache_control`; buzz simply never sent it.

At fleet scale this was a real budget item. Across matched
Terminal-Bench solo sweeps (89 tasks, `-n 20`, before the fix), the two
OpenAI-route models independently landed at ~86–87% cache reads — the
expected shape for an agentic loop, where system + tools + append-only
history repeat every turn — while the Anthropic route returned a hard 0%
on every receipt:

| Condition | Route | Input tokens | Cache reads | Cost | Cost if
uncached | Discount |
|---|---|---|---|---|---|---|
| luna (`gpt-5-6`) | OpenAI | 20,320,818 | **17.7M (87.0%)** | $6.96 |
$22.87 | **3.28×** |
| sol (`gpt-5-6`) | OpenAI | 22,312,290 | **19.2M (85.9%)** | $37.07 |
$123.35 | **3.33×** |
| opus (`claude-opus-5`) | Anthropic | 12,459,822 | **0 (0.0%)** |
$81.31 | $81.31 | **1.00×** |

Applying luna's measured 87% read rate to the opus token counts at list
prices (`input $5/M`, `cached_input $0.5/M`, `output $25/M`) puts the
opus run at **~$32.53 vs the $81.31 actually paid — a ~60% overspend on
those 49 trials (~$89 on a full sweep)**. That is an upper bound (it
prices every cached token at the 0.1× read rate and ignores the 1.25×
write premium), and the opus discount is structurally smaller than
luna/sol's because opus emits ~3.5× more uncacheable output per trial,
which sets a floor on what caching can recover.

There is also a plausible **second-order effect**: Databricks appears to
meter its per-minute rate limit on *uncached* input tokens, so the
missing cache also cost rate-limit headroom — the opus endpoint lost 63%
of its trials to fatal 429s while running alone at one-third of a GPT
endpoint's raw throughput. This is a hypothesis, not a proven mechanism
(the only zero-cache condition is also the only Anthropic endpoint), but
it is the reading that explains the throttling with one rule instead of
two.

## Post-fix results (provisional — first trials of an in-flight re-run)

On 8 tasks matched by name across the two runs, cost fell **$8.36 →
$1.77 (4.71×)** and wall-clock **12,423 s → 1,085 s (11.45×)**.

| Metric | before (`4a955a858`) | after (`3bef1f6a`) |
|---|---|---|
| Cache reads as % of input | **0.0%** | **78.7%** (still climbing
toward the ~86% steady state) |
| `cost_usd_no_cache_discount / cost_usd` | **1.00×** | **2.18×**
(tracking the projected ~2.5×) |
| Trials with a fatal 429 (same `-n 20`) | **63%** | **15–19%** |

To be clear about attribution: **~2× of that is the clean prefill saving
from caching itself**; the rest is second-order — cached requests burn
far less rate-limit budget, so they stall less and redo less destroyed
work. The 11.45× is a system-level result specific to this throttled
workspace, not a caching benchmark. Quality held (7/8 solved in each
run). A controlled low-`-n` A/B (neither arm hitting a 429), which the
`BUZZ_AGENT_PROMPT_CACHING` opt-out exists to enable, is still owed
before this becomes a published claim.

## What changed

### 1. Request caching (`llm.rs`, `config.rs`)

`anthropic_body()` emits ephemeral `cache_control` breakpoints, gated by
`BUZZ_AGENT_PROMPT_CACHING` (**default on**, `=0` to opt out):

- **Static prefix** — marker on the `system` block. Prefix order is
`tools → system → messages`, so this single marker caches **tools +
system** together. Byte-identical on every turn of a run, and survives a
context handoff (system/tools come from cfg/mcp, not `self.history`).
- **Rolling tail + leapfrog** — marker on the last block of the last
**two** messages. The append-only history re-reads the prior turn's
prefix from cache; marking two messages (not one) keeps consecutive
breakpoints inside Anthropic's **20-block lookback window** even as tool
parallelism rises, avoiding a silent full-price miss.

An empty system prompt stays a bare string (Anthropic rejects empty text
blocks), and below-threshold prefixes are silently not cached, so the
flag is safe on by default.

### 2. Surface the cache split end-to-end — the plumbing (`types.rs`,
`llm.rs`, `agent.rs`, `lib.rs`, `usage.rs`, `acp.rs`)

This is the part that makes gaps like the one above **visible** instead
of silent. A consumer that prices all of `input_tokens` at the full rate
can't tell a route that's caching from one that isn't — the total looks
right either way. So:

- `LlmResponse` gains `cached_input_tokens` (a **subset** of
`input_tokens`, never an addition); `parse_anthropic` / `parse_openai` /
`parse_responses` each populate it.
- A `usage_first()` helper reads the cache count wherever a provider
hides it — flat `cache_read_input_tokens` (Anthropic),
`prompt_tokens_details.cached_tokens` (OpenAI chat),
`input_tokens_details.cached_tokens` (Responses) — taking the **first
present value, never a sum**. Reading only flat keys is exactly why the
OpenAI route's nested `cached_tokens` had *also* been going unclaimed:
`prompt_tokens` is already inclusive, so the total looked correct while
the discount silently went unreported.
- The per-turn/per-session accumulators and the goose `usage_update`
payload now carry `accumulatedCachedInputTokens`; `buzz-acp`
deserializes it (`serde` default `0` for goose, which doesn't send it)
and logs `cached=<n>`.

### 3. Fix a Databricks MLflow-route double-count (`llm.rs`)

The Databricks MLflow route reports the flat Anthropic-spelled
`cache_read_input_tokens` *alongside* an already-inclusive
`prompt_tokens`, so the old code summed them and nearly doubled the
count — inflating both the context-budget gate and cost.
`openai_chat_input_tokens()` now reads `prompt_tokens` alone. Verified
on a live `databricks-glm-5-2` response where `prompt_tokens +
completion == total` proves inclusivity. (Anthropic's native route
genuinely *excludes* the cache fields and is still summed — the two
never collide, because `claude*` models route to the Anthropic path.)

### 4. Proxy + TLS-trust passthrough into MCP tools (`mcp.rs`) —
independent fix

`buzz-agent` `env_clear()`s each MCP child, and the allowlist carried no
proxy/TLS vars. On a proxy-only host that doesn't degrade the tools, it
**blinds** them: apt, curl, pip, git connect directly, the egress
firewall resets the socket, and the agent reports "Connection reset by
peer" — indistinguishable from a genuinely offline task. Adds both
spellings of `HTTP(S)_PROXY`/`NO_PROXY`/`ALL_PROXY` (curl/git read
lowercase; Go/Python read uppercase; libcurl ignores uppercase
`HTTP_PROXY`) plus `SSL_CERT_FILE`/`SSL_CERT_DIR` for TLS-terminating
proxies that present their own CA.

## Testing

- `cargo fmt --all -- --check`, `cargo clippy -p buzz-agent -p buzz-acp
--all-targets -- -D warnings` — clean.
- `cargo test -p buzz-agent -p buzz-acp` — **all green** (632 + 299 lib
tests plus integration suites, 0 failures). New tests cover: the three
breakpoints and the disabled/empty-system/single-message edge cases;
nested-vs-flat cache parsing for all three routes; the Databricks
inclusive-`prompt_tokens` fix; wire deserialization of
`accumulatedCachedInputTokens`; and the proxy/TLS passthrough allowlist.
- Pre-push lefthook suite green (branch-skew, rust-tests, test,
desktop-check/test/tauri).

## Relationship to the benchmark branch

These are the non-`benchmarks/` changes from
`benchmark/harness-accounting-and-solo`, lifted onto a clean base off
`main` so they can merge independently.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Atish Patel <atish@squareup.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-29 09:36:09 -04:00
klopez4212andGitHub 485d03a358 Fix mobile attachment and gallery polish (#3370)
## Summary

- align mobile message metadata and enlarge attachment-menu content
- smooth keyboard-to-camera/photo transitions and initialize the iOS
photo grid at the intended scale
- fix horizontal gallery loading, edge overflow, and end spacing

## Why

The attachment surfaces were reacting to keyboard and compact-menu
geometry during presentation, while gallery clipping and image lifecycle
behavior caused misalignment and occasional blank previews.

## Testing

- `just mobile-check`
- `flutter test` (881 passed, 1 skipped)
- native `RunnerTests` (17 passed)
- verified standalone Release build on a physical iPhone

---------

Signed-off-by: kenny lopez <klopez4212@gmail.com>
2026-07-29 07:09:08 +01:00
6300a6b1d0 fix(acp): per-runtime env defaults at spawn — isolate Hermes from configured MCP startup (#3420)
## Summary

- add a generic per-runtime env-defaults table,
`config::default_agent_env()`, mirroring the existing
`default_agent_args()` / `codex_network_env()` precedent, and merge it
once in `AcpClient::spawn` with the established precedence: **runtime
defaults < persona `extra_env` < inherited parent env**
- first (and only) row: Buzz-owned Hermes processes get
`HERMES_ACP_SKIP_CONFIGURED_MCP=1`, so Hermes does not preload unrelated
profile-configured MCP servers before answering ACP `initialize` (fixes
the 10s model-discovery timeout in #3355 — Buzz supplies session MCP
servers explicitly through `session/new`, per Hermes's documented
host-integration contract for this variable)
- normalize Windows `.cmd`/`.bat` shims alongside `.exe` in
`normalize_agent_command_identity` (npm installs resolve to those
wrappers)
- switch the `extra_env` parent-presence check from `var()` to
`var_os()` so non-UTF-8 parent values are honored

Replaces the runtime-specific approach in #3386: same behavior, but the
mechanism is generic runtime spawn metadata in `config.rs` rather than a
Hermes/ACP special case in `acp.rs`, and the seam covers every launch
path (Desktop spawn, `buzz-acp models`, CLI) because they all funnel
through `AcpClient::spawn`. ~15 lines of production code.

Fixes #3355

## Testing

- `cargo test -p buzz-acp` — **639 passed, 0 failed** (full package,
includes the new `default_agent_env_recognizes_hermes_identities` unit
test and `spawn_applies_runtime_env_defaults_with_extra_env_precedence`
integration test covering default injection, extra_env override, and
non-Hermes exclusion)
- `cargo fmt --all -- --check`, `cargo clippy -p buzz-acp --all-targets
-- -D warnings` — clean
- live-local with real Hermes v0.19.0 (`hermes-acp`): `buzz-acp models`
returned **13 models / currentModelId in 2.6–3.0s** (was a 10.0s timeout
on the first cold run without isolation); a wrapper probe confirmed the
child received `HERMES_ACP_SKIP_CONFIGURED_MCP=1` by default and `0`
when the parent env set it explicitly (operator wins)
- lefthook pre-push suite green: rust-tests, desktop-check,
desktop-test, desktop-tauri-test, mobile-test, branch-skew

No UI changes; subprocess environment behavior only.

Signed-off-by: npub1qyvc0c5kl4gqv2fd97fsk46tu378sqgy35vc83rvgfwne90sel7s0ed67d <011987e296fd5006292d2f930b574be47c7801048d1983c46c425d3c95f0cffd@buzz.block.builderlab.xyz>
Signed-off-by: Tyler Longwell <tlongwell@block.xyz>
Co-authored-by: npub1qyvc0c5kl4gqv2fd97fsk46tu378sqgy35vc83rvgfwne90sel7s0ed67d <011987e296fd5006292d2f930b574be47c7801048d1983c46c425d3c95f0cffd@buzz.block.builderlab.xyz>
Co-authored-by: Tyler Longwell <tlongwell@block.xyz>
Co-authored-by: mr-r0b0t.eth <adam.manning@pro-serveinc.com>
2026-07-28 22:51:19 -04:00
22be8bb351 fix(relay): avoid subscription lock inversion (#3413)
## Summary
- drop the `subs` DashMap guard before mutating subscription indexes
- snapshot fan-out candidate vectors so index guards are dropped before
looking up `subs`
- add concurrent fan-out/replacement regression coverage

## Why
`fan_out_scoped` previously held an index guard while `push_match`
acquired `subs`, while CLOSE and same-ID replacement held `subs` while
removing from an index. The reverse ordering made an AB/BA deadlock
reachable and could synchronously park all Tokio workers.

## Validation
- `rustup run 1.95.0 cargo test -p buzz-relay` — 769 library tests
passed, 33 ignored; 11 binary tests passed; doc tests passed
- push hooks with pinned Rust 1.95 — branch-skew, repository Rust
suites, and desktop Tauri suite passed
- `git diff --check`

## Residual risk
Fan-out now clones bounded candidate vectors before matching. This adds
allocation/copy cost proportional to the indexed candidate set, in
exchange for eliminating nested DashMap guards. This fixes the concrete
lock cycle but does not prove every observed production wedge had this
cause.

---------

Signed-off-by: npub12gtutshhh76rx0jx697f32f9tffd4hhp3hx58fp4x6u4uemkm7sqf8f757 <5217c5c2f7bfb4333e46d17c98a9255a52dadee18dcd43a43536b95e6776dfa0@buzz.block.builderlab.xyz>
Signed-off-by: npub1jh9wn95s0472h86ahapupaf7m6kx4v9sx2n0atj2hltcfer8k06s5n3pyf <95cae996907d7cab9f5dbf43c0f53edeac6ab0b032a6feae4abfd784e467b3f5@buzz.block.builderlab.xyz>
Co-authored-by: npub12gtutshhh76rx0jx697f32f9tffd4hhp3hx58fp4x6u4uemkm7sqf8f757 <5217c5c2f7bfb4333e46d17c98a9255a52dadee18dcd43a43536b95e6776dfa0@buzz.block.builderlab.xyz>
Co-authored-by: npub1jh9wn95s0472h86ahapupaf7m6kx4v9sx2n0atj2hltcfer8k06s5n3pyf <95cae996907d7cab9f5dbf43c0f53edeac6ab0b032a6feae4abfd784e467b3f5@buzz.block.builderlab.xyz>
2026-07-28 21:13:28 -04:00
90e058ebf6 feat: add explicit entry for claude-opus-5 in model config (#2831)
Fixes #2787

- Added `claude-opus-5` to `config.rs` model classification and adaptive
effort helpers.
- Updated fixture test configurations to cover `claude-opus-5`.
- Verified with `cargo test` and JS unit tests.

Signed-off-by: Apurva Shaw <apurvashaw@Apurvas-MacBook-Air.local>
Co-authored-by: Apurva Shaw <apurvashaw@Apurvas-MacBook-Air.local>
2026-07-28 22:56:20 +00:00
WesandGitHub 55a3ed7b92 fix(desktop): clear stale thread new-message pill (#3411)
## Summary
- reconcile anchored-scroll state when passive layout changes put a
thread at its physical floor
- route thread composer-padding growth and shrink through the same
hook-owned settlement path
- preserve pinned thread targets while clearing stale new-message state

## Root cause
Thread bottom state was updated primarily by native `scroll` events.
Deferred replies, viewport changes, and composer-overlay padding can
finish changing geometry after the user's last scroll—or after the
initial open pin—without another scroll event. The thread could visibly
reach the floor while `isAtBottom` and `newMessageCount` remained stale,
leaving the “N new messages” pill visible.

## Verification
- `pnpm check`
- `pnpm typecheck`
- `pnpm test` — 3,768 passed
- push hook: branch-skew, Desktop check, and Desktop full unit suite
passed

Signed-off-by: Wes <wesbillman@users.noreply.github.com>
2026-07-28 15:35:10 -07:00
9227bdf58a fix(ci): ratchet file sizes against the base tree (#3352)
## Summary

- replace the whole-tree file-size gate with a stateless differential
ratchet
- allow inherited files over 1,000 lines to hold or shrink, but never
grow
- delete the 44-entry numeric override ledger and run the same policy
across Desktop, Web, and Mobile CI
- fail closed when the local base cannot be resolved and cover policy,
Git status parsing, and base resolution in unit tests

This removes the shared mutable policy state that caused unrelated PRs
to fail after neighboring merges. It does **not** by itself prevent two
stale green PRs from becoming invalid when combined; that requires merge
queue or up-to-date branch enforcement.

### Related issue

None found. This follows the design discussion in the linked Buzz
channel.

### Testing

- `node --test scripts/check-file-sizes-core.test.mjs` (6/6)
- Desktop, Web, and Mobile ratchet entrypoints
- `just desktop-check`
- `just web-check`
- Mobile analysis
- `git diff --check`

The repository pre-push suite also exposed an unrelated existing Mobile
widget failure in `ChannelDetailPage keeps follow mode off while a tall
newest message stays visible`; it reproduces in isolation and this
branch does not touch Mobile widget behavior.

Signed-off-by: Wes <wesbillman@users.noreply.github.com>
Co-authored-by: Carl <c7ebe626f000404285d3686e1dc74cc07cc60a9754a150041ba132e14bd3e2ec@buzz.block.builderlab.xyz>
2026-07-28 22:17:36 +00:00
Will PflegerandGitHub 826bed4821 chore(ci): bump desktop smoke E2E timeout to 30 minutes (#3409)
Three main-branch runs today had shards killed at exactly 20m17s
("exceeded the maximum execution time of 20m0s"); the killed shard was
actively passing tests seconds before the cap. Shard runtime has grown
to the limit. 30 matches the other desktop jobs in the same workflow.

Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
2026-07-28 17:59:02 -04:00
12d63c67be release(chart): publish 0.1.7 (#3393)
## Why
Publish chart 0.1.7 after the feature PR merged from a fork and
therefore intentionally skipped the internal-branch auto-tag job.

## What
- Trigger the `chart-release/0.1.7` release lane
- Update the quickstart example to reference chart 0.1.7

## Risk Assessment
Low — the chart implementation is already merged and tested; this PR
creates its immutable release tag and OCI artifact.

## References
- Chart implementation: https://github.com/block/buzz/pull/3322
- `helm unittest` 0.8.2: 43/43 tests passed
- Local pre-push checks passed

Generated with Amp

Signed-off-by: David Grochowski <dgrochowski@squareup.com>
Co-authored-by: Amp <amp@ampcode.com>
chart-v0.1.7
2026-07-28 14:44:53 -07:00