The branch had ratcheted the desktop file-size gate four times. One of those
entries re-declared `src-tauri/src/lib.rs` at 1001 while a pre-existing entry
at line 58 sets 1013 — the override map is a JS `Map`, so the later key won and
the branch silently lowered a ceiling it never needed to touch (lib.rs is 920).
Splitting at existing seams removes the other two: the M1 schema migration
moves to `archive/store_migrations.rs` behind one `pub(super)` entry point, and
the kind-44200 archive tests move to `archive/mod_agent_metric_tests.rs`,
`#[path]`-included from `mod_tests.rs` so they keep reaching the shared
fixtures through `use super::*`. `mod_tests.rs` returns to its pre-branch 1208
ceiling rather than dropping it, since that debt predates this work.
`formatCoverageDate` was duplicated verbatim in both usage components; it now
lives beside the other formatters in `lib/agentUsage.ts`.
Co-authored-by: Will Pfleger <pfleger.will@gmail.com>
Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
Stack: #2035 → this PR
## What
Group the per-agent model breakdown by `(harness, model)` instead of
`model` alone, so the same model running under two different harnesses
(e.g. `claude-sonnet` via `goose` vs `claude-code`) produces two
distinct rows rather than collapsing into one.
## Changes
**Rust / SQLite**
- `agent_metric_index`: add nullable `harness TEXT` column. Schema
migration M1: `ALTER TABLE ... ADD COLUMN` + delete-then-backfill
rebuild so all existing rows get `harness` populated from retained
`archived_events.raw_json` — no data loss; ingest and backfill share the
same `from_payload` parser (frozen-plan requirement preserved).
- `AgentMetricIndexRow`: parse `harness` from
`AgentTurnMetricPayload.harness` (REQUIRED field per NIP-AM), write/read
through all store paths (`ROW_COLUMNS`, INSERT, `row_from_sql`).
- `AgentScope.models` grouping key widened from `Option<String>` to
`(Option<String>, Option<String>)` i.e. `(harness, model)`. Sort
tiebreak: harness ascending → model ascending, `None` last in each.
- `ModelUsage` wire type: add `harness: Option<String>` field.
**Frontend**
- `tauriArchive.ts` / `bridge.ts`: add `harness` to `AgentUsageModel`
type.
- `AgentUsageFocusedView`: render `harness` as a dimmed sub-label next
to each model name on breakdown rows. `null` harness (pre-migration data
or unknown) renders no label — single-harness data is visually
unchanged.
- `agentUsage.ts` `sortModelsByKnownTotal`: tiebreak updated for
compound `(harness, model)` key.
**Tests**
- Rust: collapse-fix test (same model / two harnesses → two rows),
migration test (old-shape rows get harness populated after rebuild),
harness sort coverage.
- TS: `sortModelsByKnownTotal` harness tiebreak tests, same-model
two-harness sort test.
- E2E: `bridge.ts` fixture type updated, `agent-usage.spec.ts` harness
assertion, `agent-usage-screenshots.spec.ts` shot 02 harness-label
visibility check.
## Gates
All green locally:
- `just desktop-check` ✓
- `just desktop-typecheck` ✓
- `just desktop-build` ✓
- `just desktop-test` (3487 pass, 0 fail) ✓
- `cargo test` in `desktop/src-tauri` (1575 pass, 0 fail) ✓
---------
Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
Co-authored-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
## Why
The Codex adapter install-plan test depended on the process-global
login-shell PATH cache, making it flaky under concurrent or loaded CI
runs.
## What
- Thread an explicit probe PATH through adapter install planning
- Use a controlled PATH in Codex install-plan tests
- Update the existing desktop file-size allowance for the focused seam
## Risk Assessment
Low — production behavior keeps using the same augmented PATH; only
dependency injection and deterministic tests change.
## References
- Failure:
https://github.com/block/buzz/actions/runs/30110680608/job/89539202266
- Validation: `just desktop-tauri-test`; `cd desktop && pnpm check`;
full pre-push hooks
Generated with Codex
---------
Signed-off-by: npub1x4hk035p3p9q39a3fcrd2fe30lpkrhr5dwe0cqzzjphxyyh8m0gsq4vqap <356f67c681884a0897b14e06d527317fc361dc746bb2fc0042906e6212e7dbd1@buzz.block.builderlab.xyz>
Co-authored-by: npub1x4hk035p3p9q39a3fcrd2fe30lpkrhr5dwe0cqzzjphxyyh8m0gsq4vqap <356f67c681884a0897b14e06d527317fc361dc746bb2fc0042906e6212e7dbd1@buzz.block.builderlab.xyz>
Co-authored-by: Codex <noreply@openai.com>
## Why
Installing the Codex, Claude, or Goose desktop app does not install the
command-line harness Buzz needs. The current UI makes that distinction
unclear, links some missing-CLI states to adapter documentation, and can
report a successful install from the installer exit code even when
runtime discovery still fails. On Windows, Buzz also invokes Goose's
Bash installer, which writes the executable somewhere Buzz does not
discover.
## What
- distinguish missing vendor CLIs from missing or outdated ACP adapters
in runtime metadata and UI guidance
- link Codex, Claude Code, and Goose missing-CLI states to their
official CLI installation documentation
- explain in Settings, onboarding, and agent configuration that the
desktop app alone is not sufficient
- use Goose's official PowerShell installer on Windows
- refresh PATH and rediscover the requested runtime after installation,
keeping the control retryable if the runtime is still unavailable
- add Rust and Playwright regression coverage for Windows installer
selection, CLI/adapter guidance, false-success prevention, verified
installs, and onboarding copy
## Risk Assessment
Medium. This changes desktop onboarding and runtime installation
behavior. Successful installs now require the runtime catalog to verify
availability; previously hidden discovery failures will surface as
actionable errors instead of a false success state.
## References
- [Codex CLI installation](https://developers.openai.com/codex/cli/)
- [Claude Code
installation](https://code.claude.com/docs/en/getting-started)
- [Goose
installation](https://goose-docs.ai/docs/getting-started/installation/)
- Follow-up to #2563 and #2587
## Validation
- `just desktop-typecheck`
- `just desktop-test` — 3,455 passed
- focused Rust post-install verification tests
- focused Playwright Doctor/onboarding coverage (in progress; CI and
local sequential rerun will provide final results)
Generated with Codex
---------
Signed-off-by: Atish Patel <atish@squareup.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Goose <opensource@block.xyz>
## Problem
On Windows, `install_shell_command` wraps every install command in Git
Bash `-l -c`, including the Windows-specific `powershell.exe … irm
https://chatgpt.com/codex/install.ps1 | iex` command. A Git Bash login
shell prepends its POSIX dirs (`C:\Program Files\Git\usr\bin`) to PATH,
so when Codex's `install.ps1` shells out to bare `tar -xzf C:\…`, it
resolves Git's GNU tar (`/usr/bin/tar`) instead of Windows bundled
bsdtar. GNU tar parses `C:` as a remote host:
```
tar (child): Cannot connect to C: resolve failed
gzip: stdin: unexpected end of file
/usr/bin/tar: Child returned status 128
/usr/bin/tar: Error is not recoverable: exiting now
Downloaded Codex package archive did not contain the expected package layout.
```
Claude's installer doesn't hit the same failure because it doesn't shell
out to `tar`.
## Solution
On Windows, detect `powershell.exe` install commands and spawn them
natively (`Command::new("powershell.exe")`) instead of routing them
through Git Bash. The discriminator is a case-insensitive prefix check
on the first whitespace-delimited token — minimal and precise.
The native spawn preserves everything `install_shell_command` provides
that applies:
- `NPM_CONFIG_*` / `COREPACK` env strip + managed npm prefix env
- PATH composed from managed Buzz dirs + inherited process PATH (no
POSIX login-shell dirs)
- `CREATE_NO_WINDOW` so no console flash
- stdin null, piped-drain in `run_install_command`
- The retry/backoff/annotate logic is fully shared
The `-Command` body is split correctly at the boundary
(case-insensitive) and passed as a single argument to preserve pipes and
spaces inside the installer script call.
Non-PowerShell commands (e.g. `npm install -g …` adapter steps) continue
through the existing Git Bash path unchanged.
## Tests
6 unit tests:
1. `is_powershell_command` detection (positive + negative)
2. Routing: PowerShell → native spawn on Windows, non-PowerShell → Git
Bash
3. Unix: non-Windows path returns the shell command unchanged
(compile-time cfg)
4. `-Command` body preservation (no bash args in native spawn; body is
single arg)
Full `just desktop-tauri-test` suite: 1627/1627 passing. Windows CI will
validate end-to-end.
---------
Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
Co-authored-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
Surfaces per-agent NIP-AM turn-metric usage (tokens, cost, model
breakdown, 7d/30d bucket series) in the desktop app, reading from the
Phase 2 get_agent_usage_series archive command:
- agent-usage/lib/agentUsage.ts: pure client-side helpers -- bucket
boundary construction, bigint-safe token formatting, and the ranked
overview-row projection shared by both UI surfaces.
- agent-usage/hooks.ts: react-query wrapper around
get_agent_usage_series with the app's standard retry:1 default.
- agent-usage/ui/AgentUsageSection.tsx: ranked per-agent usage rows on
the Agents overview, with a 7d/30d window switch, retry-on-error,
and a collection-off banner (with or without retained-data coverage
copy) linking to Local Archive settings.
- agent-usage/ui/AgentUsageFocusedView.tsx: per-agent drill-down
(bucketed series + model breakdown) reached via an overview row or
the profile panel's Info-tab usage ingress row; wired into
AgentsView, UserProfilePanel, and ProfilePanelContext.
- Partial badge on rows whose total is a known lower bound
(has_unknown_usage), surfaced via the accounting engine's per-field
completeness contract.
Extends the Rust archive test suite (agent_usage_tests.rs) to close
gaps in the pure accounting engine: the f64 cost ladder had no direct
outcome.cost assertions, assign_bucket_index's start-inclusive/
end-exclusive edges were only exercised indirectly, validate_request's
chrono finite-range rejection was unexercised, and the A2 ranking
tiebreak (pubkey ascending on an exact totalTokens tie) had no test
distinguishing it from insertion order.
tests/e2e/agent-usage.spec.ts (wired into playwright.config.ts's smoke
project) exercises both UI surfaces against a mocked
get_agent_usage_series: ranked rows, window switching, click-through
to the focused view from both the overview row and the profile
panel's ingress row, retry recovery from a query error, the
collection-off banner and its settings deep link, and the Partial
badge.
Co-authored-by: Will Pfleger <pfleger.will@gmail.com>
Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
Co-authored-by: Will Pfleger <pfleger.will@gmail.com>
Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
Adds the Rust backend half of the NIP-AM local agent usage feature
(Rev 3 frozen plan): a rebuildable agent_metric_index parsed from
archived kind-44200 rows, a pure per-field accounting ladder
(agent_usage.rs) computing token/cost deltas with adjacent-cumulative
preference and direct-value fallback, and the get_agent_usage_series
Tauri command wiring backfill, orphan repair, collection-enabled
detection, and A13's hasArchivedEvidence into one series response.
commit_archive now indexes kind-44200 rows in the same transaction as
the canonical event insert, and upsert_archived_event/gc_orphaned_events
return/enforce the new-row and cascade-delete invariants (A5/A6) that
persistedAgentMetrics and the index's rebuildability depend on.
get_agent_usage_series' SQLite core is split into a plain sync fn so
it can be driven directly against an in-memory Connection in tests
without a Tauri AppState.
Co-authored-by: Will Pfleger <pfleger.will@gmail.com>
Signed-off-by: Will Pfleger <pfleger.will@gmail.com>