mirror of
https://github.com/block/buzz.git
synced 2026-08-18 06:50:31 +02:00
## What this does Introduces the model-capability manifest infrastructure (Phase 1 of the Model-Capability Manifest plan v4, Thufir-approved 9/9/9). No consumer cutover — `config.rs`, `llm.rs`, `catalog.rs`, and `buzzAgentConfig.ts` are unchanged. Phase 2 wires them. **Single source of truth** replaces hand-mirrored metadata across four files: ``` scripts/model-capabilities.json → hand-curated manifest scripts/generate-model-capabilities.mjs → emits Rust + TS artifacts crates/buzz-agent/src/generated_model_capabilities.rs desktop/src/features/agents/ui/modelCapabilities.ts ``` ## Resolver contract (plan v4 §Resolver contract) Total function `resolve(provider, raw_model_id) → CapabilityResult`. Three ordered steps: 1. Provider-qualified raw exact lookup — key is `(provider, raw_model_id)`, matched before any prefix stripping. A prefixed alias never inherits an exact record. 2. Provider-scoped ordered family rules — on normalized (prefix-stripped) alias, by `match_priority` desc. 3. Per-axis provider fallback — `blank` vs `concrete_unknown`, per provider. Result is complete — every axis populated, runtime consumers never compose fields. ## Boundaries the manifest does NOT own - Transport for pure OpenAI, legacy Databricks, OpenRouter: `OpenAiApi`/`openai_request()` remain authoritative. `databricks_v2_wire_route` is DBv2-only (all other providers emit `not-applicable`). - Final display labels: `resolveModelLabel()` three-tier precedence unchanged. `registry_label` feeds only the static registry tier. - `llm.rs` scope: only `databricks_v2_route_for_model` (Phase 2). ## Test oracle (three independent layers) 1. Generated full-table coverage — `scripts/generated-model-capabilities-coverage.json`: every manifest entry + provider fallbacks. 2. Hand-authored normative corpus — `scripts/normative-corpus.json` (44 vectors): Anthropic manual-budget/adaptive families, OpenAI gpt-5 adversarial boundary cases, DBv2 segment-routing collision tests, P2-A resolver-contract vectors, P2-B blank/concrete-unknown per provider. Runs against JS resolver (`run-corpus.mjs`) and mirrored in Rust (`generated_model_capabilities_tests.rs`). 3. Schema-negative tests — `scripts/test-manifest-validator.mjs` (17 tests): every validator rule has a failing-input test. ## Reconciliation table `scripts/MODELS_DEV_RECONCILIATION.md` — all models.dev divergences dispositioned. `databricks-gpt-5-4-mini` and `databricks-gpt-5-4-nano` adopt models.dev `[low,medium,high]` (family rule adds `none+xhigh` the endpoint doesn't advertise). ## CI `.github/workflows/model-capability-regen-diff.yml`: triggers on manifest/generator/artifact changes; regenerates and fails if stale; runs JS corpus + schema-negative tests. ## Acceptance criteria (plan v4 Phase 1) - Byte-clean regen: `node scripts/generate-model-capabilities.mjs --check` passes - Rust compiles: `cargo check -p buzz-agent` - TS typechecks: `pnpm tsc --noEmit --strict` - 44/44 normative corpus vectors pass (JS interpreter) - 41/41 Rust corpus tests pass - 17/17 schema-negative tests pass (every validator rule) - Reconciliation table complete with doc citations - No consumer changes (config.rs, llm.rs, catalog.rs, buzzAgentConfig.ts untouched) --------- Signed-off-by: Will Pfleger <pfleger.will@gmail.com> Co-authored-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
2.9 KiB
2.9 KiB
Model-Capability Manifest — Mutation Evidence
Interpreter coverage: both generated interpreters are exercised per mutation fault.
- TypeScript:
scripts/run-corpus.mjsimportsresolveModelCapabilities()fromdesktop/src/features/agents/ui/modelCapabilities.tsvia--experimental-strip-types. - Rust:
cargo test -p buzz-agent -- generated_model_capabilities::tests::shared_corpus_testsdeserializes and executes every vector inscripts/normative-corpus.jsonagainstresolve_model_capabilities().
How to reproduce
# Runs generator mutations; exercises both TS and Rust interpreters per fault
node --experimental-strip-types scripts/run-mutation-evidence.mjs
# Run interpreters independently:
node --experimental-strip-types scripts/run-corpus.mjs
cargo test -p buzz-agent -- generated_model_capabilities::tests::shared_corpus_tests
Mutation run results (both interpreters)
All 7 mutations applied in isolation; manifest restored after each run. Each mutation must be detected (killed) by both interpreters for it to count as covered.
| ID | Mutation | Expected killer(s) | TS | Rust |
|---|---|---|---|---|
| M1 | Reduce claude-opus-4-7 supported_efforts to [low,medium,high] (drops xhigh+max) |
anthropic-claude-opus-4-7, dbv2-claude-prefix-stripped, dbv2-claude-route-anthropic-messages |
killed ✓ | killed ✓ |
| M2 | Add xhigh to gpt5-base supported_efforts |
openai-gpt5-base, openai-gpt5-1106-should-not-match-base, openai-gpt5-4o-matches-base, openai-gpt5-date-suffix |
killed ✓ | killed ✓ |
| M3 | Change gpt5-1 default_effort to "high" instead of "none" |
openai-gpt5.1 |
killed ✓ | killed ✓ |
| M4 | Swap dbv2-claude-code-names-segment route from anthropic-messages to openai-responses |
dbv2-goose-opus-5-is-anthropic |
killed ✓ | killed ✓ |
| M5 | Remove all three DBv2 segment rules | dbv2-goose-opus-5-is-anthropic, dbv2-consolidated-llama-not-sol, dbv2-terraform-coder-not-terra |
killed ✓ | killed ✓ |
| M6 | Change databricks_v2 concrete-unknown fallback route from mlflow-chat to openai-responses |
dbv2-concrete-unknown-mlflow-no-max |
killed ✓ | killed ✓ |
| M7 | Remove xhigh from gpt5-4 supported_efforts |
resolver-prefixed-alias-misses-exact |
killed ✓ | killed ✓ |
Summary: 7/7 mutations killed in both TS and Rust interpreters.
Coverage gaps
- Provider fallback mutations for
anthropic,openai,databricks,openrouter, and_defaultare not individually mutated. These are covered by explicit fallback vectors in the corpus foranthropic,openai, anddatabricks_v2. - Rust mutations are run by recompiling the mutated generated file per fault (via
cargo testafternode generate-model-capabilities.mjs). Compile time is acceptable for offline mutation runs; CI only runs the already-compiled shared corpus harness.