Files
buzz/scripts/MODEL_CAPABILITIES.md
T
899f55cc8d refactor(models): Phase 3 — retire old hand tables and shrink manifest apparatus (#4589)
## Summary

Retires the old hand-table authorities and transitional verification
scaffolding from the model-capability manifest arc. All production
routes now run exclusively through the generated interpreters introduced
in Phase 1 ([#3821](https://github.com/block/buzz/pull/3821)) and wired
in Phase 2 ([#3958](https://github.com/block/buzz/pull/3958)).

Stack: [#3821](https://github.com/block/buzz/pull/3821) →
[#3958](https://github.com/block/buzz/pull/3958) → this PR
Base: [#3603](https://github.com/block/buzz/pull/3603)

## What the manifest system is now

**Source of truth:** `scripts/model-capabilities.json`
**Generator:** `scripts/generate-model-capabilities.mjs` — emits Rust
and TS interpreters only (coverage JSON output removed)
**Generated interpreters:**
`crates/buzz-agent/src/generated_model_capabilities.rs` (+ normative
tests), `desktop/src/features/agents/ui/modelCapabilities.ts`
**Label registry:** `generate-databricks-model-names.py` →
`databricks_model_names.rs` / `databricksModelNames.ts`
**Permanent gates:** `scripts/normative-corpus.json` +
`scripts/run-corpus.mjs` (both-interpreter equivalence, 51 vectors),
`scripts/test-manifest-validator.mjs` (schema, 24 cases), regen-diff job
inside `ci.yml`
**One doc:** `scripts/MODEL_CAPABILITIES.md`

## Deleted

**Old hand-table authorities (production):**
- `getProviderEffortConfig_oldHandTable()` and all supporting helpers
from `desktop/src/features/agents/ui/buzzAgentConfig.ts`
- `normalize_effort_for_openai_route()`,
`_old_anthropic_thinking_config_for_databricks_v2()`, test-only
re-export wrappers from `crates/buzz-agent/src/config.rs`
- `strip_catalog_prefix()`, `anthropic_thinking_config()`,
`anthropic_model_supports_xhigh()`, `clamp_adaptive_effort()`,
`anthropic_efforts_for_model()`, `is_manual_budget_model()`,
`is_adaptive_thinking_model()`, `gpt5_token_matches()`,
`gpt5_base_matches()`, `openai_efforts_for_model()` from
`crates/buzz-agent/src/config.rs` — all superseded by generated
interpreter
- Old DBv2 body-level tests, `_OLD_DATABRICKS_V2_*` constants,
`model_name_segments()`, `_old_databricks_v2_route_for_model()`, all
Phase-2 behavioral differential test functions from
`crates/buzz-agent/src/llm.rs`

**Transitional scaffolding:**
- `scripts/run-differential.mjs` — old-vs-new JS differential harness
- `scripts/run-mutation-evidence.mjs` — one-time mutation evidence
runner
- `desktop/src/features/agents/ui/effortTable.fixture.json` — Phase-2
TS/Rust sync fixture
- `desktop/src/features/agents/ui/effortTable.fixture.test.mjs` —
fixture sync guard
- `.github/workflows/model-capability-regen-diff.yml` — standalone
workflow (steps folded into `ci.yml`)

**One-time evidence and generated snapshots:**
- `scripts/MUTATION_EVIDENCE.md`,
`scripts/MODEL_CAPABILITIES_SCHEMA.md`,
`scripts/MODELS_DEV_RECONCILIATION.md` — consolidated into
`scripts/MODEL_CAPABILITIES.md`
- `scripts/generated-model-capabilities-coverage.json` — full-table
snapshot (generator no longer emits it)
- `scripts/catalog-sample-fixture.json` — models.dev snapshot used only
by the deleted differential harness

## Verification

- `cargo test -p buzz-agent --lib` with `RUSTFLAGS="-D warnings"`:
**337/337** (clean — no dead_code warnings)
- `node --experimental-strip-types scripts/run-corpus.mjs`: **51/51**
- `node scripts/generate-model-capabilities.mjs` + regen diff: **clean
(exit 0)**
- `node --test scripts/test-manifest-validator.mjs`: **24/24**
- `just clippy`: zero warnings, zero errors

---------

Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
Co-authored-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
2026-08-03 17:04:38 -04:00

8.2 KiB

Model Capabilities Manifest

Source of truth: scripts/model-capabilities.json Generator: scripts/generate-model-capabilities.mjs Emitted artifacts:

  • crates/buzz-agent/src/generated_model_capabilities.rs
  • desktop/src/features/agents/ui/modelCapabilities.ts

How to regenerate

node scripts/generate-model-capabilities.mjs

CI regenerates and diffs on every PR that touches the manifest, generator, or generated files. Any stale generated file fails the model-capabilities job in ci.yml.

Resolver contract (plan v4)

Resolution is a total function resolve(provider, raw_model_id) → CapabilityResult. Three ordered steps:

  1. Provider-qualified raw exact lookup — key is (provider, raw_model_id), matched on the RAW ID before any prefix stripping. A prefixed alias never inherits an exact record.
  2. Provider-scoped ordered family rules — on the normalized (prefix-stripped) alias. Rules ordered by match_priority descending (higher wins). Each rule is tagged with the providers it applies to.
  3. Per-axis provider fallback — blank (empty model string) vs concrete_unknown (nonblank but unmatched), per provider.

CapabilityResult is complete — every axis is populated. Runtime consumers never compose fields from multiple tiers.

Axes (schema fields)

Axis Type Notes
registry_label string | null Optional static display label. Feeds resolveModelLabel() registry tier only.
thinking_mode enum manual-budget | adaptive | omit-fields | none | not-applicable
supported_efforts ThinkingEffort[] Non-empty. UI effort dropdown options.
default_effort ThinkingEffort | null null = "Inherit" (Anthropic manual-budget models).
databricks_v2_wire_route enum openai-responses | anthropic-messages | mlflow-chat | route-unknown | not-applicable
normalization_policy enum none | openai-standard | openai-clamp-max-to-xhigh

thinking_mode values

Value Meaning
manual-budget thinking:{type:"enabled", budget_tokens} -- claude-3*, claude-opus-4-5
adaptive thinking:{type:"adaptive"} + output_config:{effort} -- opus-4-6+, sonnet-4-6+, etc.
omit-fields Unknown Anthropic model -- omit thinking fields rather than guess request shape
none Non-Anthropic-routed model -- thinking fields not applicable
not-applicable Provider does not use Anthropic thinking API

databricks_v2_wire_route values

Scoped to DBv2 only. All non-DBv2 providers emit not-applicable. Transport for pure OpenAI, legacy Databricks, and OpenRouter is selected by OpenAiApi / openai_request() at runtime.

Value Meaning
openai-responses /ai-gateway/openai/v1/responses
anthropic-messages /ai-gateway/anthropic/v1/messages
mlflow-chat /ai-gateway/mlflow/v1/chat/completions
route-unknown DBv2 blank model -- route not yet determinable
not-applicable Not a DBv2 provider

Family rule match kinds

Kind Semantics
exact Case-insensitive exact string equality on normalized alias
prefix Normalized alias starts with match_value
gpt5-token Boundary-aware token: present at end-of-string or followed by - (not digit/letter)
gpt5-base Like gpt5-token but also rejects -<1-3 digit> suffixes (version-number rejection)
segment Normalized alias contains match_value as a full alphanumeric segment (split on non-alnum)
segment-prefix Any segment of the normalized alias starts with match_value

Boundaries the manifest does NOT own

  • Transport/endpoint selection for pure OpenAI, legacy Databricks, OpenRouter: OpenAiApi and openai_request() remain authoritative. The databricks_v2_wire_route axis is DBv2-only.
  • Final display labels: resolveModelLabel(discovered_name, registry_label, raw_id) three-tier precedence is authoritative. The manifest's registry_label feeds only the static registry tier.
  • llm.rs replacement scope: only databricks_v2_route_for_model. Other dispatch paths remain.

Reconciliation policy (plan v4 §Behavior policy)

Not purely behavior-preserving. models.dev reasoning_options become exact overrides. Each divergence from family rule results is reconciled against provider docs and either:

  • (a) adopted as an intentional correction with its own test + exact record, or
  • (b) rejected with a curation note in the exact record.

models.dev reconciliation table

Source queried: https://models.dev/api.json (2026-07-31) Payload SHA-256: d5a4974cd69f19b0f67713acaa6bb3b16e920defdc07ecbdf6b0a936181bb0e0 Verbatim source snapshot SHA-256 (catalog-sample-fixture.json, deleted in Phase 3): dc4092a04392f258bea65de2cef53cb1902dce1779dc2b1b2e21fb56774f2d78

databricks-gpt-5-4-mini

Family rule (gpt5-4) models.dev Disposition
supported_efforts [none, low, medium, high, xhigh] [low, medium, high] ADOPT

Provider-advertised wins per plan F1 policy. Source: providers.databricks.models["databricks-gpt-5-4-mini"].reasoning_options (retrieved 2026-07-31).

databricks-gpt-5-4-nano

Family rule (gpt5-4) models.dev Disposition
supported_efforts [none, low, medium, high, xhigh] [low, medium, high] ADOPT

Same as databricks-gpt-5-4-mini.

databricks-gpt-5-6-sol

Family rule (gpt5-6) models.dev Disposition
supported_efforts [none, low, medium, high, xhigh, max] [low, medium, high, max] ADOPT

Provider-advertised wins. Source: providers.databricks.models["databricks-gpt-5-6-sol"].reasoning_options (retrieved 2026-07-31).

databricks-gpt-5-5

Family rule (gpt5-5) models.dev Disposition
supported_efforts [none, low, medium, high, xhigh] [low, medium, high] ADOPT

Provider-advertised wins. Source: providers.databricks.models["databricks-gpt-5-5"].reasoning_options (retrieved 2026-07-31).

databricks-claude-opus-4-7

Family rule models.dev Disposition
reasoning_options type effort-based budget_tokens NO EFFORT DIVERGENCE

models.dev advertises a different capability axis (extended thinking token budget), not an effort-level selector. No effort divergence to reconcile. Source: providers.databricks.models ["databricks-claude-opus-4-7"].reasoning_options (retrieved 2026-07-31).

Non-divergences (confirmed consistent)

Model family Source Status
claude-opus-4-7, claude-opus-4-8 Anthropic extended-thinking docs (July 2025) ok
claude-sonnet-5.*, claude-fable-5, claude-mythos-5 Anthropic extended-thinking docs (July 2025) ok
claude-opus-4-6, claude-sonnet-4-6, claude-mythos-preview Anthropic extended-thinking docs (July 2025) ok
claude-3* Anthropic extended-thinking docs (July 2025) ok
gpt-5-pro, gpt-5.6, gpt-5.5, gpt-5.4, gpt-5.1, gpt-5 OpenAI reasoning guide (July 2025) ok

Mutation evidence (historical record)

Mutation testing was run at Phase 2 completion (2026-07-31). 7 generator mutations were applied in isolation against both TS and Rust interpreters. All 7 were killed by both interpreters (7/7). The mutation runner (scripts/run-mutation-evidence.mjs) was deleted in Phase 3; the normative corpus (scripts/normative-corpus.json) that kills these mutations continues to run in CI.

Adding a new model family

  1. Add a family_rules entry with a new unique id, appropriate match_kind, providers, match_priority, and all capability axes.
  2. Run node scripts/generate-model-capabilities.mjs to regenerate artifacts.
  3. CI verifies byte-clean regeneration.
  4. The normative corpus (scripts/normative-corpus.json) may need new vectors.

Adding an exact model override

  1. Add an exact_records entry with provider + raw_model_id (the full raw ID, no prefix stripping). Include a _reconciliation note and doc citation.
  2. Run node scripts/generate-model-capabilities.mjs -- completeness validator will fail if any axis cannot be resolved.
  3. Regenerate and commit.