## What Phase 2 consumer cutover targeting the `duncan/databricks-model-label-registry` umbrella branch. Wires `crates/**` and `desktop/**` consumers to the generated capability module introduced in Phase 1 (#3821), while keeping old and new paths both live for differential testing. Phase 3 removes the old paths. ## Commits (boundary-separated) ### feat(agent): Phase 2a — wire Rust consumers to generated capability module (`crates/**`, `scripts/**`) - `catalog.rs`: `DATABRICKS_V2_KNOWN_MODELS` re-exported from the generated module — single source of truth. - `llm.rs`: `databricks_v2_route_for_model` delegates to `resolve_model_capabilities("databricks_v2", model)`. Old segment-based classifier preserved as `#[cfg(test)] _old_*` for the differential harness. New `databricks_v2_route_differential_old_vs_new` test confirms 100% agreement on all 20 route vectors. - `config.rs`: new `effort_table_fixture_differential_old_vs_new` test runs `resolve_model_capabilities` over the 36-entry `effortTable.fixture.json` and asserts old/new agree modulo a doc-cited allowlist (4 F1 corrections). - `scripts/run-differential.mjs`: JS differential harness over effortTable fixture + normative corpus + catalog-sample fixture. 85 checks, 0 unexpected divergences (5 allowlisted: 4 F1 corrections + goose-opus-5 anthropic route correction). - `scripts/MODELS_DEV_RECONCILIATION.md`: deferred MINOR from Phase 1 — 8 trailing-double-space line breaks replaced with `<br>`. ### feat(desktop): Phase 2b — cut TS consumers to generated model-capabilities module (`desktop/**`) - `buzzAgentConfig.ts`: adds `getProviderEffortConfigFromManifest(provider, model?)` — thin wrapper over `resolveModelCapabilities()` from `modelCapabilities.ts`. Maps `supportedEfforts → validValues` and `defaultEffort → defaultValue` (null preserved for manual-budget/Inherit). Old `getProviderEffortConfig()` and all hand-tables stay live for the differential harness; Phase 3 retires them. - `formatAgentModelLabel.ts`: registry-label lookup re-pointed from hand-maintained `databricksModelNames.ts` import to generated `DATABRICKS_MODEL_NAMES` exported from `modelCapabilities.ts`. Same Map shape, identical contents, behavior unchanged. ### fix(scripts): add ts-esm-loader and fix allowlist coverage in run-differential (`scripts/**`) - `scripts/ts-esm-loader.mjs`: minimal ESM custom loader that resolves extensionless relative TS imports. Required because Phase 2b's `buzzAgentConfig.ts` imports `modelCapabilities` without `.ts` extension — which Node's `--experimental-strip-types` runner cannot resolve without a hook. - `scripts/run-differential.mjs`: shebang updated to self-bootstrap with the loader; fixes the `totalAllowlisted` counter (was declared but never incremented — always printed `0 allowlisted`). Replaced with per-axis hit tracking: reports exercised slot count (`N/total`) in summary; fails with `STALE_ALLOWLIST` if any declared entry fires zero divergences, preventing stale entries from silently masking future regressions. ## Verification - `cargo test -p buzz-agent --lib`: 426/426 - Corpus: 45/45 · schema-negative: 24/24 · `--check` byte-clean - Differential: 85 checks, 0 unexpected divergences, 6/6 allowlist slots exercised - Desktop: 3847/3847 · typecheck clean · biome clean - Mobile: 1019 pass, 1 skipped — same 5 flaky tests in `mobile/test/features/channels/` that reproduce at the umbrella base; zero mobile files in this branch range - `git diff --check`: clean ## What Remains (Phase 3) Remove old hand-maintained paths: `_old_*` functions in `llm.rs`/`config.rs`, old `getProviderEffortConfig` tables in `buzzAgentConfig.ts`, old `databricksModelNames.ts` import in `formatAgentModelLabel.ts`, old `databricks_model_names.rs` module. --------- Signed-off-by: Will Pfleger <pfleger.will@gmail.com> Signed-off-by: npub1g8493u0xfsjrvflg4n08ezd7vec99mnwzlv0qgwpr9d7gvjwhuzqx59rhw <41ea58f1e64c243627e8acde7c89be667052ee6e17d8f021c1195be4324ebf04@buzz.block.builderlab.xyz> Signed-off-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz> Co-authored-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz> Co-authored-by: npub1g8493u0xfsjrvflg4n08ezd7vec99mnwzlv0qgwpr9d7gvjwhuzqx59rhw <41ea58f1e64c243627e8acde7c89be667052ee6e17d8f021c1195be4324ebf04@buzz.block.builderlab.xyz>
8.3 KiB
models.dev Reasoning Options Reconciliation Table
Source queried: https://models.dev/api.json (2026-07-31)
Payload SHA-256: d5a4974cd69f19b0f67713acaa6bb3b16e920defdc07ecbdf6b0a936181bb0e0
Policy (plan v4 §Behavior policy): models.dev reasoning_options become exact overrides.
Each divergence from the current family rule result is reconciled here: either (a) adopted as an
intentional correction or (b) rejected with a curation note.
Verbatim source snapshot: scripts/catalog-sample-fixture.json — verbatim id, name, and
nested reasoning_options objects captured from the live API without transformation.
Re-verify hash: curl -s https://models.dev/api.json | sha256sum
Divergences
databricks-gpt-5-4-mini
| Current family rule (gpt5-4) | models.dev | Disposition | |
|---|---|---|---|
supported_efforts |
[none, low, medium, high, xhigh] |
[low, medium, high] |
ADOPT |
Rationale: The Databricks AI Gateway v2 endpoint for databricks-gpt-5-4-mini explicitly
advertises only [low, medium, high] in its reasoning_options. The family rule's none and
xhigh are derived from the upstream OpenAI GPT-5.4 spec, which this Databricks endpoint does
not expose. Provider-advertised wins per plan F1 policy.
Source: https://models.dev/api.json — retrieved 2026-07-31; providers.databricks.models["databricks-gpt-5-4-mini"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]
Snapshot: scripts/catalog-sample-fixture.json key "databricks-gpt-5-4-mini"
Test vector: resolver-exact-raw-id-hit in scripts/normative-corpus.json
databricks-gpt-5-4-nano
| Current family rule (gpt5-4) | models.dev | Disposition | |
|---|---|---|---|
supported_efforts |
[none, low, medium, high, xhigh] |
[low, medium, high] |
ADOPT |
Rationale: Same as databricks-gpt-5-4-mini. The nano variant exposes the same restricted
effort set. Provider-advertised wins.
Source: https://models.dev/api.json — retrieved 2026-07-31; providers.databricks.models["databricks-gpt-5-4-nano"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]
Snapshot: scripts/catalog-sample-fixture.json key "databricks-gpt-5-4-nano"
databricks-gpt-5-6-sol
| Current family rule (gpt5-6) | models.dev | Disposition | |
|---|---|---|---|
supported_efforts |
[none, low, medium, high, xhigh, max] |
[low, medium, high, max] |
ADOPT |
Rationale: The Databricks AI Gateway v2 endpoint for databricks-gpt-5-6-sol advertises only
[low, medium, high, max] in its reasoning_options. The family rule's none and xhigh are
derived from the upstream OpenAI GPT-5.6 spec, which this Databricks endpoint does not expose.
Provider-advertised wins per plan F1 policy.
Source: https://models.dev/api.json — retrieved 2026-07-31; providers.databricks.models["databricks-gpt-5-6-sol"].reasoning_options = [{"type":"effort","values":["low","medium","high","max"]}]
Snapshot: scripts/catalog-sample-fixture.json key "databricks-gpt-5-6-sol"
databricks-gpt-5-5
| Current family rule (gpt5-5) | models.dev | Disposition | |
|---|---|---|---|
supported_efforts |
[none, low, medium, high, xhigh] |
[low, medium, high] |
ADOPT |
Rationale: The Databricks AI Gateway v2 endpoint for databricks-gpt-5-5 advertises only
[low, medium, high] in its reasoning_options. The family rule's none and xhigh are
derived from the upstream OpenAI GPT-5.5 spec, which this Databricks endpoint does not expose.
Provider-advertised wins per plan F1 policy.
Source: https://models.dev/api.json — retrieved 2026-07-31; providers.databricks.models["databricks-gpt-5-5"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]
Snapshot: scripts/catalog-sample-fixture.json key "databricks-gpt-5-5"
databricks-claude-opus-4-7
| Current family rule (anthropic-adaptive-xhigh-opus-4-7) | models.dev | Disposition | |
|---|---|---|---|
reasoning_options type |
effort-based | budget_tokens |
NO EFFORT DIVERGENCE |
Rationale: models.dev advertises reasoning_options=[{"type":"budget_tokens","min":1024}] —
a different capability axis (extended thinking token budget), not an effort-level selector.
There is no effort divergence to reconcile. The effort capabilities for this model come from the
anthropic-adaptive-xhigh-opus-4-7 family rule (Anthropic extended-thinking support table).
Source: https://models.dev/api.json — retrieved 2026-07-31; providers.databricks.models["databricks-claude-opus-4-7"].reasoning_options = [{"type":"budget_tokens","min":1024}]
Snapshot: scripts/catalog-sample-fixture.json key "databricks-claude-opus-4-7"
Non-divergences (confirmed consistent)
The following models were checked against models.dev or provider docs and found consistent with the manifest family rules. No exact records needed.
| Model family | Source | Checked against | Status |
|---|---|---|---|
claude-opus-4-7 |
https://platform.claude.com/docs/en/build-with-claude/extended-thinking | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
claude-opus-4-8 |
https://platform.claude.com/docs/en/build-with-claude/extended-thinking | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
claude-sonnet-5.* |
https://platform.claude.com/docs/en/build-with-claude/extended-thinking | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
claude-fable-5 |
https://platform.claude.com/docs/en/build-with-claude/extended-thinking | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
claude-mythos-5 |
https://platform.claude.com/docs/en/build-with-claude/extended-thinking | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
claude-opus-4-6 |
https://platform.claude.com/docs/en/build-with-claude/extended-thinking | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
claude-sonnet-4-6 |
https://platform.claude.com/docs/en/build-with-claude/extended-thinking | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
claude-mythos-preview |
https://platform.claude.com/docs/en/build-with-claude/extended-thinking | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
claude-3* |
https://platform.claude.com/docs/en/build-with-claude/extended-thinking | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
gpt-5-pro |
https://platform.openai.com/docs/guides/reasoning | OpenAI reasoning guide (July 2025) | ✓ Consistent |
gpt-5.6 |
https://platform.openai.com/docs/guides/reasoning | OpenAI reasoning guide (July 2025) | ✓ Consistent |
gpt-5.5 |
https://platform.openai.com/docs/guides/reasoning | OpenAI reasoning guide (July 2025) | ✓ Consistent |
gpt-5.4 |
https://platform.openai.com/docs/guides/reasoning | OpenAI reasoning guide (July 2025) | ✓ Consistent |
gpt-5.1 |
https://platform.openai.com/docs/guides/reasoning | OpenAI reasoning guide (July 2025) | ✓ Consistent |
gpt-5 (base) |
https://platform.openai.com/docs/guides/reasoning | OpenAI reasoning guide (July 2025) | ✓ Consistent |