Files
buzz/scripts/MODELS_DEV_RECONCILIATION.md
T
cc00060ea3 feat(agent): Phase 2 — wire Rust and TS consumers to generated model-capabilities module (#3958)
## What

Phase 2 consumer cutover targeting the
`duncan/databricks-model-label-registry` umbrella branch. Wires
`crates/**` and `desktop/**` consumers to the generated capability
module introduced in Phase 1 (#3821), while keeping old and new paths
both live for differential testing. Phase 3 removes the old paths.

## Commits (boundary-separated)

### feat(agent): Phase 2a — wire Rust consumers to generated capability
module (`crates/**`, `scripts/**`)

- `catalog.rs`: `DATABRICKS_V2_KNOWN_MODELS` re-exported from the
generated module — single source of truth.
- `llm.rs`: `databricks_v2_route_for_model` delegates to
`resolve_model_capabilities("databricks_v2", model)`. Old segment-based
classifier preserved as `#[cfg(test)] _old_*` for the differential
harness. New `databricks_v2_route_differential_old_vs_new` test confirms
100% agreement on all 20 route vectors.
- `config.rs`: new `effort_table_fixture_differential_old_vs_new` test
runs `resolve_model_capabilities` over the 36-entry
`effortTable.fixture.json` and asserts old/new agree modulo a doc-cited
allowlist (4 F1 corrections).
- `scripts/run-differential.mjs`: JS differential harness over
effortTable fixture + normative corpus + catalog-sample fixture. 85
checks, 0 unexpected divergences (5 allowlisted: 4 F1 corrections +
goose-opus-5 anthropic route correction).
- `scripts/MODELS_DEV_RECONCILIATION.md`: deferred MINOR from Phase 1 —
8 trailing-double-space line breaks replaced with `<br>`.

### feat(desktop): Phase 2b — cut TS consumers to generated
model-capabilities module (`desktop/**`)

- `buzzAgentConfig.ts`: adds
`getProviderEffortConfigFromManifest(provider, model?)` — thin wrapper
over `resolveModelCapabilities()` from `modelCapabilities.ts`. Maps
`supportedEfforts → validValues` and `defaultEffort → defaultValue`
(null preserved for manual-budget/Inherit). Old
`getProviderEffortConfig()` and all hand-tables stay live for the
differential harness; Phase 3 retires them.
- `formatAgentModelLabel.ts`: registry-label lookup re-pointed from
hand-maintained `databricksModelNames.ts` import to generated
`DATABRICKS_MODEL_NAMES` exported from `modelCapabilities.ts`. Same Map
shape, identical contents, behavior unchanged.

### fix(scripts): add ts-esm-loader and fix allowlist coverage in
run-differential (`scripts/**`)

- `scripts/ts-esm-loader.mjs`: minimal ESM custom loader that resolves
extensionless relative TS imports. Required because Phase 2b's
`buzzAgentConfig.ts` imports `modelCapabilities` without `.ts` extension
— which Node's `--experimental-strip-types` runner cannot resolve
without a hook.
- `scripts/run-differential.mjs`: shebang updated to self-bootstrap with
the loader; fixes the `totalAllowlisted` counter (was declared but never
incremented — always printed `0 allowlisted`). Replaced with per-axis
hit tracking: reports exercised slot count (`N/total`) in summary; fails
with `STALE_ALLOWLIST` if any declared entry fires zero divergences,
preventing stale entries from silently masking future regressions.

## Verification

- `cargo test -p buzz-agent --lib`: 426/426
- Corpus: 45/45 · schema-negative: 24/24 · `--check` byte-clean
- Differential: 85 checks, 0 unexpected divergences, 6/6 allowlist slots
exercised
- Desktop: 3847/3847 · typecheck clean · biome clean
- Mobile: 1019 pass, 1 skipped — same 5 flaky tests in
`mobile/test/features/channels/` that reproduce at the umbrella base;
zero mobile files in this branch range
- `git diff --check`: clean

## What Remains (Phase 3)

Remove old hand-maintained paths: `_old_*` functions in
`llm.rs`/`config.rs`, old `getProviderEffortConfig` tables in
`buzzAgentConfig.ts`, old `databricksModelNames.ts` import in
`formatAgentModelLabel.ts`, old `databricks_model_names.rs` module.

---------

Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
Signed-off-by: npub1g8493u0xfsjrvflg4n08ezd7vec99mnwzlv0qgwpr9d7gvjwhuzqx59rhw <41ea58f1e64c243627e8acde7c89be667052ee6e17d8f021c1195be4324ebf04@buzz.block.builderlab.xyz>
Signed-off-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
Co-authored-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
Co-authored-by: npub1g8493u0xfsjrvflg4n08ezd7vec99mnwzlv0qgwpr9d7gvjwhuzqx59rhw <41ea58f1e64c243627e8acde7c89be667052ee6e17d8f021c1195be4324ebf04@buzz.block.builderlab.xyz>
2026-08-03 14:58:34 -04:00

8.3 KiB

models.dev Reasoning Options Reconciliation Table

Source queried: https://models.dev/api.json (2026-07-31)
Payload SHA-256: d5a4974cd69f19b0f67713acaa6bb3b16e920defdc07ecbdf6b0a936181bb0e0
Policy (plan v4 §Behavior policy): models.dev reasoning_options become exact overrides. Each divergence from the current family rule result is reconciled here: either (a) adopted as an intentional correction or (b) rejected with a curation note.

Verbatim source snapshot: scripts/catalog-sample-fixture.json — verbatim id, name, and nested reasoning_options objects captured from the live API without transformation. Re-verify hash: curl -s https://models.dev/api.json | sha256sum

Divergences

databricks-gpt-5-4-mini

Current family rule (gpt5-4) models.dev Disposition
supported_efforts [none, low, medium, high, xhigh] [low, medium, high] ADOPT

Rationale: The Databricks AI Gateway v2 endpoint for databricks-gpt-5-4-mini explicitly advertises only [low, medium, high] in its reasoning_options. The family rule's none and xhigh are derived from the upstream OpenAI GPT-5.4 spec, which this Databricks endpoint does not expose. Provider-advertised wins per plan F1 policy.

Source: https://models.dev/api.json — retrieved 2026-07-31; providers.databricks.models["databricks-gpt-5-4-mini"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]
Snapshot: scripts/catalog-sample-fixture.json key "databricks-gpt-5-4-mini"
Test vector: resolver-exact-raw-id-hit in scripts/normative-corpus.json


databricks-gpt-5-4-nano

Current family rule (gpt5-4) models.dev Disposition
supported_efforts [none, low, medium, high, xhigh] [low, medium, high] ADOPT

Rationale: Same as databricks-gpt-5-4-mini. The nano variant exposes the same restricted effort set. Provider-advertised wins.

Source: https://models.dev/api.json — retrieved 2026-07-31; providers.databricks.models["databricks-gpt-5-4-nano"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]
Snapshot: scripts/catalog-sample-fixture.json key "databricks-gpt-5-4-nano"


databricks-gpt-5-6-sol

Current family rule (gpt5-6) models.dev Disposition
supported_efforts [none, low, medium, high, xhigh, max] [low, medium, high, max] ADOPT

Rationale: The Databricks AI Gateway v2 endpoint for databricks-gpt-5-6-sol advertises only [low, medium, high, max] in its reasoning_options. The family rule's none and xhigh are derived from the upstream OpenAI GPT-5.6 spec, which this Databricks endpoint does not expose. Provider-advertised wins per plan F1 policy.

Source: https://models.dev/api.json — retrieved 2026-07-31; providers.databricks.models["databricks-gpt-5-6-sol"].reasoning_options = [{"type":"effort","values":["low","medium","high","max"]}]
Snapshot: scripts/catalog-sample-fixture.json key "databricks-gpt-5-6-sol"


databricks-gpt-5-5

Current family rule (gpt5-5) models.dev Disposition
supported_efforts [none, low, medium, high, xhigh] [low, medium, high] ADOPT

Rationale: The Databricks AI Gateway v2 endpoint for databricks-gpt-5-5 advertises only [low, medium, high] in its reasoning_options. The family rule's none and xhigh are derived from the upstream OpenAI GPT-5.5 spec, which this Databricks endpoint does not expose. Provider-advertised wins per plan F1 policy.

Source: https://models.dev/api.json — retrieved 2026-07-31; providers.databricks.models["databricks-gpt-5-5"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]
Snapshot: scripts/catalog-sample-fixture.json key "databricks-gpt-5-5"


databricks-claude-opus-4-7

Current family rule (anthropic-adaptive-xhigh-opus-4-7) models.dev Disposition
reasoning_options type effort-based budget_tokens NO EFFORT DIVERGENCE

Rationale: models.dev advertises reasoning_options=[{"type":"budget_tokens","min":1024}] — a different capability axis (extended thinking token budget), not an effort-level selector. There is no effort divergence to reconcile. The effort capabilities for this model come from the anthropic-adaptive-xhigh-opus-4-7 family rule (Anthropic extended-thinking support table).

Source: https://models.dev/api.json — retrieved 2026-07-31; providers.databricks.models["databricks-claude-opus-4-7"].reasoning_options = [{"type":"budget_tokens","min":1024}]
Snapshot: scripts/catalog-sample-fixture.json key "databricks-claude-opus-4-7"


Non-divergences (confirmed consistent)

The following models were checked against models.dev or provider docs and found consistent with the manifest family rules. No exact records needed.

Model family Source Checked against Status
claude-opus-4-7 https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
claude-opus-4-8 https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
claude-sonnet-5.* https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
claude-fable-5 https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
claude-mythos-5 https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
claude-opus-4-6 https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
claude-sonnet-4-6 https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
claude-mythos-preview https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
claude-3* https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
gpt-5-pro https://platform.openai.com/docs/guides/reasoning OpenAI reasoning guide (July 2025) ✓ Consistent
gpt-5.6 https://platform.openai.com/docs/guides/reasoning OpenAI reasoning guide (July 2025) ✓ Consistent
gpt-5.5 https://platform.openai.com/docs/guides/reasoning OpenAI reasoning guide (July 2025) ✓ Consistent
gpt-5.4 https://platform.openai.com/docs/guides/reasoning OpenAI reasoning guide (July 2025) ✓ Consistent
gpt-5.1 https://platform.openai.com/docs/guides/reasoning OpenAI reasoning guide (July 2025) ✓ Consistent
gpt-5 (base) https://platform.openai.com/docs/guides/reasoning OpenAI reasoning guide (July 2025) ✓ Consistent