Files
buzz/scripts/MODELS_DEV_RECONCILIATION.md
T
cc00060ea3 feat(agent): Phase 2 — wire Rust and TS consumers to generated model-capabilities module (#3958)
## What

Phase 2 consumer cutover targeting the
`duncan/databricks-model-label-registry` umbrella branch. Wires
`crates/**` and `desktop/**` consumers to the generated capability
module introduced in Phase 1 (#3821), while keeping old and new paths
both live for differential testing. Phase 3 removes the old paths.

## Commits (boundary-separated)

### feat(agent): Phase 2a — wire Rust consumers to generated capability
module (`crates/**`, `scripts/**`)

- `catalog.rs`: `DATABRICKS_V2_KNOWN_MODELS` re-exported from the
generated module — single source of truth.
- `llm.rs`: `databricks_v2_route_for_model` delegates to
`resolve_model_capabilities("databricks_v2", model)`. Old segment-based
classifier preserved as `#[cfg(test)] _old_*` for the differential
harness. New `databricks_v2_route_differential_old_vs_new` test confirms
100% agreement on all 20 route vectors.
- `config.rs`: new `effort_table_fixture_differential_old_vs_new` test
runs `resolve_model_capabilities` over the 36-entry
`effortTable.fixture.json` and asserts old/new agree modulo a doc-cited
allowlist (4 F1 corrections).
- `scripts/run-differential.mjs`: JS differential harness over
effortTable fixture + normative corpus + catalog-sample fixture. 85
checks, 0 unexpected divergences (5 allowlisted: 4 F1 corrections +
goose-opus-5 anthropic route correction).
- `scripts/MODELS_DEV_RECONCILIATION.md`: deferred MINOR from Phase 1 —
8 trailing-double-space line breaks replaced with `<br>`.

### feat(desktop): Phase 2b — cut TS consumers to generated
model-capabilities module (`desktop/**`)

- `buzzAgentConfig.ts`: adds
`getProviderEffortConfigFromManifest(provider, model?)` — thin wrapper
over `resolveModelCapabilities()` from `modelCapabilities.ts`. Maps
`supportedEfforts → validValues` and `defaultEffort → defaultValue`
(null preserved for manual-budget/Inherit). Old
`getProviderEffortConfig()` and all hand-tables stay live for the
differential harness; Phase 3 retires them.
- `formatAgentModelLabel.ts`: registry-label lookup re-pointed from
hand-maintained `databricksModelNames.ts` import to generated
`DATABRICKS_MODEL_NAMES` exported from `modelCapabilities.ts`. Same Map
shape, identical contents, behavior unchanged.

### fix(scripts): add ts-esm-loader and fix allowlist coverage in
run-differential (`scripts/**`)

- `scripts/ts-esm-loader.mjs`: minimal ESM custom loader that resolves
extensionless relative TS imports. Required because Phase 2b's
`buzzAgentConfig.ts` imports `modelCapabilities` without `.ts` extension
— which Node's `--experimental-strip-types` runner cannot resolve
without a hook.
- `scripts/run-differential.mjs`: shebang updated to self-bootstrap with
the loader; fixes the `totalAllowlisted` counter (was declared but never
incremented — always printed `0 allowlisted`). Replaced with per-axis
hit tracking: reports exercised slot count (`N/total`) in summary; fails
with `STALE_ALLOWLIST` if any declared entry fires zero divergences,
preventing stale entries from silently masking future regressions.

## Verification

- `cargo test -p buzz-agent --lib`: 426/426
- Corpus: 45/45 · schema-negative: 24/24 · `--check` byte-clean
- Differential: 85 checks, 0 unexpected divergences, 6/6 allowlist slots
exercised
- Desktop: 3847/3847 · typecheck clean · biome clean
- Mobile: 1019 pass, 1 skipped — same 5 flaky tests in
`mobile/test/features/channels/` that reproduce at the umbrella base;
zero mobile files in this branch range
- `git diff --check`: clean

## What Remains (Phase 3)

Remove old hand-maintained paths: `_old_*` functions in
`llm.rs`/`config.rs`, old `getProviderEffortConfig` tables in
`buzzAgentConfig.ts`, old `databricksModelNames.ts` import in
`formatAgentModelLabel.ts`, old `databricks_model_names.rs` module.

---------

Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
Signed-off-by: npub1g8493u0xfsjrvflg4n08ezd7vec99mnwzlv0qgwpr9d7gvjwhuzqx59rhw <41ea58f1e64c243627e8acde7c89be667052ee6e17d8f021c1195be4324ebf04@buzz.block.builderlab.xyz>
Signed-off-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
Co-authored-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
Co-authored-by: npub1g8493u0xfsjrvflg4n08ezd7vec99mnwzlv0qgwpr9d7gvjwhuzqx59rhw <41ea58f1e64c243627e8acde7c89be667052ee6e17d8f021c1195be4324ebf04@buzz.block.builderlab.xyz>
2026-08-03 14:58:34 -04:00

116 lines
8.3 KiB
Markdown

# models.dev Reasoning Options Reconciliation Table
**Source queried**: https://models.dev/api.json (2026-07-31)<br>
**Payload SHA-256**: `d5a4974cd69f19b0f67713acaa6bb3b16e920defdc07ecbdf6b0a936181bb0e0`<br>
**Policy (plan v4 §Behavior policy)**: models.dev `reasoning_options` become exact overrides.
Each divergence from the current family rule result is reconciled here: either (a) adopted as an
intentional correction or (b) rejected with a curation note.
**Verbatim source snapshot**: `scripts/catalog-sample-fixture.json` — verbatim `id`, `name`, and
nested `reasoning_options` objects captured from the live API without transformation.
Re-verify hash: `curl -s https://models.dev/api.json | sha256sum`
## Divergences
### `databricks-gpt-5-4-mini`
| | Current family rule (gpt5-4) | models.dev | Disposition |
|---|---|---|---|
| `supported_efforts` | `[none, low, medium, high, xhigh]` | `[low, medium, high]` | **ADOPT** |
**Rationale**: The Databricks AI Gateway v2 endpoint for `databricks-gpt-5-4-mini` explicitly
advertises only `[low, medium, high]` in its `reasoning_options`. The family rule's `none` and
`xhigh` are derived from the upstream OpenAI GPT-5.4 spec, which this Databricks endpoint does
not expose. Provider-advertised wins per plan F1 policy.
**Source**: [https://models.dev/api.json](https://models.dev/api.json) — retrieved 2026-07-31; `providers.databricks.models["databricks-gpt-5-4-mini"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]`<br>
**Snapshot**: `scripts/catalog-sample-fixture.json` key `"databricks-gpt-5-4-mini"`<br>
**Test vector**: `resolver-exact-raw-id-hit` in `scripts/normative-corpus.json`
---
### `databricks-gpt-5-4-nano`
| | Current family rule (gpt5-4) | models.dev | Disposition |
|---|---|---|---|
| `supported_efforts` | `[none, low, medium, high, xhigh]` | `[low, medium, high]` | **ADOPT** |
**Rationale**: Same as `databricks-gpt-5-4-mini`. The nano variant exposes the same restricted
effort set. Provider-advertised wins.
**Source**: [https://models.dev/api.json](https://models.dev/api.json) — retrieved 2026-07-31; `providers.databricks.models["databricks-gpt-5-4-nano"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]`<br>
**Snapshot**: `scripts/catalog-sample-fixture.json` key `"databricks-gpt-5-4-nano"`
---
### `databricks-gpt-5-6-sol`
| | Current family rule (gpt5-6) | models.dev | Disposition |
|---|---|---|---|
| `supported_efforts` | `[none, low, medium, high, xhigh, max]` | `[low, medium, high, max]` | **ADOPT** |
**Rationale**: The Databricks AI Gateway v2 endpoint for `databricks-gpt-5-6-sol` advertises only
`[low, medium, high, max]` in its `reasoning_options`. The family rule's `none` and `xhigh` are
derived from the upstream OpenAI GPT-5.6 spec, which this Databricks endpoint does not expose.
Provider-advertised wins per plan F1 policy.
**Source**: [https://models.dev/api.json](https://models.dev/api.json) — retrieved 2026-07-31; `providers.databricks.models["databricks-gpt-5-6-sol"].reasoning_options = [{"type":"effort","values":["low","medium","high","max"]}]`<br>
**Snapshot**: `scripts/catalog-sample-fixture.json` key `"databricks-gpt-5-6-sol"`
---
### `databricks-gpt-5-5`
| | Current family rule (gpt5-5) | models.dev | Disposition |
|---|---|---|---|
| `supported_efforts` | `[none, low, medium, high, xhigh]` | `[low, medium, high]` | **ADOPT** |
**Rationale**: The Databricks AI Gateway v2 endpoint for `databricks-gpt-5-5` advertises only
`[low, medium, high]` in its `reasoning_options`. The family rule's `none` and `xhigh` are
derived from the upstream OpenAI GPT-5.5 spec, which this Databricks endpoint does not expose.
Provider-advertised wins per plan F1 policy.
**Source**: [https://models.dev/api.json](https://models.dev/api.json) — retrieved 2026-07-31; `providers.databricks.models["databricks-gpt-5-5"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]`<br>
**Snapshot**: `scripts/catalog-sample-fixture.json` key `"databricks-gpt-5-5"`
---
### `databricks-claude-opus-4-7`
| | Current family rule (anthropic-adaptive-xhigh-opus-4-7) | models.dev | Disposition |
|---|---|---|---|
| `reasoning_options` type | effort-based | `budget_tokens` | **NO EFFORT DIVERGENCE** |
**Rationale**: models.dev advertises `reasoning_options=[{"type":"budget_tokens","min":1024}]` —
a different capability axis (extended thinking token budget), not an effort-level selector.
There is no effort divergence to reconcile. The effort capabilities for this model come from the
`anthropic-adaptive-xhigh-opus-4-7` family rule (Anthropic extended-thinking support table).
**Source**: [https://models.dev/api.json](https://models.dev/api.json) — retrieved 2026-07-31; `providers.databricks.models["databricks-claude-opus-4-7"].reasoning_options = [{"type":"budget_tokens","min":1024}]`<br>
**Snapshot**: `scripts/catalog-sample-fixture.json` key `"databricks-claude-opus-4-7"`
---
## Non-divergences (confirmed consistent)
The following models were checked against models.dev or provider docs and found consistent with
the manifest family rules. No exact records needed.
| Model family | Source | Checked against | Status |
|---|---|---|---|
| `claude-opus-4-7` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
| `claude-opus-4-8` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
| `claude-sonnet-5.*` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
| `claude-fable-5` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
| `claude-mythos-5` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
| `claude-opus-4-6` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
| `claude-sonnet-4-6` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
| `claude-mythos-preview` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
| `claude-3*` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
| `gpt-5-pro` | [https://platform.openai.com/docs/guides/reasoning](https://platform.openai.com/docs/guides/reasoning) | OpenAI reasoning guide (July 2025) | ✓ Consistent |
| `gpt-5.6` | [https://platform.openai.com/docs/guides/reasoning](https://platform.openai.com/docs/guides/reasoning) | OpenAI reasoning guide (July 2025) | ✓ Consistent |
| `gpt-5.5` | [https://platform.openai.com/docs/guides/reasoning](https://platform.openai.com/docs/guides/reasoning) | OpenAI reasoning guide (July 2025) | ✓ Consistent |
| `gpt-5.4` | [https://platform.openai.com/docs/guides/reasoning](https://platform.openai.com/docs/guides/reasoning) | OpenAI reasoning guide (July 2025) | ✓ Consistent |
| `gpt-5.1` | [https://platform.openai.com/docs/guides/reasoning](https://platform.openai.com/docs/guides/reasoning) | OpenAI reasoning guide (July 2025) | ✓ Consistent |
| `gpt-5` (base) | [https://platform.openai.com/docs/guides/reasoning](https://platform.openai.com/docs/guides/reasoning) | OpenAI reasoning guide (July 2025) | ✓ Consistent |