feat(agent): Phase 2 — wire Rust and TS consumers to generated model-capabilities module (#3958)

## What

Phase 2 consumer cutover targeting the
`duncan/databricks-model-label-registry` umbrella branch. Wires
`crates/**` and `desktop/**` consumers to the generated capability
module introduced in Phase 1 (#3821), while keeping old and new paths
both live for differential testing. Phase 3 removes the old paths.

## Commits (boundary-separated)

### feat(agent): Phase 2a — wire Rust consumers to generated capability
module (`crates/**`, `scripts/**`)

- `catalog.rs`: `DATABRICKS_V2_KNOWN_MODELS` re-exported from the
generated module — single source of truth.
- `llm.rs`: `databricks_v2_route_for_model` delegates to
`resolve_model_capabilities("databricks_v2", model)`. Old segment-based
classifier preserved as `#[cfg(test)] _old_*` for the differential
harness. New `databricks_v2_route_differential_old_vs_new` test confirms
100% agreement on all 20 route vectors.
- `config.rs`: new `effort_table_fixture_differential_old_vs_new` test
runs `resolve_model_capabilities` over the 36-entry
`effortTable.fixture.json` and asserts old/new agree modulo a doc-cited
allowlist (4 F1 corrections).
- `scripts/run-differential.mjs`: JS differential harness over
effortTable fixture + normative corpus + catalog-sample fixture. 85
checks, 0 unexpected divergences (5 allowlisted: 4 F1 corrections +
goose-opus-5 anthropic route correction).
- `scripts/MODELS_DEV_RECONCILIATION.md`: deferred MINOR from Phase 1 —
8 trailing-double-space line breaks replaced with `<br>`.

### feat(desktop): Phase 2b — cut TS consumers to generated
model-capabilities module (`desktop/**`)

- `buzzAgentConfig.ts`: adds
`getProviderEffortConfigFromManifest(provider, model?)` — thin wrapper
over `resolveModelCapabilities()` from `modelCapabilities.ts`. Maps
`supportedEfforts → validValues` and `defaultEffort → defaultValue`
(null preserved for manual-budget/Inherit). Old
`getProviderEffortConfig()` and all hand-tables stay live for the
differential harness; Phase 3 retires them.
- `formatAgentModelLabel.ts`: registry-label lookup re-pointed from
hand-maintained `databricksModelNames.ts` import to generated
`DATABRICKS_MODEL_NAMES` exported from `modelCapabilities.ts`. Same Map
shape, identical contents, behavior unchanged.

### fix(scripts): add ts-esm-loader and fix allowlist coverage in
run-differential (`scripts/**`)

- `scripts/ts-esm-loader.mjs`: minimal ESM custom loader that resolves
extensionless relative TS imports. Required because Phase 2b's
`buzzAgentConfig.ts` imports `modelCapabilities` without `.ts` extension
— which Node's `--experimental-strip-types` runner cannot resolve
without a hook.
- `scripts/run-differential.mjs`: shebang updated to self-bootstrap with
the loader; fixes the `totalAllowlisted` counter (was declared but never
incremented — always printed `0 allowlisted`). Replaced with per-axis
hit tracking: reports exercised slot count (`N/total`) in summary; fails
with `STALE_ALLOWLIST` if any declared entry fires zero divergences,
preventing stale entries from silently masking future regressions.

## Verification

- `cargo test -p buzz-agent --lib`: 426/426
- Corpus: 45/45 · schema-negative: 24/24 · `--check` byte-clean
- Differential: 85 checks, 0 unexpected divergences, 6/6 allowlist slots
exercised
- Desktop: 3847/3847 · typecheck clean · biome clean
- Mobile: 1019 pass, 1 skipped — same 5 flaky tests in
`mobile/test/features/channels/` that reproduce at the umbrella base;
zero mobile files in this branch range
- `git diff --check`: clean

## What Remains (Phase 3)

Remove old hand-maintained paths: `_old_*` functions in
`llm.rs`/`config.rs`, old `getProviderEffortConfig` tables in
`buzzAgentConfig.ts`, old `databricksModelNames.ts` import in
`formatAgentModelLabel.ts`, old `databricks_model_names.rs` module.

---------

Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
Signed-off-by: npub1g8493u0xfsjrvflg4n08ezd7vec99mnwzlv0qgwpr9d7gvjwhuzqx59rhw <41ea58f1e64c243627e8acde7c89be667052ee6e17d8f021c1195be4324ebf04@buzz.block.builderlab.xyz>
Signed-off-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
Co-authored-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
Co-authored-by: npub1g8493u0xfsjrvflg4n08ezd7vec99mnwzlv0qgwpr9d7gvjwhuzqx59rhw <41ea58f1e64c243627e8acde7c89be667052ee6e17d8f021c1195be4324ebf04@buzz.block.builderlab.xyz>
This commit is contained in:
Will Pfleger
2026-08-03 14:58:34 -04:00
committed by GitHub
co-authored by npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 npub1g8493u0xfsjrvflg4n08ezd7vec99mnwzlv0qgwpr9d7gvjwhuzqx59rhw
parent 847f2c5749
commit cc00060ea3
25 changed files with 2905 additions and 157 deletions
+8 -8
View File
@@ -1,7 +1,7 @@
# models.dev Reasoning Options Reconciliation Table
**Source queried**: https://models.dev/api.json (2026-07-31)
**Payload SHA-256**: `d5a4974cd69f19b0f67713acaa6bb3b16e920defdc07ecbdf6b0a936181bb0e0`
**Source queried**: https://models.dev/api.json (2026-07-31)<br>
**Payload SHA-256**: `d5a4974cd69f19b0f67713acaa6bb3b16e920defdc07ecbdf6b0a936181bb0e0`<br>
**Policy (plan v4 §Behavior policy)**: models.dev `reasoning_options` become exact overrides.
Each divergence from the current family rule result is reconciled here: either (a) adopted as an
intentional correction or (b) rejected with a curation note.
@@ -23,8 +23,8 @@ advertises only `[low, medium, high]` in its `reasoning_options`. The family rul
`xhigh` are derived from the upstream OpenAI GPT-5.4 spec, which this Databricks endpoint does
not expose. Provider-advertised wins per plan F1 policy.
**Source**: [https://models.dev/api.json](https://models.dev/api.json) — retrieved 2026-07-31; `providers.databricks.models["databricks-gpt-5-4-mini"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]`
**Snapshot**: `scripts/catalog-sample-fixture.json` key `"databricks-gpt-5-4-mini"`
**Source**: [https://models.dev/api.json](https://models.dev/api.json) — retrieved 2026-07-31; `providers.databricks.models["databricks-gpt-5-4-mini"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]`<br>
**Snapshot**: `scripts/catalog-sample-fixture.json` key `"databricks-gpt-5-4-mini"`<br>
**Test vector**: `resolver-exact-raw-id-hit` in `scripts/normative-corpus.json`
---
@@ -38,7 +38,7 @@ not expose. Provider-advertised wins per plan F1 policy.
**Rationale**: Same as `databricks-gpt-5-4-mini`. The nano variant exposes the same restricted
effort set. Provider-advertised wins.
**Source**: [https://models.dev/api.json](https://models.dev/api.json) — retrieved 2026-07-31; `providers.databricks.models["databricks-gpt-5-4-nano"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]`
**Source**: [https://models.dev/api.json](https://models.dev/api.json) — retrieved 2026-07-31; `providers.databricks.models["databricks-gpt-5-4-nano"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]`<br>
**Snapshot**: `scripts/catalog-sample-fixture.json` key `"databricks-gpt-5-4-nano"`
---
@@ -54,7 +54,7 @@ effort set. Provider-advertised wins.
derived from the upstream OpenAI GPT-5.6 spec, which this Databricks endpoint does not expose.
Provider-advertised wins per plan F1 policy.
**Source**: [https://models.dev/api.json](https://models.dev/api.json) — retrieved 2026-07-31; `providers.databricks.models["databricks-gpt-5-6-sol"].reasoning_options = [{"type":"effort","values":["low","medium","high","max"]}]`
**Source**: [https://models.dev/api.json](https://models.dev/api.json) — retrieved 2026-07-31; `providers.databricks.models["databricks-gpt-5-6-sol"].reasoning_options = [{"type":"effort","values":["low","medium","high","max"]}]`<br>
**Snapshot**: `scripts/catalog-sample-fixture.json` key `"databricks-gpt-5-6-sol"`
---
@@ -70,7 +70,7 @@ Provider-advertised wins per plan F1 policy.
derived from the upstream OpenAI GPT-5.5 spec, which this Databricks endpoint does not expose.
Provider-advertised wins per plan F1 policy.
**Source**: [https://models.dev/api.json](https://models.dev/api.json) — retrieved 2026-07-31; `providers.databricks.models["databricks-gpt-5-5"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]`
**Source**: [https://models.dev/api.json](https://models.dev/api.json) — retrieved 2026-07-31; `providers.databricks.models["databricks-gpt-5-5"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]`<br>
**Snapshot**: `scripts/catalog-sample-fixture.json` key `"databricks-gpt-5-5"`
---
@@ -86,7 +86,7 @@ a different capability axis (extended thinking token budget), not an effort-leve
There is no effort divergence to reconcile. The effort capabilities for this model come from the
`anthropic-adaptive-xhigh-opus-4-7` family rule (Anthropic extended-thinking support table).
**Source**: [https://models.dev/api.json](https://models.dev/api.json) — retrieved 2026-07-31; `providers.databricks.models["databricks-claude-opus-4-7"].reasoning_options = [{"type":"budget_tokens","min":1024}]`
**Source**: [https://models.dev/api.json](https://models.dev/api.json) — retrieved 2026-07-31; `providers.databricks.models["databricks-claude-opus-4-7"].reasoning_options = [{"type":"budget_tokens","min":1024}]`<br>
**Snapshot**: `scripts/catalog-sample-fixture.json` key `"databricks-claude-opus-4-7"`
---
@@ -601,6 +601,50 @@
}
}
},
{
"note": "family rule openai-gpt5-pro / provider databricks",
"provider": "databricks",
"model": "gpt-5-pro",
"resolved": {
"registry_label": "GPT-5 Pro",
"thinking_mode": "none",
"supported_efforts": [
"high"
],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable",
"normalization_policy": "openai-standard",
"_provenance": {
"source": "family",
"rule_id": "openai-gpt5-pro",
"rule_priority": 20,
"normalized_alias": "gpt-5-pro",
"raw_model_id": "gpt-5-pro"
}
}
},
{
"note": "family rule openai-gpt5-pro alias gpt5-pro / provider databricks",
"provider": "databricks",
"model": "gpt5-pro",
"resolved": {
"registry_label": "GPT-5 Pro",
"thinking_mode": "none",
"supported_efforts": [
"high"
],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable",
"normalization_policy": "openai-standard",
"_provenance": {
"source": "family",
"rule_id": "openai-gpt5-pro",
"rule_priority": 20,
"normalized_alias": "gpt5-pro",
"raw_model_id": "gpt5-pro"
}
}
},
{
"note": "family rule openai-gpt5-pro / provider databricks_v2",
"provider": "databricks_v2",
@@ -753,6 +797,114 @@
}
}
},
{
"note": "family rule openai-gpt5-6 / provider databricks",
"provider": "databricks",
"model": "gpt-5.6",
"resolved": {
"registry_label": "GPT-5.6",
"thinking_mode": "none",
"supported_efforts": [
"none",
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable",
"normalization_policy": "openai-standard",
"_provenance": {
"source": "family",
"rule_id": "openai-gpt5-6",
"rule_priority": 15,
"normalized_alias": "gpt-5.6",
"raw_model_id": "gpt-5.6"
}
}
},
{
"note": "family rule openai-gpt5-6 alias gpt5.6 / provider databricks",
"provider": "databricks",
"model": "gpt5.6",
"resolved": {
"registry_label": "GPT-5.6",
"thinking_mode": "none",
"supported_efforts": [
"none",
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable",
"normalization_policy": "openai-standard",
"_provenance": {
"source": "family",
"rule_id": "openai-gpt5-6",
"rule_priority": 15,
"normalized_alias": "gpt5.6",
"raw_model_id": "gpt5.6"
}
}
},
{
"note": "family rule openai-gpt5-6 alias gpt-5-6 / provider databricks",
"provider": "databricks",
"model": "gpt-5-6",
"resolved": {
"registry_label": "GPT-5.6",
"thinking_mode": "none",
"supported_efforts": [
"none",
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable",
"normalization_policy": "openai-standard",
"_provenance": {
"source": "family",
"rule_id": "openai-gpt5-6",
"rule_priority": 15,
"normalized_alias": "gpt-5-6",
"raw_model_id": "gpt-5-6"
}
}
},
{
"note": "family rule openai-gpt5-6 alias gpt5-6 / provider databricks",
"provider": "databricks",
"model": "gpt5-6",
"resolved": {
"registry_label": "GPT-5.6",
"thinking_mode": "none",
"supported_efforts": [
"none",
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable",
"normalization_policy": "openai-standard",
"_provenance": {
"source": "family",
"rule_id": "openai-gpt5-6",
"rule_priority": 15,
"normalized_alias": "gpt5-6",
"raw_model_id": "gpt5-6"
}
}
},
{
"note": "family rule openai-gpt5-6 / provider databricks_v2",
"provider": "databricks_v2",
@@ -965,6 +1117,110 @@
}
}
},
{
"note": "family rule openai-gpt5-5 / provider databricks",
"provider": "databricks",
"model": "gpt-5.5",
"resolved": {
"registry_label": "GPT-5.5",
"thinking_mode": "none",
"supported_efforts": [
"none",
"low",
"medium",
"high",
"xhigh"
],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable",
"normalization_policy": "openai-standard",
"_provenance": {
"source": "family",
"rule_id": "openai-gpt5-5",
"rule_priority": 15,
"normalized_alias": "gpt-5.5",
"raw_model_id": "gpt-5.5"
}
}
},
{
"note": "family rule openai-gpt5-5 alias gpt5.5 / provider databricks",
"provider": "databricks",
"model": "gpt5.5",
"resolved": {
"registry_label": "GPT-5.5",
"thinking_mode": "none",
"supported_efforts": [
"none",
"low",
"medium",
"high",
"xhigh"
],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable",
"normalization_policy": "openai-standard",
"_provenance": {
"source": "family",
"rule_id": "openai-gpt5-5",
"rule_priority": 15,
"normalized_alias": "gpt5.5",
"raw_model_id": "gpt5.5"
}
}
},
{
"note": "family rule openai-gpt5-5 alias gpt-5-5 / provider databricks",
"provider": "databricks",
"model": "gpt-5-5",
"resolved": {
"registry_label": "GPT-5.5",
"thinking_mode": "none",
"supported_efforts": [
"none",
"low",
"medium",
"high",
"xhigh"
],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable",
"normalization_policy": "openai-standard",
"_provenance": {
"source": "family",
"rule_id": "openai-gpt5-5",
"rule_priority": 15,
"normalized_alias": "gpt-5-5",
"raw_model_id": "gpt-5-5"
}
}
},
{
"note": "family rule openai-gpt5-5 alias gpt5-5 / provider databricks",
"provider": "databricks",
"model": "gpt5-5",
"resolved": {
"registry_label": "GPT-5.5",
"thinking_mode": "none",
"supported_efforts": [
"none",
"low",
"medium",
"high",
"xhigh"
],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable",
"normalization_policy": "openai-standard",
"_provenance": {
"source": "family",
"rule_id": "openai-gpt5-5",
"rule_priority": 15,
"normalized_alias": "gpt5-5",
"raw_model_id": "gpt5-5"
}
}
},
{
"note": "family rule openai-gpt5-5 / provider databricks_v2",
"provider": "databricks_v2",
@@ -1173,6 +1429,110 @@
}
}
},
{
"note": "family rule openai-gpt5-4 / provider databricks",
"provider": "databricks",
"model": "gpt-5.4",
"resolved": {
"registry_label": "GPT-5.4",
"thinking_mode": "none",
"supported_efforts": [
"none",
"low",
"medium",
"high",
"xhigh"
],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable",
"normalization_policy": "openai-standard",
"_provenance": {
"source": "family",
"rule_id": "openai-gpt5-4",
"rule_priority": 15,
"normalized_alias": "gpt-5.4",
"raw_model_id": "gpt-5.4"
}
}
},
{
"note": "family rule openai-gpt5-4 alias gpt5.4 / provider databricks",
"provider": "databricks",
"model": "gpt5.4",
"resolved": {
"registry_label": "GPT-5.4",
"thinking_mode": "none",
"supported_efforts": [
"none",
"low",
"medium",
"high",
"xhigh"
],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable",
"normalization_policy": "openai-standard",
"_provenance": {
"source": "family",
"rule_id": "openai-gpt5-4",
"rule_priority": 15,
"normalized_alias": "gpt5.4",
"raw_model_id": "gpt5.4"
}
}
},
{
"note": "family rule openai-gpt5-4 alias gpt-5-4 / provider databricks",
"provider": "databricks",
"model": "gpt-5-4",
"resolved": {
"registry_label": "GPT-5.4",
"thinking_mode": "none",
"supported_efforts": [
"none",
"low",
"medium",
"high",
"xhigh"
],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable",
"normalization_policy": "openai-standard",
"_provenance": {
"source": "family",
"rule_id": "openai-gpt5-4",
"rule_priority": 15,
"normalized_alias": "gpt-5-4",
"raw_model_id": "gpt-5-4"
}
}
},
{
"note": "family rule openai-gpt5-4 alias gpt5-4 / provider databricks",
"provider": "databricks",
"model": "gpt5-4",
"resolved": {
"registry_label": "GPT-5.4",
"thinking_mode": "none",
"supported_efforts": [
"none",
"low",
"medium",
"high",
"xhigh"
],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable",
"normalization_policy": "openai-standard",
"_provenance": {
"source": "family",
"rule_id": "openai-gpt5-4",
"rule_priority": 15,
"normalized_alias": "gpt5-4",
"raw_model_id": "gpt5-4"
}
}
},
{
"note": "family rule openai-gpt5-4 / provider databricks_v2",
"provider": "databricks_v2",
@@ -1377,6 +1737,106 @@
}
}
},
{
"note": "family rule openai-gpt5-1 / provider databricks",
"provider": "databricks",
"model": "gpt-5.1",
"resolved": {
"registry_label": "GPT-5.1",
"thinking_mode": "none",
"supported_efforts": [
"none",
"low",
"medium",
"high"
],
"default_effort": "none",
"databricks_v2_wire_route": "not-applicable",
"normalization_policy": "openai-standard",
"_provenance": {
"source": "family",
"rule_id": "openai-gpt5-1",
"rule_priority": 15,
"normalized_alias": "gpt-5.1",
"raw_model_id": "gpt-5.1"
}
}
},
{
"note": "family rule openai-gpt5-1 alias gpt5.1 / provider databricks",
"provider": "databricks",
"model": "gpt5.1",
"resolved": {
"registry_label": "GPT-5.1",
"thinking_mode": "none",
"supported_efforts": [
"none",
"low",
"medium",
"high"
],
"default_effort": "none",
"databricks_v2_wire_route": "not-applicable",
"normalization_policy": "openai-standard",
"_provenance": {
"source": "family",
"rule_id": "openai-gpt5-1",
"rule_priority": 15,
"normalized_alias": "gpt5.1",
"raw_model_id": "gpt5.1"
}
}
},
{
"note": "family rule openai-gpt5-1 alias gpt-5-1 / provider databricks",
"provider": "databricks",
"model": "gpt-5-1",
"resolved": {
"registry_label": "GPT-5.1",
"thinking_mode": "none",
"supported_efforts": [
"none",
"low",
"medium",
"high"
],
"default_effort": "none",
"databricks_v2_wire_route": "not-applicable",
"normalization_policy": "openai-standard",
"_provenance": {
"source": "family",
"rule_id": "openai-gpt5-1",
"rule_priority": 15,
"normalized_alias": "gpt-5-1",
"raw_model_id": "gpt-5-1"
}
}
},
{
"note": "family rule openai-gpt5-1 alias gpt5-1 / provider databricks",
"provider": "databricks",
"model": "gpt5-1",
"resolved": {
"registry_label": "GPT-5.1",
"thinking_mode": "none",
"supported_efforts": [
"none",
"low",
"medium",
"high"
],
"default_effort": "none",
"databricks_v2_wire_route": "not-applicable",
"normalization_policy": "openai-standard",
"_provenance": {
"source": "family",
"rule_id": "openai-gpt5-1",
"rule_priority": 15,
"normalized_alias": "gpt5-1",
"raw_model_id": "gpt5-1"
}
}
},
{
"note": "family rule openai-gpt5-1 / provider databricks_v2",
"provider": "databricks_v2",
@@ -1527,6 +1987,56 @@
}
}
},
{
"note": "family rule openai-gpt5-base / provider databricks",
"provider": "databricks",
"model": "gpt-5",
"resolved": {
"registry_label": "GPT-5",
"thinking_mode": "none",
"supported_efforts": [
"minimal",
"low",
"medium",
"high"
],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable",
"normalization_policy": "openai-standard",
"_provenance": {
"source": "family",
"rule_id": "openai-gpt5-base",
"rule_priority": 10,
"normalized_alias": "gpt-5",
"raw_model_id": "gpt-5"
}
}
},
{
"note": "family rule openai-gpt5-base alias gpt5 / provider databricks",
"provider": "databricks",
"model": "gpt5",
"resolved": {
"registry_label": "GPT-5",
"thinking_mode": "none",
"supported_efforts": [
"minimal",
"low",
"medium",
"high"
],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable",
"normalization_policy": "openai-standard",
"_provenance": {
"source": "family",
"rule_id": "openai-gpt5-base",
"rule_priority": 10,
"normalized_alias": "gpt5",
"raw_model_id": "gpt5"
}
}
},
{
"note": "family rule openai-gpt5-base / provider databricks_v2",
"provider": "databricks_v2",
+6
View File
@@ -257,6 +257,7 @@
],
"providers": [
"openai",
"databricks",
"databricks_v2"
],
"match_priority": 20,
@@ -280,6 +281,7 @@
],
"providers": [
"openai",
"databricks",
"databricks_v2"
],
"match_priority": 15,
@@ -308,6 +310,7 @@
],
"providers": [
"openai",
"databricks",
"databricks_v2"
],
"match_priority": 15,
@@ -335,6 +338,7 @@
],
"providers": [
"openai",
"databricks",
"databricks_v2"
],
"match_priority": 15,
@@ -362,6 +366,7 @@
],
"providers": [
"openai",
"databricks",
"databricks_v2"
],
"match_priority": 15,
@@ -386,6 +391,7 @@
],
"providers": [
"openai",
"databricks",
"databricks_v2"
],
"match_priority": 10,
+360 -50
View File
@@ -9,7 +9,11 @@
"raw_model_id": "claude-3-7-sonnet-20250219",
"expect": {
"thinking_mode": "manual-budget",
"supported_efforts": ["low", "medium", "high"],
"supported_efforts": [
"low",
"medium",
"high"
],
"default_effort": null,
"databricks_v2_wire_route": "not-applicable"
}
@@ -20,7 +24,11 @@
"raw_model_id": "claude-opus-4-5",
"expect": {
"thinking_mode": "manual-budget",
"supported_efforts": ["low", "medium", "high"],
"supported_efforts": [
"low",
"medium",
"high"
],
"default_effort": null,
"databricks_v2_wire_route": "not-applicable"
}
@@ -31,7 +39,13 @@
"raw_model_id": "claude-opus-4-7",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
"supported_efforts": [
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
@@ -42,7 +56,13 @@
"raw_model_id": "claude-opus-4-8",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
"supported_efforts": [
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
@@ -53,7 +73,13 @@
"raw_model_id": "claude-sonnet-5-20260101",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
"supported_efforts": [
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
@@ -64,7 +90,13 @@
"raw_model_id": "claude-fable-5",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
"supported_efforts": [
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
@@ -75,7 +107,13 @@
"raw_model_id": "claude-mythos-5",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
"supported_efforts": [
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
@@ -86,7 +124,12 @@
"raw_model_id": "claude-opus-4-6",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "max"],
"supported_efforts": [
"low",
"medium",
"high",
"max"
],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
@@ -97,7 +140,12 @@
"raw_model_id": "claude-sonnet-4-6",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "max"],
"supported_efforts": [
"low",
"medium",
"high",
"max"
],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
@@ -108,7 +156,12 @@
"raw_model_id": "claude-mythos-preview",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "max"],
"supported_efforts": [
"low",
"medium",
"high",
"max"
],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
@@ -122,7 +175,13 @@
"raw_model_id": "",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
"supported_efforts": [
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
@@ -133,7 +192,13 @@
"raw_model_id": "claude-ultra-9000",
"expect": {
"thinking_mode": "omit-fields",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
"supported_efforts": [
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
@@ -147,7 +212,9 @@
"raw_model_id": "gpt-5-pro",
"expect": {
"thinking_mode": "none",
"supported_efforts": ["high"],
"supported_efforts": [
"high"
],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
@@ -158,7 +225,14 @@
"raw_model_id": "gpt-5.6",
"expect": {
"thinking_mode": "none",
"supported_efforts": ["none", "low", "medium", "high", "xhigh", "max"],
"supported_efforts": [
"none",
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable"
}
@@ -169,7 +243,14 @@
"raw_model_id": "gpt-5-6",
"expect": {
"thinking_mode": "none",
"supported_efforts": ["none", "low", "medium", "high", "xhigh", "max"],
"supported_efforts": [
"none",
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable"
}
@@ -180,7 +261,13 @@
"raw_model_id": "gpt-5.5",
"expect": {
"thinking_mode": "none",
"supported_efforts": ["none", "low", "medium", "high", "xhigh"],
"supported_efforts": [
"none",
"low",
"medium",
"high",
"xhigh"
],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable"
}
@@ -191,7 +278,13 @@
"raw_model_id": "gpt-5.4",
"expect": {
"thinking_mode": "none",
"supported_efforts": ["none", "low", "medium", "high", "xhigh"],
"supported_efforts": [
"none",
"low",
"medium",
"high",
"xhigh"
],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable"
}
@@ -202,7 +295,12 @@
"raw_model_id": "gpt-5.1",
"expect": {
"thinking_mode": "none",
"supported_efforts": ["none", "low", "medium", "high"],
"supported_efforts": [
"none",
"low",
"medium",
"high"
],
"default_effort": "none",
"databricks_v2_wire_route": "not-applicable"
}
@@ -213,13 +311,18 @@
"raw_model_id": "gpt-5",
"expect": {
"thinking_mode": "none",
"supported_efforts": ["minimal", "low", "medium", "high"],
"supported_efforts": [
"minimal",
"low",
"medium",
"high"
],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"_group": "OpenAI adversarial — gpt5 boundary-aware matching (ported from config.rs tests)"
"_group": "OpenAI adversarial \u2014 gpt5 boundary-aware matching (ported from config.rs tests)"
},
{
"id": "openai-gpt5-1106-should-not-match-base",
@@ -227,7 +330,12 @@
"raw_model_id": "gpt-5-1106",
"_note": "gpt-5-1106: '-1106' is a 4-digit date segment, NOT a short version (gpt5-base rejects only 1-3 digit suffixes). Must match base table [minimal,low,medium,high], NOT fall through to unknown.",
"expect": {
"supported_efforts": ["minimal", "low", "medium", "high"]
"supported_efforts": [
"minimal",
"low",
"medium",
"high"
]
}
},
{
@@ -236,7 +344,12 @@
"raw_model_id": "gpt-5-4o",
"_note": "gpt-5-4o: '4o' after '-' is NOT a short numeric suffix (it contains a letter). Must match gpt5-base. Crucially, must NOT match gpt-5.4 (the '4' is followed by 'o', not boundary char).",
"expect": {
"supported_efforts": ["minimal", "low", "medium", "high"]
"supported_efforts": [
"minimal",
"low",
"medium",
"high"
]
}
},
{
@@ -245,7 +358,9 @@
"raw_model_id": "gpt-5-pro",
"_note": "gpt-5-pro should hit gpt5-pro rule (priority 20), NOT gpt-5 base.",
"expect": {
"supported_efforts": ["high"],
"supported_efforts": [
"high"
],
"default_effort": "high"
}
},
@@ -253,22 +368,34 @@
"id": "openai-multi-digit-version-gpt5-10",
"provider": "openai",
"raw_model_id": "gpt-5-10",
"_note": "gpt-5-10 — two-digit suffix prevents gpt5-base match. Falls through to unknown.",
"_note": "gpt-5-10 \u2014 two-digit suffix prevents gpt5-base match. Falls through to unknown.",
"expect": {
"supported_efforts": ["none", "minimal", "low", "medium", "high", "xhigh"]
"supported_efforts": [
"none",
"minimal",
"low",
"medium",
"high",
"xhigh"
]
}
},
{
"id": "openai-gpt5-date-suffix",
"provider": "openai",
"raw_model_id": "gpt-5-20260101",
"_note": "gpt-5-20260101 — long numeric suffix after base: '20260101' is 8 digits, beyond 1-3 digit reject, should hit gpt5-base.",
"_note": "gpt-5-20260101 \u2014 long numeric suffix after base: '20260101' is 8 digits, beyond 1-3 digit reject, should hit gpt5-base.",
"expect": {
"supported_efforts": ["minimal", "low", "medium", "high"]
"supported_efforts": [
"minimal",
"low",
"medium",
"high"
]
}
},
{
"_group": "DatabricksV2 — segment-based routing (ported from llm.rs tests)"
"_group": "DatabricksV2 \u2014 segment-based routing (ported from llm.rs tests)"
},
{
"id": "dbv2-gpt5-route-openai-responses",
@@ -276,7 +403,13 @@
"raw_model_id": "gpt-5.5",
"expect": {
"databricks_v2_wire_route": "openai-responses",
"supported_efforts": ["none", "low", "medium", "high", "xhigh"]
"supported_efforts": [
"none",
"low",
"medium",
"high",
"xhigh"
]
}
},
{
@@ -286,14 +419,20 @@
"expect": {
"databricks_v2_wire_route": "anthropic-messages",
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"]
"supported_efforts": [
"low",
"medium",
"high",
"xhigh",
"max"
]
}
},
{
"id": "dbv2-claude-prefix-stripped",
"provider": "databricks_v2",
"raw_model_id": "databricks-claude-opus-4-7",
"_note": "databricks- prefix stripped → claude-opus-4-7 → Anthropic route",
"_note": "databricks- prefix stripped \u2192 claude-opus-4-7 \u2192 Anthropic route",
"expect": {
"databricks_v2_wire_route": "anthropic-messages",
"thinking_mode": "adaptive"
@@ -303,29 +442,41 @@
"id": "dbv2-goose-claude-prefix-stripped",
"provider": "databricks_v2",
"raw_model_id": "goose-claude-fable-5",
"_note": "goose- prefix stripped → claude-fable-5 → Anthropic adaptive+xhigh",
"_note": "goose- prefix stripped \u2192 claude-fable-5 \u2192 Anthropic adaptive+xhigh",
"expect": {
"databricks_v2_wire_route": "anthropic-messages",
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"]
"supported_efforts": [
"low",
"medium",
"high",
"xhigh",
"max"
]
}
},
{
"id": "dbv2-team-prefix-stripped",
"provider": "databricks_v2",
"raw_model_id": "team-x-claude-opus-4-7",
"_note": "team-x- prefix stripped → claude-opus-4-7 → Anthropic route",
"_note": "team-x- prefix stripped \u2192 claude-opus-4-7 \u2192 Anthropic route",
"expect": {
"databricks_v2_wire_route": "anthropic-messages",
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"]
"supported_efforts": [
"low",
"medium",
"high",
"xhigh",
"max"
]
}
},
{
"id": "dbv2-consolidated-llama-not-sol",
"provider": "databricks_v2",
"raw_model_id": "consolidated-llama",
"_note": "segment test: 'sol' is a SUBSTRING of 'consolidated' — must NOT match DATABRICKS_V2_OPENAI_CODE_NAMES 'sol'. Falls through to mlflow-chat.",
"_note": "segment test: 'sol' is a SUBSTRING of 'consolidated' \u2014 must NOT match DATABRICKS_V2_OPENAI_CODE_NAMES 'sol'. Falls through to mlflow-chat.",
"expect": {
"databricks_v2_wire_route": "mlflow-chat"
}
@@ -334,7 +485,7 @@
"id": "dbv2-terraform-coder-not-terra",
"provider": "databricks_v2",
"raw_model_id": "terraform-coder",
"_note": "segment test: 'terra' is a prefix of 'terraform' — must NOT match 'terra' code name. Falls through to mlflow-chat.",
"_note": "segment test: 'terra' is a prefix of 'terraform' \u2014 must NOT match 'terra' code name. Falls through to mlflow-chat.",
"expect": {
"databricks_v2_wire_route": "mlflow-chat"
}
@@ -367,7 +518,7 @@
}
},
{
"_group": "P2-A resolver-contract vectors (plan v4 §Resolver contract)"
"_group": "P2-A resolver-contract vectors (plan v4 \u00a7Resolver contract)"
},
{
"id": "resolver-exact-raw-id-hit",
@@ -375,16 +526,26 @@
"raw_model_id": "databricks-gpt-5-4-mini",
"_note": "Exact record exists. Must return exact Databricks override: low|medium|high (not family's none+xhigh).",
"expect": {
"supported_efforts": ["low", "medium", "high"]
"supported_efforts": [
"low",
"medium",
"high"
]
}
},
{
"id": "resolver-prefixed-alias-misses-exact",
"provider": "databricks_v2",
"raw_model_id": "team-x-databricks-gpt-5-4-mini",
"_note": "Prefixed alias of an exact ID. Raw exact lookup MUST miss (key is team-x-..., not databricks-...). Falls to family rules (gpt5-4 family → none+xhigh).",
"_note": "Prefixed alias of an exact ID. Raw exact lookup MUST miss (key is team-x-..., not databricks-...). Falls to family rules (gpt5-4 family \u2192 none+xhigh).",
"expect": {
"supported_efforts": ["none", "low", "medium", "high", "xhigh"]
"supported_efforts": [
"none",
"low",
"medium",
"high",
"xhigh"
]
}
},
{
@@ -400,9 +561,14 @@
"id": "resolver-exact-efforts-plus-family-route",
"provider": "databricks_v2",
"raw_model_id": "databricks-gpt-5-6-sol",
"_note": "Exact record with efforts from models.dev (low|medium|high|max — provider-advertised, no none/xhigh). Route materialized from gpt5-6 family rule (openai-responses). Must return both, complete.",
"_note": "Exact record with efforts from models.dev (low|medium|high|max \u2014 provider-advertised, no none/xhigh). Route materialized from gpt5-6 family rule (openai-responses). Must return both, complete.",
"expect": {
"supported_efforts": ["low", "medium", "high", "max"],
"supported_efforts": [
"low",
"medium",
"high",
"max"
],
"databricks_v2_wire_route": "openai-responses"
}
},
@@ -410,9 +576,13 @@
"id": "dbv2-gpt5-5-exact-override",
"provider": "databricks_v2",
"raw_model_id": "databricks-gpt-5-5",
"_note": "Exact record adopts models.dev advertised set [low,medium,high]. Family rule (gpt5-5) has none+xhigh — provider-advertised wins per plan F1.",
"_note": "Exact record adopts models.dev advertised set [low,medium,high]. Family rule (gpt5-5) has none+xhigh \u2014 provider-advertised wins per plan F1.",
"expect": {
"supported_efforts": ["low", "medium", "high"],
"supported_efforts": [
"low",
"medium",
"high"
],
"databricks_v2_wire_route": "openai-responses"
}
},
@@ -426,7 +596,15 @@
"_note": "DBv2 blank: route-unknown, all 7 efforts, default medium.",
"expect": {
"databricks_v2_wire_route": "route-unknown",
"supported_efforts": ["none", "minimal", "low", "medium", "high", "xhigh", "max"],
"supported_efforts": [
"none",
"minimal",
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "medium"
}
},
@@ -437,7 +615,14 @@
"_note": "DBv2 concrete-unknown: mlflow-chat, all-except-max (6 efforts).",
"expect": {
"databricks_v2_wire_route": "mlflow-chat",
"supported_efforts": ["none", "minimal", "low", "medium", "high", "xhigh"]
"supported_efforts": [
"none",
"minimal",
"low",
"medium",
"high",
"xhigh"
]
}
},
{
@@ -447,7 +632,14 @@
"_note": "OpenAI blank: not-applicable route, all-except-max, medium default.",
"expect": {
"databricks_v2_wire_route": "not-applicable",
"supported_efforts": ["none", "minimal", "low", "medium", "high", "xhigh"],
"supported_efforts": [
"none",
"minimal",
"low",
"medium",
"high",
"xhigh"
],
"default_effort": "medium"
}
},
@@ -458,7 +650,14 @@
"_note": "OpenAI concrete unknown (unverified family): not-applicable route, all-except-max, medium default.",
"expect": {
"databricks_v2_wire_route": "not-applicable",
"supported_efforts": ["none", "minimal", "low", "medium", "high", "xhigh"],
"supported_efforts": [
"none",
"minimal",
"low",
"medium",
"high",
"xhigh"
],
"default_effort": "medium"
}
},
@@ -469,7 +668,13 @@
"_note": "Anthropic blank: assume adaptive with full support (incl. xhigh).",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
"supported_efforts": [
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "high"
}
},
@@ -481,5 +686,110 @@
"expect": {
"thinking_mode": "omit-fields"
}
},
{
"_group": "Legacy Databricks provider (P3 effort rules)"
},
{
"id": "databricks-gpt5-pro-effort",
"provider": "databricks",
"raw_model_id": "databricks-gpt-5-pro",
"_note": "Legacy databricks GPT-5 Pro: same effort set as openai/gpt-5-pro \u2014 only [high], default high. Wire route not-applicable.",
"expect": {
"databricks_v2_wire_route": "not-applicable",
"supported_efforts": [
"high"
],
"default_effort": "high"
}
},
{
"id": "databricks-gpt5-6-effort",
"provider": "databricks",
"raw_model_id": "databricks-gpt-5.6",
"_note": "Legacy databricks GPT-5.6: same effort set as openai/gpt-5.6 \u2014 [none,low,medium,high,xhigh,max], default medium. Wire route not-applicable.",
"expect": {
"databricks_v2_wire_route": "not-applicable",
"supported_efforts": [
"none",
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "medium"
}
},
{
"id": "databricks-gpt5-1-effort",
"provider": "databricks",
"raw_model_id": "databricks-gpt-5.1",
"_note": "Legacy databricks GPT-5.1: same effort set as openai/gpt-5.1 \u2014 [none,low,medium,high], default none. Wire route not-applicable.",
"expect": {
"databricks_v2_wire_route": "not-applicable",
"supported_efforts": [
"none",
"low",
"medium",
"high"
],
"default_effort": "none"
}
},
{
"_group": "openai-compat alias canonicalization (Thufir P3 corrective action 1)",
"_note": "Rust normalizes openai-compat \u2192 Provider::OpenAi before reaching normalize_effort_for_provider. TS PROVIDER_ALIASES must match so the UI effort table equals the Rust request behavior. Interpreters must canonicalize openai-compat \u2192 openai before resolving; the expected values are identical to the corresponding openai vectors."
},
{
"id": "openai-compat-gpt-5-pro",
"provider": "openai-compat",
"raw_model_id": "gpt-5-pro",
"_note": "openai-compat/gpt-5-pro must resolve identically to openai/gpt-5-pro: [high] only, default high.",
"expect": {
"thinking_mode": "none",
"supported_efforts": [
"high"
],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "openai-compat-gpt-5-5",
"provider": "openai-compat",
"raw_model_id": "gpt-5.5",
"_note": "openai-compat/gpt-5.5 must resolve identically to openai/gpt-5.5: [none,low,medium,high,xhigh], default medium.",
"expect": {
"thinking_mode": "none",
"supported_efforts": [
"none",
"low",
"medium",
"high",
"xhigh"
],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "openai-compat-empty-model",
"provider": "openai-compat",
"raw_model_id": "",
"_note": "openai-compat with blank model: resolves identically to openai unknown \u2014 all-except-max, default medium.",
"expect": {
"thinking_mode": "none",
"supported_efforts": [
"none",
"minimal",
"low",
"medium",
"high",
"xhigh"
],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable"
}
}
]
+14 -1
View File
@@ -30,6 +30,19 @@ const corpus = JSON.parse(
readFileSync(join(repoRoot, "scripts", "normative-corpus.json"), "utf8"),
);
// ----- Provider alias canonicalization -----
// Mirrors production canonicalizeProvider() in desktop/src/features/agents/lib/formatAgentModelLabel.ts.
// Applied before every generated lookup so alias vectors (e.g. "openai-compat") pass both interpreters.
const PROVIDER_ALIASES = {
"databricks-v2": "databricks_v2",
"openai-compat": "openai",
};
function canonicalizeProvider(provider) {
const normalized = (provider ?? "").trim().toLowerCase();
return PROVIDER_ALIASES[normalized] ?? normalized;
}
// ----- Run corpus -----
let passed = 0;
@@ -41,7 +54,7 @@ for (const entry of corpus) {
if (!entry.expect) continue;
// resolveModelCapabilities returns camelCase keys (registryLabel, thinkingMode, etc.)
const result = resolveModelCapabilities(entry.provider, entry.raw_model_id);
const result = resolveModelCapabilities(canonicalizeProvider(entry.provider), entry.raw_model_id);
const expect = entry.expect;
const failures = [];
+239
View File
@@ -0,0 +1,239 @@
#!/usr/bin/env node
/**
* Phase-2 differential harness — compare old buzzAgentConfig.ts effort logic with
* the new generated modelCapabilities.ts interpreter over:
* 1. The 36-entry effortTable.fixture.json (cross-boundary Rust/TS fixture)
* 2. The 45-vector normative corpus (scripts/normative-corpus.json)
* 3. The catalog-sample fixture (scripts/catalog-sample-fixture.json)
*
* Equality is required except for entries in the committed allowlist of intentional
* F1 corrections (models.dev provider-capability reconciliations).
*
* Usage: node --experimental-strip-types scripts/run-differential.mjs [--verbose]
* Exits 0 on all-pass (modulo allowlist), 1 on unexpected divergence or unexercised allowlist entry.
*/
import { readFileSync } from "node:fs";
import { join, dirname } from "node:path";
import { fileURLToPath } from "node:url";
const __dirname = dirname(fileURLToPath(import.meta.url));
const repoRoot = join(__dirname, "..");
const VERBOSE = process.argv.includes("--verbose");
// ---------------------------------------------------------------------------
// Import both interpreters
// ---------------------------------------------------------------------------
// NEW: generated capability module
const { resolveModelCapabilities: resolveNew } = await import(
join(repoRoot, "desktop", "src", "features", "agents", "ui", "modelCapabilities.ts")
);
// OLD: buzzAgentConfig.ts effort config
const { getProviderEffortConfig_oldHandTable: getOldEffortConfig } = await import(
join(repoRoot, "desktop", "src", "features", "agents", "ui", "buzzAgentConfig.ts")
);
// ---------------------------------------------------------------------------
// Intentional corrections allowlist (Phase 1 F1 reconciliations)
// Each entry: { provider, raw_model_id, reason }
// ---------------------------------------------------------------------------
const ALLOWLIST = [
{
provider: "databricks_v2",
raw_model_id: "databricks-gpt-5-5",
axes: ["supported_efforts"],
reason: "Phase 1 ADOPT: models.dev d5a4974c advertises [low,medium,high]; old returns [none,low,medium,high,xhigh]",
},
{
provider: "databricks_v2",
raw_model_id: "databricks-gpt-5-4-mini",
axes: ["supported_efforts"],
reason: "Phase 1 ADOPT: models.dev advertises [low,medium,high]; old returns [none,low,medium,high,xhigh]",
},
{
provider: "databricks_v2",
raw_model_id: "databricks-gpt-5-4-nano",
axes: ["supported_efforts"],
reason: "Phase 1 ADOPT: models.dev advertises [low,medium,high]; old returns [none,low,medium,high,xhigh]",
},
{
provider: "databricks_v2",
raw_model_id: "databricks-gpt-5-6-sol",
axes: ["supported_efforts"],
reason: "Phase 1 ADOPT: models.dev advertises [low,medium,high,max]; old returns [none,low,medium,high,xhigh,max]",
},
{
provider: "databricks_v2",
raw_model_id: "goose-opus-5",
axes: ["supported_efforts", "default_effort"],
reason: "Phase 1 correction: 'opus' is a named DBv2 segment → anthropic-messages route; old config.rs disagreed with llm.rs (corpus note dbv2-goose-opus-5-is-anthropic). Generated adopts anthropic adaptive-xhigh capabilities consistent with the wire route.",
},
];
// Track which allowlist entries are actually exercised (suppressed a divergence).
// Keyed as "provider:raw_model_id:axis".
const allowlistHits = new Set();
function isAllowlisted(provider, rawModelId, axis) {
const entry = ALLOWLIST.find(
(e) =>
e.provider === provider &&
e.raw_model_id === rawModelId &&
e.axes.includes(axis),
);
if (entry) {
allowlistHits.add(`${provider}:${rawModelId}:${axis}`);
return true;
}
return false;
}
// ---------------------------------------------------------------------------
// Comparison helpers
// ---------------------------------------------------------------------------
/**
* Compare effort axes from both interpreters for one (provider, model) pair.
* Returns array of divergence objects.
*/
function compareEffortAxes(provider, model) {
const newResult = resolveNew(provider, model);
const oldResult = getOldEffortConfig(provider, model);
const divergences = [];
// supported_efforts
const newEfforts = newResult.supportedEfforts ?? [];
const oldEfforts = oldResult?.validValues ?? [];
if (JSON.stringify(newEfforts) !== JSON.stringify(oldEfforts)) {
if (!isAllowlisted(provider, model, "supported_efforts")) {
divergences.push({
axis: "supported_efforts",
old: oldEfforts,
new: newEfforts,
});
}
}
// default_effort
const newDefault = newResult.defaultEffort ?? null;
const oldDefault = oldResult?.defaultValue ?? null;
if (newDefault !== oldDefault) {
if (!isAllowlisted(provider, model, "default_effort")) {
divergences.push({
axis: "default_effort",
old: oldDefault,
new: newDefault,
});
}
}
return divergences;
}
// ---------------------------------------------------------------------------
// Test suites
// ---------------------------------------------------------------------------
let totalChecks = 0;
let totalDivergences = 0;
function runCheck(label, provider, model) {
totalChecks++;
const divs = compareEffortAxes(provider, model);
if (divs.length > 0) {
totalDivergences += divs.length;
for (const d of divs) {
console.error(
`DIVERGE [${label}] provider=${provider} model=${model} axis=${d.axis}\n` +
` old: ${JSON.stringify(d.old)}\n` +
` new: ${JSON.stringify(d.new)}`,
);
}
} else if (VERBOSE) {
console.log(`OK [${label}] provider=${provider} model=${model}`);
}
}
// 1. effortTable.fixture.json
console.log("--- effortTable.fixture.json ---");
const fixture = JSON.parse(
readFileSync(
join(repoRoot, "desktop", "src", "features", "agents", "ui", "effortTable.fixture.json"),
"utf8",
),
);
for (const entry of fixture) {
if (!entry.provider) continue;
runCheck("fixture", entry.provider, entry.model ?? "");
}
// 2. normative-corpus.json (effort axes only)
console.log("--- normative-corpus.json ---");
const corpus = JSON.parse(
readFileSync(join(repoRoot, "scripts", "normative-corpus.json"), "utf8"),
);
for (const entry of corpus) {
if (entry._group) continue;
if (!entry.provider || !entry.expect) continue;
if (!entry.expect.supported_efforts && !entry.expect.default_effort) continue;
runCheck("corpus", entry.provider, entry.raw_model_id ?? "");
}
// 3. catalog-sample-fixture.json (exact records from pinned models.dev payload)
console.log("--- catalog-sample-fixture.json ---");
const catalogFixture = JSON.parse(
readFileSync(join(repoRoot, "scripts", "catalog-sample-fixture.json"), "utf8"),
);
for (const ep of catalogFixture.endpoints ?? []) {
if (!ep.name) continue;
// All catalog endpoints are databricks_v2 provider
runCheck("catalog-sample", "databricks_v2", ep.name);
}
// ---------------------------------------------------------------------------
// Summary
// ---------------------------------------------------------------------------
// Count total allowlist axis slots expected to be hit
const totalAllowlistSlots = ALLOWLIST.reduce((n, e) => n + e.axes.length, 0);
const allowlistHitCount = allowlistHits.size;
// Detect stale allowlist entries (declared but never actually suppressed a divergence)
const staleEntries = [];
for (const entry of ALLOWLIST) {
for (const axis of entry.axes) {
const key = `${entry.provider}:${entry.raw_model_id}:${axis}`;
if (!allowlistHits.has(key)) {
staleEntries.push({ ...entry, axis });
}
}
}
console.log(
`\nDifferential: ${totalChecks} checks, ${totalDivergences} unexpected divergences, ${allowlistHitCount}/${totalAllowlistSlots} allowlist slots exercised`,
);
if (staleEntries.length > 0) {
for (const e of staleEntries) {
console.error(
`STALE_ALLOWLIST provider=${e.provider} model=${e.raw_model_id} axis=${e.axis} — entry never fired; remove or update it`,
);
}
}
if (totalDivergences > 0) {
console.error(
`FAIL: ${totalDivergences} unexpected divergence(s) — see output above`,
);
process.exit(1);
} else if (staleEntries.length > 0) {
console.error(
`FAIL: ${staleEntries.length} stale allowlist entry(ies) — entries that never suppress a divergence mask future regressions`,
);
process.exit(1);
} else {
console.log("PASS: old and new effort logic agree on all non-allowlisted entries");
}