Files
buzz/scripts/MODELS_DEV_RECONCILIATION.md
npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7andWill Pfleger d28b9bf411 feat(agent): Phase 2a — wire Rust consumers to generated capability module
Cut config.rs, catalog.rs, and llm.rs over to the generated capability
module from Phase 1. Both old and new paths are preserved; Phase 3
removes the old authorities.

Changes:
- catalog.rs: replace hand-typed DATABRICKS_V2_KNOWN_MODELS literal with
  a re-export of generated_model_capabilities::DATABRICKS_V2_KNOWN_MODELS
- llm.rs: databricks_v2_route_for_model now delegates to
  resolve_model_capabilities("databricks_v2", model). Old segment-based
  classifier and constants moved to #[cfg(test)] under _old_* names for
  the differential harness. New differential test confirms old/new agree
  on all 20 route test vectors (empty allowlist — logic is identical).
- config.rs: add effort_table_fixture_differential_old_vs_new test that
  runs resolve_model_capabilities over the 36-entry effortTable.fixture.json
  and asserts old/new agree mod a doc-cited allowlist of 4 intentional F1
  corrections (gpt-5-5, gpt-5-4-mini, gpt-5-4-nano, gpt-5-6-sol).
- scripts/run-differential.mjs: new JS differential harness running the
  old buzzAgentConfig.ts effort logic vs new modelCapabilities.ts over the
  effortTable fixture (36 entries), normative corpus (45 vectors), and
  catalog-sample fixture (14 endpoints). Passes with allowlist of 5 entries
  (4 models.dev F1 corrections + goose-opus-5 anthropic route correction).
- scripts/MODELS_DEV_RECONCILIATION.md: replace 8 trailing-double-space
  Markdown line breaks with <br> (deferred MINOR from Phase 1 review).

Verification:
- cargo test -p buzz-agent --lib: 426/426 (424 existing + 2 new differential)
- node run-corpus.mjs: 45/45
- node test-manifest-validator.mjs: 24/24 schema-negative
- generate-model-capabilities.mjs --check: byte-clean
- node run-differential.mjs: 85 checks, 0 unexpected divergences
- git diff --check: clean

Co-authored-by: Will Pfleger <pfleger.will@gmail.com>
Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
2026-07-31 12:53:43 -04:00

8.3 KiB

models.dev Reasoning Options Reconciliation Table

Source queried: https://models.dev/api.json (2026-07-31)
Payload SHA-256: d5a4974cd69f19b0f67713acaa6bb3b16e920defdc07ecbdf6b0a936181bb0e0
Policy (plan v4 §Behavior policy): models.dev reasoning_options become exact overrides. Each divergence from the current family rule result is reconciled here: either (a) adopted as an intentional correction or (b) rejected with a curation note.

Verbatim source snapshot: scripts/catalog-sample-fixture.json — verbatim id, name, and nested reasoning_options objects captured from the live API without transformation. Re-verify hash: curl -s https://models.dev/api.json | sha256sum

Divergences

databricks-gpt-5-4-mini

Current family rule (gpt5-4) models.dev Disposition
supported_efforts [none, low, medium, high, xhigh] [low, medium, high] ADOPT

Rationale: The Databricks AI Gateway v2 endpoint for databricks-gpt-5-4-mini explicitly advertises only [low, medium, high] in its reasoning_options. The family rule's none and xhigh are derived from the upstream OpenAI GPT-5.4 spec, which this Databricks endpoint does not expose. Provider-advertised wins per plan F1 policy.

Source: https://models.dev/api.json — retrieved 2026-07-31; providers.databricks.models["databricks-gpt-5-4-mini"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]
Snapshot: scripts/catalog-sample-fixture.json key "databricks-gpt-5-4-mini"
Test vector: resolver-exact-raw-id-hit in scripts/normative-corpus.json


databricks-gpt-5-4-nano

Current family rule (gpt5-4) models.dev Disposition
supported_efforts [none, low, medium, high, xhigh] [low, medium, high] ADOPT

Rationale: Same as databricks-gpt-5-4-mini. The nano variant exposes the same restricted effort set. Provider-advertised wins.

Source: https://models.dev/api.json — retrieved 2026-07-31; providers.databricks.models["databricks-gpt-5-4-nano"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]
Snapshot: scripts/catalog-sample-fixture.json key "databricks-gpt-5-4-nano"


databricks-gpt-5-6-sol

Current family rule (gpt5-6) models.dev Disposition
supported_efforts [none, low, medium, high, xhigh, max] [low, medium, high, max] ADOPT

Rationale: The Databricks AI Gateway v2 endpoint for databricks-gpt-5-6-sol advertises only [low, medium, high, max] in its reasoning_options. The family rule's none and xhigh are derived from the upstream OpenAI GPT-5.6 spec, which this Databricks endpoint does not expose. Provider-advertised wins per plan F1 policy.

Source: https://models.dev/api.json — retrieved 2026-07-31; providers.databricks.models["databricks-gpt-5-6-sol"].reasoning_options = [{"type":"effort","values":["low","medium","high","max"]}]
Snapshot: scripts/catalog-sample-fixture.json key "databricks-gpt-5-6-sol"


databricks-gpt-5-5

Current family rule (gpt5-5) models.dev Disposition
supported_efforts [none, low, medium, high, xhigh] [low, medium, high] ADOPT

Rationale: The Databricks AI Gateway v2 endpoint for databricks-gpt-5-5 advertises only [low, medium, high] in its reasoning_options. The family rule's none and xhigh are derived from the upstream OpenAI GPT-5.5 spec, which this Databricks endpoint does not expose. Provider-advertised wins per plan F1 policy.

Source: https://models.dev/api.json — retrieved 2026-07-31; providers.databricks.models["databricks-gpt-5-5"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]
Snapshot: scripts/catalog-sample-fixture.json key "databricks-gpt-5-5"


databricks-claude-opus-4-7

Current family rule (anthropic-adaptive-xhigh-opus-4-7) models.dev Disposition
reasoning_options type effort-based budget_tokens NO EFFORT DIVERGENCE

Rationale: models.dev advertises reasoning_options=[{"type":"budget_tokens","min":1024}] — a different capability axis (extended thinking token budget), not an effort-level selector. There is no effort divergence to reconcile. The effort capabilities for this model come from the anthropic-adaptive-xhigh-opus-4-7 family rule (Anthropic extended-thinking support table).

Source: https://models.dev/api.json — retrieved 2026-07-31; providers.databricks.models["databricks-claude-opus-4-7"].reasoning_options = [{"type":"budget_tokens","min":1024}]
Snapshot: scripts/catalog-sample-fixture.json key "databricks-claude-opus-4-7"


Non-divergences (confirmed consistent)

The following models were checked against models.dev or provider docs and found consistent with the manifest family rules. No exact records needed.

Model family Source Checked against Status
claude-opus-4-7 https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
claude-opus-4-8 https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
claude-sonnet-5.* https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
claude-fable-5 https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
claude-mythos-5 https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
claude-opus-4-6 https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
claude-sonnet-4-6 https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
claude-mythos-preview https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
claude-3* https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
gpt-5-pro https://platform.openai.com/docs/guides/reasoning OpenAI reasoning guide (July 2025) ✓ Consistent
gpt-5.6 https://platform.openai.com/docs/guides/reasoning OpenAI reasoning guide (July 2025) ✓ Consistent
gpt-5.5 https://platform.openai.com/docs/guides/reasoning OpenAI reasoning guide (July 2025) ✓ Consistent
gpt-5.4 https://platform.openai.com/docs/guides/reasoning OpenAI reasoning guide (July 2025) ✓ Consistent
gpt-5.1 https://platform.openai.com/docs/guides/reasoning OpenAI reasoning guide (July 2025) ✓ Consistent
gpt-5 (base) https://platform.openai.com/docs/guides/reasoning OpenAI reasoning guide (July 2025) ✓ Consistent