Files
buzz/scripts/MODELS_DEV_RECONCILIATION.md
T
4d47f48143 feat(agent): Phase 1 — model-capability manifest, generator, and test oracle (#3821)
## What this does

Introduces the model-capability manifest infrastructure (Phase 1 of the
Model-Capability Manifest plan v4, Thufir-approved 9/9/9). No consumer
cutover — `config.rs`, `llm.rs`, `catalog.rs`, and `buzzAgentConfig.ts`
are unchanged. Phase 2 wires them.

**Single source of truth** replaces hand-mirrored metadata across four
files:

```
scripts/model-capabilities.json         → hand-curated manifest
scripts/generate-model-capabilities.mjs → emits Rust + TS artifacts
crates/buzz-agent/src/generated_model_capabilities.rs
desktop/src/features/agents/ui/modelCapabilities.ts
```

## Resolver contract (plan v4 §Resolver contract)

Total function `resolve(provider, raw_model_id) → CapabilityResult`.
Three ordered steps:

1. Provider-qualified raw exact lookup — key is `(provider,
raw_model_id)`, matched before any prefix stripping. A prefixed alias
never inherits an exact record.
2. Provider-scoped ordered family rules — on normalized
(prefix-stripped) alias, by `match_priority` desc.
3. Per-axis provider fallback — `blank` vs `concrete_unknown`, per
provider.

Result is complete — every axis populated, runtime consumers never
compose fields.

## Boundaries the manifest does NOT own

- Transport for pure OpenAI, legacy Databricks, OpenRouter:
`OpenAiApi`/`openai_request()` remain authoritative.
`databricks_v2_wire_route` is DBv2-only (all other providers emit
`not-applicable`).
- Final display labels: `resolveModelLabel()` three-tier precedence
unchanged. `registry_label` feeds only the static registry tier.
- `llm.rs` scope: only `databricks_v2_route_for_model` (Phase 2).

## Test oracle (three independent layers)

1. Generated full-table coverage —
`scripts/generated-model-capabilities-coverage.json`: every manifest
entry + provider fallbacks.
2. Hand-authored normative corpus — `scripts/normative-corpus.json` (44
vectors): Anthropic manual-budget/adaptive families, OpenAI gpt-5
adversarial boundary cases, DBv2 segment-routing collision tests, P2-A
resolver-contract vectors, P2-B blank/concrete-unknown per provider.
Runs against JS resolver (`run-corpus.mjs`) and mirrored in Rust
(`generated_model_capabilities_tests.rs`).
3. Schema-negative tests — `scripts/test-manifest-validator.mjs` (17
tests): every validator rule has a failing-input test.

## Reconciliation table

`scripts/MODELS_DEV_RECONCILIATION.md` — all models.dev divergences
dispositioned. `databricks-gpt-5-4-mini` and `databricks-gpt-5-4-nano`
adopt models.dev `[low,medium,high]` (family rule adds `none+xhigh` the
endpoint doesn't advertise).

## CI

`.github/workflows/model-capability-regen-diff.yml`: triggers on
manifest/generator/artifact changes; regenerates and fails if stale;
runs JS corpus + schema-negative tests.

## Acceptance criteria (plan v4 Phase 1)

- Byte-clean regen: `node scripts/generate-model-capabilities.mjs
--check` passes
- Rust compiles: `cargo check -p buzz-agent`
- TS typechecks: `pnpm tsc --noEmit --strict`
- 44/44 normative corpus vectors pass (JS interpreter)
- 41/41 Rust corpus tests pass
- 17/17 schema-negative tests pass (every validator rule)
- Reconciliation table complete with doc citations
- No consumer changes (config.rs, llm.rs, catalog.rs, buzzAgentConfig.ts
untouched)

---------

Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
Co-authored-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
2026-07-31 12:39:37 -04:00

8.3 KiB

models.dev Reasoning Options Reconciliation Table

Source queried: https://models.dev/api.json (2026-07-31)
Payload SHA-256: d5a4974cd69f19b0f67713acaa6bb3b16e920defdc07ecbdf6b0a936181bb0e0
Policy (plan v4 §Behavior policy): models.dev reasoning_options become exact overrides. Each divergence from the current family rule result is reconciled here: either (a) adopted as an intentional correction or (b) rejected with a curation note.

Verbatim source snapshot: scripts/catalog-sample-fixture.json — verbatim id, name, and nested reasoning_options objects captured from the live API without transformation. Re-verify hash: curl -s https://models.dev/api.json | sha256sum

Divergences

databricks-gpt-5-4-mini

Current family rule (gpt5-4) models.dev Disposition
supported_efforts [none, low, medium, high, xhigh] [low, medium, high] ADOPT

Rationale: The Databricks AI Gateway v2 endpoint for databricks-gpt-5-4-mini explicitly advertises only [low, medium, high] in its reasoning_options. The family rule's none and xhigh are derived from the upstream OpenAI GPT-5.4 spec, which this Databricks endpoint does not expose. Provider-advertised wins per plan F1 policy.

Source: https://models.dev/api.json — retrieved 2026-07-31; providers.databricks.models["databricks-gpt-5-4-mini"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]
Snapshot: scripts/catalog-sample-fixture.json key "databricks-gpt-5-4-mini"
Test vector: resolver-exact-raw-id-hit in scripts/normative-corpus.json


databricks-gpt-5-4-nano

Current family rule (gpt5-4) models.dev Disposition
supported_efforts [none, low, medium, high, xhigh] [low, medium, high] ADOPT

Rationale: Same as databricks-gpt-5-4-mini. The nano variant exposes the same restricted effort set. Provider-advertised wins.

Source: https://models.dev/api.json — retrieved 2026-07-31; providers.databricks.models["databricks-gpt-5-4-nano"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]
Snapshot: scripts/catalog-sample-fixture.json key "databricks-gpt-5-4-nano"


databricks-gpt-5-6-sol

Current family rule (gpt5-6) models.dev Disposition
supported_efforts [none, low, medium, high, xhigh, max] [low, medium, high, max] ADOPT

Rationale: The Databricks AI Gateway v2 endpoint for databricks-gpt-5-6-sol advertises only [low, medium, high, max] in its reasoning_options. The family rule's none and xhigh are derived from the upstream OpenAI GPT-5.6 spec, which this Databricks endpoint does not expose. Provider-advertised wins per plan F1 policy.

Source: https://models.dev/api.json — retrieved 2026-07-31; providers.databricks.models["databricks-gpt-5-6-sol"].reasoning_options = [{"type":"effort","values":["low","medium","high","max"]}]
Snapshot: scripts/catalog-sample-fixture.json key "databricks-gpt-5-6-sol"


databricks-gpt-5-5

Current family rule (gpt5-5) models.dev Disposition
supported_efforts [none, low, medium, high, xhigh] [low, medium, high] ADOPT

Rationale: The Databricks AI Gateway v2 endpoint for databricks-gpt-5-5 advertises only [low, medium, high] in its reasoning_options. The family rule's none and xhigh are derived from the upstream OpenAI GPT-5.5 spec, which this Databricks endpoint does not expose. Provider-advertised wins per plan F1 policy.

Source: https://models.dev/api.json — retrieved 2026-07-31; providers.databricks.models["databricks-gpt-5-5"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]
Snapshot: scripts/catalog-sample-fixture.json key "databricks-gpt-5-5"


databricks-claude-opus-4-7

Current family rule (anthropic-adaptive-xhigh-opus-4-7) models.dev Disposition
reasoning_options type effort-based budget_tokens NO EFFORT DIVERGENCE

Rationale: models.dev advertises reasoning_options=[{"type":"budget_tokens","min":1024}] — a different capability axis (extended thinking token budget), not an effort-level selector. There is no effort divergence to reconcile. The effort capabilities for this model come from the anthropic-adaptive-xhigh-opus-4-7 family rule (Anthropic extended-thinking support table).

Source: https://models.dev/api.json — retrieved 2026-07-31; providers.databricks.models["databricks-claude-opus-4-7"].reasoning_options = [{"type":"budget_tokens","min":1024}]
Snapshot: scripts/catalog-sample-fixture.json key "databricks-claude-opus-4-7"


Non-divergences (confirmed consistent)

The following models were checked against models.dev or provider docs and found consistent with the manifest family rules. No exact records needed.

Model family Source Checked against Status
claude-opus-4-7 https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
claude-opus-4-8 https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
claude-sonnet-5.* https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
claude-fable-5 https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
claude-mythos-5 https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
claude-opus-4-6 https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
claude-sonnet-4-6 https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
claude-mythos-preview https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
claude-3* https://platform.claude.com/docs/en/build-with-claude/extended-thinking Anthropic extended-thinking support table (July 2025) ✓ Consistent
gpt-5-pro https://platform.openai.com/docs/guides/reasoning OpenAI reasoning guide (July 2025) ✓ Consistent
gpt-5.6 https://platform.openai.com/docs/guides/reasoning OpenAI reasoning guide (July 2025) ✓ Consistent
gpt-5.5 https://platform.openai.com/docs/guides/reasoning OpenAI reasoning guide (July 2025) ✓ Consistent
gpt-5.4 https://platform.openai.com/docs/guides/reasoning OpenAI reasoning guide (July 2025) ✓ Consistent
gpt-5.1 https://platform.openai.com/docs/guides/reasoning OpenAI reasoning guide (July 2025) ✓ Consistent
gpt-5 (base) https://platform.openai.com/docs/guides/reasoning OpenAI reasoning guide (July 2025) ✓ Consistent