Files
buzz/scripts/normative-corpus.json
T
4d47f48143 feat(agent): Phase 1 — model-capability manifest, generator, and test oracle (#3821)
## What this does

Introduces the model-capability manifest infrastructure (Phase 1 of the
Model-Capability Manifest plan v4, Thufir-approved 9/9/9). No consumer
cutover — `config.rs`, `llm.rs`, `catalog.rs`, and `buzzAgentConfig.ts`
are unchanged. Phase 2 wires them.

**Single source of truth** replaces hand-mirrored metadata across four
files:

```
scripts/model-capabilities.json         → hand-curated manifest
scripts/generate-model-capabilities.mjs → emits Rust + TS artifacts
crates/buzz-agent/src/generated_model_capabilities.rs
desktop/src/features/agents/ui/modelCapabilities.ts
```

## Resolver contract (plan v4 §Resolver contract)

Total function `resolve(provider, raw_model_id) → CapabilityResult`.
Three ordered steps:

1. Provider-qualified raw exact lookup — key is `(provider,
raw_model_id)`, matched before any prefix stripping. A prefixed alias
never inherits an exact record.
2. Provider-scoped ordered family rules — on normalized
(prefix-stripped) alias, by `match_priority` desc.
3. Per-axis provider fallback — `blank` vs `concrete_unknown`, per
provider.

Result is complete — every axis populated, runtime consumers never
compose fields.

## Boundaries the manifest does NOT own

- Transport for pure OpenAI, legacy Databricks, OpenRouter:
`OpenAiApi`/`openai_request()` remain authoritative.
`databricks_v2_wire_route` is DBv2-only (all other providers emit
`not-applicable`).
- Final display labels: `resolveModelLabel()` three-tier precedence
unchanged. `registry_label` feeds only the static registry tier.
- `llm.rs` scope: only `databricks_v2_route_for_model` (Phase 2).

## Test oracle (three independent layers)

1. Generated full-table coverage —
`scripts/generated-model-capabilities-coverage.json`: every manifest
entry + provider fallbacks.
2. Hand-authored normative corpus — `scripts/normative-corpus.json` (44
vectors): Anthropic manual-budget/adaptive families, OpenAI gpt-5
adversarial boundary cases, DBv2 segment-routing collision tests, P2-A
resolver-contract vectors, P2-B blank/concrete-unknown per provider.
Runs against JS resolver (`run-corpus.mjs`) and mirrored in Rust
(`generated_model_capabilities_tests.rs`).
3. Schema-negative tests — `scripts/test-manifest-validator.mjs` (17
tests): every validator rule has a failing-input test.

## Reconciliation table

`scripts/MODELS_DEV_RECONCILIATION.md` — all models.dev divergences
dispositioned. `databricks-gpt-5-4-mini` and `databricks-gpt-5-4-nano`
adopt models.dev `[low,medium,high]` (family rule adds `none+xhigh` the
endpoint doesn't advertise).

## CI

`.github/workflows/model-capability-regen-diff.yml`: triggers on
manifest/generator/artifact changes; regenerates and fails if stale;
runs JS corpus + schema-negative tests.

## Acceptance criteria (plan v4 Phase 1)

- Byte-clean regen: `node scripts/generate-model-capabilities.mjs
--check` passes
- Rust compiles: `cargo check -p buzz-agent`
- TS typechecks: `pnpm tsc --noEmit --strict`
- 44/44 normative corpus vectors pass (JS interpreter)
- 41/41 Rust corpus tests pass
- 17/17 schema-negative tests pass (every validator rule)
- Reconciliation table complete with doc citations
- No consumer changes (config.rs, llm.rs, catalog.rs, buzzAgentConfig.ts
untouched)

---------

Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
Co-authored-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
2026-07-31 12:39:37 -04:00

486 lines
16 KiB
JSON

[
{
"_group": "Anthropic exact family rules",
"_note": "All require thinking_mode=manual-budget or adaptive, correct supported_efforts, databricks_v2_wire_route=not-applicable"
},
{
"id": "anthropic-claude-3-family",
"provider": "anthropic",
"raw_model_id": "claude-3-7-sonnet-20250219",
"expect": {
"thinking_mode": "manual-budget",
"supported_efforts": ["low", "medium", "high"],
"default_effort": null,
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "anthropic-claude-opus-4-5",
"provider": "anthropic",
"raw_model_id": "claude-opus-4-5",
"expect": {
"thinking_mode": "manual-budget",
"supported_efforts": ["low", "medium", "high"],
"default_effort": null,
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "anthropic-claude-opus-4-7",
"provider": "anthropic",
"raw_model_id": "claude-opus-4-7",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "anthropic-claude-opus-4-8",
"provider": "anthropic",
"raw_model_id": "claude-opus-4-8",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "anthropic-claude-sonnet-5",
"provider": "anthropic",
"raw_model_id": "claude-sonnet-5-20260101",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "anthropic-claude-fable-5",
"provider": "anthropic",
"raw_model_id": "claude-fable-5",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "anthropic-claude-mythos-5",
"provider": "anthropic",
"raw_model_id": "claude-mythos-5",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "anthropic-claude-opus-4-6",
"provider": "anthropic",
"raw_model_id": "claude-opus-4-6",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "max"],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "anthropic-claude-sonnet-4-6",
"provider": "anthropic",
"raw_model_id": "claude-sonnet-4-6",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "max"],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "anthropic-claude-mythos-preview",
"provider": "anthropic",
"raw_model_id": "claude-mythos-preview",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "max"],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"_group": "Anthropic unknowns and fallbacks"
},
{
"id": "anthropic-unknown-blank",
"provider": "anthropic",
"raw_model_id": "",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "anthropic-unknown-concrete",
"provider": "anthropic",
"raw_model_id": "claude-ultra-9000",
"expect": {
"thinking_mode": "omit-fields",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"_group": "OpenAI exact family rules"
},
{
"id": "openai-gpt5-pro",
"provider": "openai",
"raw_model_id": "gpt-5-pro",
"expect": {
"thinking_mode": "none",
"supported_efforts": ["high"],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "openai-gpt5.6",
"provider": "openai",
"raw_model_id": "gpt-5.6",
"expect": {
"thinking_mode": "none",
"supported_efforts": ["none", "low", "medium", "high", "xhigh", "max"],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "openai-gpt5-6-dashed",
"provider": "openai",
"raw_model_id": "gpt-5-6",
"expect": {
"thinking_mode": "none",
"supported_efforts": ["none", "low", "medium", "high", "xhigh", "max"],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "openai-gpt5.5",
"provider": "openai",
"raw_model_id": "gpt-5.5",
"expect": {
"thinking_mode": "none",
"supported_efforts": ["none", "low", "medium", "high", "xhigh"],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "openai-gpt5.4",
"provider": "openai",
"raw_model_id": "gpt-5.4",
"expect": {
"thinking_mode": "none",
"supported_efforts": ["none", "low", "medium", "high", "xhigh"],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "openai-gpt5.1",
"provider": "openai",
"raw_model_id": "gpt-5.1",
"expect": {
"thinking_mode": "none",
"supported_efforts": ["none", "low", "medium", "high"],
"default_effort": "none",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "openai-gpt5-base",
"provider": "openai",
"raw_model_id": "gpt-5",
"expect": {
"thinking_mode": "none",
"supported_efforts": ["minimal", "low", "medium", "high"],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"_group": "OpenAI adversarial — gpt5 boundary-aware matching (ported from config.rs tests)"
},
{
"id": "openai-gpt5-1106-should-not-match-base",
"provider": "openai",
"raw_model_id": "gpt-5-1106",
"_note": "gpt-5-1106: '-1106' is a 4-digit date segment, NOT a short version (gpt5-base rejects only 1-3 digit suffixes). Must match base table [minimal,low,medium,high], NOT fall through to unknown.",
"expect": {
"supported_efforts": ["minimal", "low", "medium", "high"]
}
},
{
"id": "openai-gpt5-4o-matches-base",
"provider": "openai",
"raw_model_id": "gpt-5-4o",
"_note": "gpt-5-4o: '4o' after '-' is NOT a short numeric suffix (it contains a letter). Must match gpt5-base. Crucially, must NOT match gpt-5.4 (the '4' is followed by 'o', not boundary char).",
"expect": {
"supported_efforts": ["minimal", "low", "medium", "high"]
}
},
{
"id": "openai-gpt5-pro-not-matching-gpt5-base",
"provider": "openai",
"raw_model_id": "gpt-5-pro",
"_note": "gpt-5-pro should hit gpt5-pro rule (priority 20), NOT gpt-5 base.",
"expect": {
"supported_efforts": ["high"],
"default_effort": "high"
}
},
{
"id": "openai-multi-digit-version-gpt5-10",
"provider": "openai",
"raw_model_id": "gpt-5-10",
"_note": "gpt-5-10 — two-digit suffix prevents gpt5-base match. Falls through to unknown.",
"expect": {
"supported_efforts": ["none", "minimal", "low", "medium", "high", "xhigh"]
}
},
{
"id": "openai-gpt5-date-suffix",
"provider": "openai",
"raw_model_id": "gpt-5-20260101",
"_note": "gpt-5-20260101 — long numeric suffix after base: '20260101' is 8 digits, beyond 1-3 digit reject, should hit gpt5-base.",
"expect": {
"supported_efforts": ["minimal", "low", "medium", "high"]
}
},
{
"_group": "DatabricksV2 — segment-based routing (ported from llm.rs tests)"
},
{
"id": "dbv2-gpt5-route-openai-responses",
"provider": "databricks_v2",
"raw_model_id": "gpt-5.5",
"expect": {
"databricks_v2_wire_route": "openai-responses",
"supported_efforts": ["none", "low", "medium", "high", "xhigh"]
}
},
{
"id": "dbv2-claude-route-anthropic-messages",
"provider": "databricks_v2",
"raw_model_id": "claude-opus-4-7",
"expect": {
"databricks_v2_wire_route": "anthropic-messages",
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"]
}
},
{
"id": "dbv2-claude-prefix-stripped",
"provider": "databricks_v2",
"raw_model_id": "databricks-claude-opus-4-7",
"_note": "databricks- prefix stripped → claude-opus-4-7 → Anthropic route",
"expect": {
"databricks_v2_wire_route": "anthropic-messages",
"thinking_mode": "adaptive"
}
},
{
"id": "dbv2-goose-claude-prefix-stripped",
"provider": "databricks_v2",
"raw_model_id": "goose-claude-fable-5",
"_note": "goose- prefix stripped → claude-fable-5 → Anthropic adaptive+xhigh",
"expect": {
"databricks_v2_wire_route": "anthropic-messages",
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"]
}
},
{
"id": "dbv2-team-prefix-stripped",
"provider": "databricks_v2",
"raw_model_id": "team-x-claude-opus-4-7",
"_note": "team-x- prefix stripped → claude-opus-4-7 → Anthropic route",
"expect": {
"databricks_v2_wire_route": "anthropic-messages",
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"]
}
},
{
"id": "dbv2-consolidated-llama-not-sol",
"provider": "databricks_v2",
"raw_model_id": "consolidated-llama",
"_note": "segment test: 'sol' is a SUBSTRING of 'consolidated' — must NOT match DATABRICKS_V2_OPENAI_CODE_NAMES 'sol'. Falls through to mlflow-chat.",
"expect": {
"databricks_v2_wire_route": "mlflow-chat"
}
},
{
"id": "dbv2-terraform-coder-not-terra",
"provider": "databricks_v2",
"raw_model_id": "terraform-coder",
"_note": "segment test: 'terra' is a prefix of 'terraform' — must NOT match 'terra' code name. Falls through to mlflow-chat.",
"expect": {
"databricks_v2_wire_route": "mlflow-chat"
}
},
{
"id": "dbv2-corpus-reranker-not-opus",
"provider": "databricks_v2",
"raw_model_id": "corpus-reranker",
"_note": "segment test: 'opus' is NOT a segment of corpus-reranker (segments: corpus, reranker). mlflow-chat.",
"expect": {
"databricks_v2_wire_route": "mlflow-chat"
}
},
{
"id": "dbv2-octopus-model-not-opus",
"provider": "databricks_v2",
"raw_model_id": "octopus-model",
"_note": "segment test: 'opus' is not a segment of octopus-model (segments: octopus, model). mlflow-chat.",
"expect": {
"databricks_v2_wire_route": "mlflow-chat"
}
},
{
"id": "dbv2-goose-opus-5-is-anthropic",
"provider": "databricks_v2",
"raw_model_id": "goose-opus-5",
"_note": "'opus' IS a named segment of goose-opus-5 (segments: goose, opus, 5). Routes Anthropic. Key test: agrees with llm.rs but disagreed with old config.rs.",
"expect": {
"databricks_v2_wire_route": "anthropic-messages"
}
},
{
"_group": "P2-A resolver-contract vectors (plan v4 §Resolver contract)"
},
{
"id": "resolver-exact-raw-id-hit",
"provider": "databricks_v2",
"raw_model_id": "databricks-gpt-5-4-mini",
"_note": "Exact record exists. Must return exact Databricks override: low|medium|high (not family's none+xhigh).",
"expect": {
"supported_efforts": ["low", "medium", "high"]
}
},
{
"id": "resolver-prefixed-alias-misses-exact",
"provider": "databricks_v2",
"raw_model_id": "team-x-databricks-gpt-5-4-mini",
"_note": "Prefixed alias of an exact ID. Raw exact lookup MUST miss (key is team-x-..., not databricks-...). Falls to family rules (gpt5-4 family → none+xhigh).",
"expect": {
"supported_efforts": ["none", "low", "medium", "high", "xhigh"]
}
},
{
"id": "resolver-cross-provider-misses-exact",
"provider": "openai",
"raw_model_id": "databricks-gpt-5-4-mini",
"_note": "Same raw ID but different provider. Exact record is databricks_v2-scoped; must miss. Falls to openai family rules.",
"expect": {
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "resolver-exact-efforts-plus-family-route",
"provider": "databricks_v2",
"raw_model_id": "databricks-gpt-5-6-sol",
"_note": "Exact record with efforts from models.dev (low|medium|high|max — provider-advertised, no none/xhigh). Route materialized from gpt5-6 family rule (openai-responses). Must return both, complete.",
"expect": {
"supported_efforts": ["low", "medium", "high", "max"],
"databricks_v2_wire_route": "openai-responses"
}
},
{
"id": "dbv2-gpt5-5-exact-override",
"provider": "databricks_v2",
"raw_model_id": "databricks-gpt-5-5",
"_note": "Exact record adopts models.dev advertised set [low,medium,high]. Family rule (gpt5-5) has none+xhigh — provider-advertised wins per plan F1.",
"expect": {
"supported_efforts": ["low", "medium", "high"],
"databricks_v2_wire_route": "openai-responses"
}
},
{
"_group": "Blank vs concrete-unknown per provider (P2-B fallback vectors)"
},
{
"id": "dbv2-blank-all7-route-unknown",
"provider": "databricks_v2",
"raw_model_id": "",
"_note": "DBv2 blank: route-unknown, all 7 efforts, default medium.",
"expect": {
"databricks_v2_wire_route": "route-unknown",
"supported_efforts": ["none", "minimal", "low", "medium", "high", "xhigh", "max"],
"default_effort": "medium"
}
},
{
"id": "dbv2-concrete-unknown-mlflow-no-max",
"provider": "databricks_v2",
"raw_model_id": "some-unknown-model-xyz",
"_note": "DBv2 concrete-unknown: mlflow-chat, all-except-max (6 efforts).",
"expect": {
"databricks_v2_wire_route": "mlflow-chat",
"supported_efforts": ["none", "minimal", "low", "medium", "high", "xhigh"]
}
},
{
"id": "openai-blank-all-except-max",
"provider": "openai",
"raw_model_id": "",
"_note": "OpenAI blank: not-applicable route, all-except-max, medium default.",
"expect": {
"databricks_v2_wire_route": "not-applicable",
"supported_efforts": ["none", "minimal", "low", "medium", "high", "xhigh"],
"default_effort": "medium"
}
},
{
"id": "openai-concrete-unknown-all-except-max",
"provider": "openai",
"raw_model_id": "gpt-4o",
"_note": "OpenAI concrete unknown (unverified family): not-applicable route, all-except-max, medium default.",
"expect": {
"databricks_v2_wire_route": "not-applicable",
"supported_efforts": ["none", "minimal", "low", "medium", "high", "xhigh"],
"default_effort": "medium"
}
},
{
"id": "anthropic-blank-adaptive-full",
"provider": "anthropic",
"raw_model_id": "",
"_note": "Anthropic blank: assume adaptive with full support (incl. xhigh).",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
"default_effort": "high"
}
},
{
"id": "anthropic-concrete-unknown-omit-fields",
"provider": "anthropic",
"raw_model_id": "claude-ultra-9000",
"_note": "Anthropic concrete-unknown: omit-fields (never guess request shape).",
"expect": {
"thinking_mode": "omit-fields"
}
}
]