feat(agent): Phase 1 — model-capability manifest, generator, and test oracle (#3821)

## What this does

Introduces the model-capability manifest infrastructure (Phase 1 of the
Model-Capability Manifest plan v4, Thufir-approved 9/9/9). No consumer
cutover — `config.rs`, `llm.rs`, `catalog.rs`, and `buzzAgentConfig.ts`
are unchanged. Phase 2 wires them.

**Single source of truth** replaces hand-mirrored metadata across four
files:

```
scripts/model-capabilities.json         → hand-curated manifest
scripts/generate-model-capabilities.mjs → emits Rust + TS artifacts
crates/buzz-agent/src/generated_model_capabilities.rs
desktop/src/features/agents/ui/modelCapabilities.ts
```

## Resolver contract (plan v4 §Resolver contract)

Total function `resolve(provider, raw_model_id) → CapabilityResult`.
Three ordered steps:

1. Provider-qualified raw exact lookup — key is `(provider,
raw_model_id)`, matched before any prefix stripping. A prefixed alias
never inherits an exact record.
2. Provider-scoped ordered family rules — on normalized
(prefix-stripped) alias, by `match_priority` desc.
3. Per-axis provider fallback — `blank` vs `concrete_unknown`, per
provider.

Result is complete — every axis populated, runtime consumers never
compose fields.

## Boundaries the manifest does NOT own

- Transport for pure OpenAI, legacy Databricks, OpenRouter:
`OpenAiApi`/`openai_request()` remain authoritative.
`databricks_v2_wire_route` is DBv2-only (all other providers emit
`not-applicable`).
- Final display labels: `resolveModelLabel()` three-tier precedence
unchanged. `registry_label` feeds only the static registry tier.
- `llm.rs` scope: only `databricks_v2_route_for_model` (Phase 2).

## Test oracle (three independent layers)

1. Generated full-table coverage —
`scripts/generated-model-capabilities-coverage.json`: every manifest
entry + provider fallbacks.
2. Hand-authored normative corpus — `scripts/normative-corpus.json` (44
vectors): Anthropic manual-budget/adaptive families, OpenAI gpt-5
adversarial boundary cases, DBv2 segment-routing collision tests, P2-A
resolver-contract vectors, P2-B blank/concrete-unknown per provider.
Runs against JS resolver (`run-corpus.mjs`) and mirrored in Rust
(`generated_model_capabilities_tests.rs`).
3. Schema-negative tests — `scripts/test-manifest-validator.mjs` (17
tests): every validator rule has a failing-input test.

## Reconciliation table

`scripts/MODELS_DEV_RECONCILIATION.md` — all models.dev divergences
dispositioned. `databricks-gpt-5-4-mini` and `databricks-gpt-5-4-nano`
adopt models.dev `[low,medium,high]` (family rule adds `none+xhigh` the
endpoint doesn't advertise).

## CI

`.github/workflows/model-capability-regen-diff.yml`: triggers on
manifest/generator/artifact changes; regenerates and fails if stale;
runs JS corpus + schema-negative tests.

## Acceptance criteria (plan v4 Phase 1)

- Byte-clean regen: `node scripts/generate-model-capabilities.mjs
--check` passes
- Rust compiles: `cargo check -p buzz-agent`
- TS typechecks: `pnpm tsc --noEmit --strict`
- 44/44 normative corpus vectors pass (JS interpreter)
- 41/41 Rust corpus tests pass
- 17/17 schema-negative tests pass (every validator rule)
- Reconciliation table complete with doc citations
- No consumer changes (config.rs, llm.rs, catalog.rs, buzzAgentConfig.ts
untouched)

---------

Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
Co-authored-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
This commit is contained in:
Will Pfleger
2026-07-31 12:39:37 -04:00
committed by GitHub
co-authored by npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7
parent 485b36280a
commit 4d47f48143
16 changed files with 9079 additions and 0 deletions
+115
View File
@@ -0,0 +1,115 @@
# models.dev Reasoning Options Reconciliation Table
**Source queried**: https://models.dev/api.json (2026-07-31)
**Payload SHA-256**: `d5a4974cd69f19b0f67713acaa6bb3b16e920defdc07ecbdf6b0a936181bb0e0`
**Policy (plan v4 §Behavior policy)**: models.dev `reasoning_options` become exact overrides.
Each divergence from the current family rule result is reconciled here: either (a) adopted as an
intentional correction or (b) rejected with a curation note.
**Verbatim source snapshot**: `scripts/catalog-sample-fixture.json` — verbatim `id`, `name`, and
nested `reasoning_options` objects captured from the live API without transformation.
Re-verify hash: `curl -s https://models.dev/api.json | sha256sum`
## Divergences
### `databricks-gpt-5-4-mini`
| | Current family rule (gpt5-4) | models.dev | Disposition |
|---|---|---|---|
| `supported_efforts` | `[none, low, medium, high, xhigh]` | `[low, medium, high]` | **ADOPT** |
**Rationale**: The Databricks AI Gateway v2 endpoint for `databricks-gpt-5-4-mini` explicitly
advertises only `[low, medium, high]` in its `reasoning_options`. The family rule's `none` and
`xhigh` are derived from the upstream OpenAI GPT-5.4 spec, which this Databricks endpoint does
not expose. Provider-advertised wins per plan F1 policy.
**Source**: [https://models.dev/api.json](https://models.dev/api.json) — retrieved 2026-07-31; `providers.databricks.models["databricks-gpt-5-4-mini"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]`
**Snapshot**: `scripts/catalog-sample-fixture.json` key `"databricks-gpt-5-4-mini"`
**Test vector**: `resolver-exact-raw-id-hit` in `scripts/normative-corpus.json`
---
### `databricks-gpt-5-4-nano`
| | Current family rule (gpt5-4) | models.dev | Disposition |
|---|---|---|---|
| `supported_efforts` | `[none, low, medium, high, xhigh]` | `[low, medium, high]` | **ADOPT** |
**Rationale**: Same as `databricks-gpt-5-4-mini`. The nano variant exposes the same restricted
effort set. Provider-advertised wins.
**Source**: [https://models.dev/api.json](https://models.dev/api.json) — retrieved 2026-07-31; `providers.databricks.models["databricks-gpt-5-4-nano"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]`
**Snapshot**: `scripts/catalog-sample-fixture.json` key `"databricks-gpt-5-4-nano"`
---
### `databricks-gpt-5-6-sol`
| | Current family rule (gpt5-6) | models.dev | Disposition |
|---|---|---|---|
| `supported_efforts` | `[none, low, medium, high, xhigh, max]` | `[low, medium, high, max]` | **ADOPT** |
**Rationale**: The Databricks AI Gateway v2 endpoint for `databricks-gpt-5-6-sol` advertises only
`[low, medium, high, max]` in its `reasoning_options`. The family rule's `none` and `xhigh` are
derived from the upstream OpenAI GPT-5.6 spec, which this Databricks endpoint does not expose.
Provider-advertised wins per plan F1 policy.
**Source**: [https://models.dev/api.json](https://models.dev/api.json) — retrieved 2026-07-31; `providers.databricks.models["databricks-gpt-5-6-sol"].reasoning_options = [{"type":"effort","values":["low","medium","high","max"]}]`
**Snapshot**: `scripts/catalog-sample-fixture.json` key `"databricks-gpt-5-6-sol"`
---
### `databricks-gpt-5-5`
| | Current family rule (gpt5-5) | models.dev | Disposition |
|---|---|---|---|
| `supported_efforts` | `[none, low, medium, high, xhigh]` | `[low, medium, high]` | **ADOPT** |
**Rationale**: The Databricks AI Gateway v2 endpoint for `databricks-gpt-5-5` advertises only
`[low, medium, high]` in its `reasoning_options`. The family rule's `none` and `xhigh` are
derived from the upstream OpenAI GPT-5.5 spec, which this Databricks endpoint does not expose.
Provider-advertised wins per plan F1 policy.
**Source**: [https://models.dev/api.json](https://models.dev/api.json) — retrieved 2026-07-31; `providers.databricks.models["databricks-gpt-5-5"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]`
**Snapshot**: `scripts/catalog-sample-fixture.json` key `"databricks-gpt-5-5"`
---
### `databricks-claude-opus-4-7`
| | Current family rule (anthropic-adaptive-xhigh-opus-4-7) | models.dev | Disposition |
|---|---|---|---|
| `reasoning_options` type | effort-based | `budget_tokens` | **NO EFFORT DIVERGENCE** |
**Rationale**: models.dev advertises `reasoning_options=[{"type":"budget_tokens","min":1024}]` —
a different capability axis (extended thinking token budget), not an effort-level selector.
There is no effort divergence to reconcile. The effort capabilities for this model come from the
`anthropic-adaptive-xhigh-opus-4-7` family rule (Anthropic extended-thinking support table).
**Source**: [https://models.dev/api.json](https://models.dev/api.json) — retrieved 2026-07-31; `providers.databricks.models["databricks-claude-opus-4-7"].reasoning_options = [{"type":"budget_tokens","min":1024}]`
**Snapshot**: `scripts/catalog-sample-fixture.json` key `"databricks-claude-opus-4-7"`
---
## Non-divergences (confirmed consistent)
The following models were checked against models.dev or provider docs and found consistent with
the manifest family rules. No exact records needed.
| Model family | Source | Checked against | Status |
|---|---|---|---|
| `claude-opus-4-7` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
| `claude-opus-4-8` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
| `claude-sonnet-5.*` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
| `claude-fable-5` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
| `claude-mythos-5` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
| `claude-opus-4-6` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
| `claude-sonnet-4-6` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
| `claude-mythos-preview` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
| `claude-3*` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
| `gpt-5-pro` | [https://platform.openai.com/docs/guides/reasoning](https://platform.openai.com/docs/guides/reasoning) | OpenAI reasoning guide (July 2025) | ✓ Consistent |
| `gpt-5.6` | [https://platform.openai.com/docs/guides/reasoning](https://platform.openai.com/docs/guides/reasoning) | OpenAI reasoning guide (July 2025) | ✓ Consistent |
| `gpt-5.5` | [https://platform.openai.com/docs/guides/reasoning](https://platform.openai.com/docs/guides/reasoning) | OpenAI reasoning guide (July 2025) | ✓ Consistent |
| `gpt-5.4` | [https://platform.openai.com/docs/guides/reasoning](https://platform.openai.com/docs/guides/reasoning) | OpenAI reasoning guide (July 2025) | ✓ Consistent |
| `gpt-5.1` | [https://platform.openai.com/docs/guides/reasoning](https://platform.openai.com/docs/guides/reasoning) | OpenAI reasoning guide (July 2025) | ✓ Consistent |
| `gpt-5` (base) | [https://platform.openai.com/docs/guides/reasoning](https://platform.openai.com/docs/guides/reasoning) | OpenAI reasoning guide (July 2025) | ✓ Consistent |
+103
View File
@@ -0,0 +1,103 @@
# Model Capabilities Manifest — Schema Reference
**Source of truth**: `scripts/model-capabilities.json`
**Generator**: `scripts/generate-model-capabilities.mjs`
**Emitted artifacts**:
- `crates/buzz-agent/src/generated_model_capabilities.rs`
- `desktop/src/features/agents/ui/modelCapabilities.ts`
- `scripts/generated-model-capabilities-coverage.json` (test fixture)
## Resolver contract (plan v4)
Resolution is a total function `resolve(provider, raw_model_id) → CapabilityResult`.
Three ordered steps:
1. **Provider-qualified raw exact lookup** — key is `(provider, raw_model_id)`, matched on the
RAW ID **before any prefix stripping**. A prefixed alias never inherits an exact record.
2. **Provider-scoped ordered family rules** — on the normalized (prefix-stripped) alias.
Rules ordered by `match_priority` descending (higher wins). Each rule is tagged with the
providers it applies to.
3. **Per-axis provider fallback** — `blank` (empty model string) vs `concrete_unknown`
(nonblank but unmatched), per provider.
`CapabilityResult` is complete — every axis is populated. Runtime consumers never compose fields
from multiple tiers.
## Axes (schema fields)
| Axis | Type | Notes |
|------|------|-------|
| `registry_label` | `string \| null` | Optional static display label. Feeds `resolveModelLabel()` registry tier only. |
| `thinking_mode` | enum | `manual-budget \| adaptive \| omit-fields \| none \| not-applicable` |
| `supported_efforts` | `ThinkingEffort[]` | Non-empty. UI effort dropdown options. |
| `default_effort` | `ThinkingEffort \| null` | null = "Inherit" (Anthropic manual-budget models). |
| `databricks_v2_wire_route` | enum | `openai-responses \| anthropic-messages \| mlflow-chat \| route-unknown \| not-applicable` |
| `normalization_policy` | enum | `none \| openai-standard \| openai-clamp-max-to-xhigh` |
### `thinking_mode` values
| Value | Meaning |
|-------|---------|
| `manual-budget` | `thinking:{type:"enabled", budget_tokens}` — claude-3*, claude-opus-4-5 |
| `adaptive` | `thinking:{type:"adaptive"}` + `output_config:{effort}` — opus-4-6+, sonnet-4-6+, etc. |
| `omit-fields` | Unknown Anthropic model — omit thinking fields rather than guess request shape |
| `none` | Non-Anthropic-routed model — thinking fields not applicable |
| `not-applicable` | Provider does not use Anthropic thinking API |
### `databricks_v2_wire_route` values
Scoped to DBv2 only. All non-DBv2 providers emit `not-applicable`.
Transport for pure OpenAI, legacy Databricks, and OpenRouter is selected by `OpenAiApi` /
`openai_request()` at runtime.
| Value | Meaning |
|-------|---------|
| `openai-responses` | `/ai-gateway/openai/v1/responses` |
| `anthropic-messages` | `/ai-gateway/anthropic/v1/messages` |
| `mlflow-chat` | `/ai-gateway/mlflow/v1/chat/completions` |
| `route-unknown` | DBv2 blank model — route not yet determinable |
| `not-applicable` | Not a DBv2 provider |
## Family rule match kinds
| Kind | Semantics |
|------|-----------|
| `exact` | Case-insensitive exact string equality on normalized alias |
| `prefix` | Normalized alias starts with match_value |
| `gpt5-token` | Boundary-aware token: present at end-of-string or followed by `-` (not digit/letter) |
| `gpt5-base` | Like gpt5-token but also rejects `-<1-3 digit>` suffixes (version-number rejection) |
| `segment` | Normalized alias contains match_value as a full alphanumeric segment (split on non-alnum) |
| `segment-prefix` | Any segment of the normalized alias starts with match_value |
## Boundaries the manifest does NOT own
- **Transport/endpoint selection for pure OpenAI, legacy Databricks, OpenRouter**: `OpenAiApi` and
`openai_request()` remain authoritative. The `databricks_v2_wire_route` axis is DBv2-only.
- **Final display labels**: `resolveModelLabel(discovered_name, registry_label, raw_id)` three-tier
precedence is authoritative. The manifest's `registry_label` feeds only the static registry tier.
- **`llm.rs` replacement scope**: only `databricks_v2_route_for_model`. Other dispatch paths remain.
## Reconciliation policy (plan v4 §Behavior policy)
Not purely behavior-preserving. `models.dev` `reasoning_options` become exact overrides. Each
divergence from family rule results is reconciled against provider docs and either:
- (a) **adopted** as an intentional correction with its own test + exact record, or
- (b) **rejected** with a curation note in the exact record.
See reconciliation table: `scripts/MODELS_DEV_RECONCILIATION.md`.
## Adding a new model family
1. Add a `family_rules` entry with a new unique `id`, appropriate `match_kind`, `providers`,
`match_priority`, and all capability axes.
2. Run `node scripts/generate-model-capabilities.mjs` to regenerate artifacts.
3. CI `model-capability-regen-diff` job verifies byte-clean regeneration.
4. The normative corpus (`scripts/normative-corpus.json`) may need new vectors.
## Adding an exact model override
1. Add an `exact_records` entry with `provider` + `raw_model_id` (the full raw ID, no prefix
stripping). Include a `_reconciliation` note and doc citation.
2. Run `node scripts/generate-model-capabilities.mjs` — completeness validator will fail if any
axis cannot be resolved.
3. Regenerate and commit.
+45
View File
@@ -0,0 +1,45 @@
# Model-Capability Manifest — Mutation Evidence
**Interpreter coverage**: both generated interpreters are exercised per mutation fault.
- **TypeScript**: `scripts/run-corpus.mjs` imports `resolveModelCapabilities()` from
`desktop/src/features/agents/ui/modelCapabilities.ts` via `--experimental-strip-types`.
- **Rust**: `cargo test -p buzz-agent -- generated_model_capabilities::tests::shared_corpus_tests`
deserializes and executes every vector in `scripts/normative-corpus.json` against
`resolve_model_capabilities()`.
## How to reproduce
```sh
# Runs generator mutations; exercises both TS and Rust interpreters per fault
node --experimental-strip-types scripts/run-mutation-evidence.mjs
# Run interpreters independently:
node --experimental-strip-types scripts/run-corpus.mjs
cargo test -p buzz-agent -- generated_model_capabilities::tests::shared_corpus_tests
```
## Mutation run results (both interpreters)
All 7 mutations applied in isolation; manifest restored after each run.
Each mutation must be detected (killed) by **both** interpreters for it to count as covered.
| ID | Mutation | Expected killer(s) | TS | Rust |
|----|----------|--------------------|----|------|
| M1 | Reduce `claude-opus-4-7` `supported_efforts` to `[low,medium,high]` (drops xhigh+max) | `anthropic-claude-opus-4-7`, `dbv2-claude-prefix-stripped`, `dbv2-claude-route-anthropic-messages` | **killed ✓** | **killed ✓** |
| M2 | Add `xhigh` to `gpt5-base` `supported_efforts` | `openai-gpt5-base`, `openai-gpt5-1106-should-not-match-base`, `openai-gpt5-4o-matches-base`, `openai-gpt5-date-suffix` | **killed ✓** | **killed ✓** |
| M3 | Change `gpt5-1` `default_effort` to `"high"` instead of `"none"` | `openai-gpt5.1` | **killed ✓** | **killed ✓** |
| M4 | Swap `dbv2-claude-code-names-segment` route from `anthropic-messages` to `openai-responses` | `dbv2-goose-opus-5-is-anthropic` | **killed ✓** | **killed ✓** |
| M5 | Remove all three DBv2 segment rules | `dbv2-goose-opus-5-is-anthropic`, `dbv2-consolidated-llama-not-sol`, `dbv2-terraform-coder-not-terra` | **killed ✓** | **killed ✓** |
| M6 | Change `databricks_v2` concrete-unknown fallback route from `mlflow-chat` to `openai-responses` | `dbv2-concrete-unknown-mlflow-no-max` | **killed ✓** | **killed ✓** |
| M7 | Remove `xhigh` from `gpt5-4` `supported_efforts` | `resolver-prefixed-alias-misses-exact` | **killed ✓** | **killed ✓** |
**Summary: 7/7 mutations killed in both TS and Rust interpreters.**
## Coverage gaps
- Provider fallback mutations for `anthropic`, `openai`, `databricks`, `openrouter`, and
`_default` are not individually mutated. These are covered by explicit fallback vectors
in the corpus for `anthropic`, `openai`, and `databricks_v2`.
- Rust mutations are run by recompiling the mutated generated file per fault (via `cargo
test` after `node generate-model-capabilities.mjs`). Compile time is acceptable for
offline mutation runs; CI only runs the already-compiled shared corpus harness.
+134
View File
@@ -0,0 +1,134 @@
{
"_comment": "Verbatim models.dev snapshot for differential harness (plan v4 §Oracle). Contains exact records captured from the live API for exact-override entries. Verbatim: name and reasoning_options are reproduced without transformation.",
"_source_url": "https://models.dev/api.json",
"_retrieval_date": "2026-07-31",
"_payload_sha256": "d5a4974cd69f19b0f67713acaa6bb3b16e920defdc07ecbdf6b0a936181bb0e0",
"_retrieval_note": "Full payload SHA-256 computed over the raw response body of GET https://models.dev/api.json (no transforms). Re-verify: curl -s https://models.dev/api.json | sha256sum",
"_models_dev_records": {
"databricks-gpt-5-5": {
"id": "databricks-gpt-5-5",
"name": "GPT-5.5",
"reasoning_options": [
{
"type": "effort",
"values": [
"low",
"medium",
"high"
]
}
]
},
"databricks-gpt-5-4-mini": {
"id": "databricks-gpt-5-4-mini",
"name": "GPT-5.4 mini",
"reasoning_options": [
{
"type": "effort",
"values": [
"low",
"medium",
"high"
]
}
]
},
"databricks-gpt-5-4-nano": {
"id": "databricks-gpt-5-4-nano",
"name": "GPT-5.4 nano",
"reasoning_options": [
{
"type": "effort",
"values": [
"low",
"medium",
"high"
]
}
]
},
"databricks-gpt-5-6-sol": {
"id": "databricks-gpt-5-6-sol",
"name": "GPT-5.6 Sol",
"reasoning_options": [
{
"type": "effort",
"values": [
"low",
"medium",
"high",
"max"
]
}
]
},
"databricks-claude-opus-4-7": {
"id": "databricks-claude-opus-4-7",
"name": "Claude Opus 4.7",
"reasoning_options": [
{
"type": "budget_tokens",
"min": 1024
}
]
}
},
"endpoints": [
{
"name": "databricks-gpt-5-5",
"note": "DATABRICKS_V2_KNOWN_MODELS entry; gpt5-5 family; openai-responses route"
},
{
"name": "databricks-gpt-5-4-mini",
"note": "exact record; models.dev override: low|medium|high (not family rule none+xhigh)"
},
{
"name": "databricks-gpt-5-4-nano",
"note": "exact record; models.dev override: low|medium|high"
},
{
"name": "databricks-gpt-5-6-sol",
"note": "exact record; models.dev source: low|medium|high|max (adopted as-is)"
},
{
"name": "databricks-claude-opus-4-7",
"note": "DATABRICKS_V2_KNOWN_MODELS entry; anthropic adaptive xhigh-capable; anthropic-messages route"
},
{
"name": "goose-claude-fable-5",
"note": "goose- prefix stripped; claude-fable-5 → anthropic adaptive xhigh-capable; anthropic-messages"
},
{
"name": "goose-claude-sonnet-5-20260101",
"note": "goose- prefix stripped; claude-sonnet-5 family; anthropic adaptive xhigh-capable"
},
{
"name": "goose-opus-5",
"note": "'opus' segment → anthropic-messages route; effort: fallback (prefix-stripped alias 'opus-5' not recognized Claude family)"
},
{
"name": "consolidated-llama",
"note": "segment test: 'sol' is substring of 'consolidated', NOT a segment → mlflow-chat"
},
{
"name": "terraform-coder",
"note": "segment test: 'terra' is prefix of 'terraform', NOT a segment → mlflow-chat"
},
{
"name": "corpus-reranker",
"note": "segment test: 'opus' is NOT a segment of 'corpus-reranker' → mlflow-chat"
},
{
"name": "octopus-model",
"note": "segment test: 'opus' is NOT a segment of 'octopus-model' → mlflow-chat"
},
{
"name": "llama-3-70b",
"note": "concrete non-Claude non-GPT → mlflow-chat; effort: all-except-max"
},
{
"name": "",
"note": "blank model → route-unknown; all 7 efforts; default medium"
}
]
}
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+862
View File
@@ -0,0 +1,862 @@
{
"$schema": "./model-capabilities-schema.json",
"_comment": "Hand-curated model capability manifest. Edit here; run scripts/generate-model-capabilities.mjs to regenerate artifacts.",
"_generated_by": "scripts/generate-model-capabilities.mjs",
"_sources": {
"models_dev": "https://models.dev/api.json (retrieved 2026-07-31, SHA-256 d5a4974cd69f19b0f67713acaa6bb3b16e920defdc07ecbdf6b0a936181bb0e0)",
"anthropic_thinking": "https://platform.claude.com/docs/en/build-with-claude/extended-thinking (July 2025)",
"anthropic_effort": "https://platform.claude.com/docs/en/build-with-claude/effort (July 2025)",
"openai_reasoning": "https://platform.openai.com/docs/guides/reasoning (July 2025)",
"goose_known_models": "goose revision 6789d4af (crates/goose-providers/src/databricks_v2.rs:41-42) \u2014 two IDs: databricks-gpt-5-5, databricks-claude-opus-4-7"
},
"family_tokens": [
"claude-",
"gpt-"
],
"family_rules": [
{
"id": "anthropic-manual-budget-claude3",
"match_kind": "prefix",
"match_value": "claude-3",
"providers": [
"anthropic",
"databricks_v2"
],
"match_priority": 10,
"thinking_mode": "manual-budget",
"supported_efforts": [
"low",
"medium",
"high"
],
"default_effort": null,
"databricks_v2_wire_route": "anthropic-messages",
"normalization_policy": "none"
},
{
"id": "anthropic-manual-budget-opus-4-5",
"match_kind": "exact",
"match_value": "claude-opus-4-5",
"providers": [
"anthropic",
"databricks_v2"
],
"match_priority": 10,
"thinking_mode": "manual-budget",
"supported_efforts": [
"low",
"medium",
"high"
],
"default_effort": null,
"databricks_v2_wire_route": "anthropic-messages",
"normalization_policy": "none",
"registry_label": "Claude Opus 4.5"
},
{
"id": "anthropic-adaptive-xhigh-opus-4-7",
"match_kind": "prefix",
"match_value": "claude-opus-4-7",
"providers": [
"anthropic",
"databricks_v2"
],
"match_priority": 10,
"thinking_mode": "adaptive",
"supported_efforts": [
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "high",
"databricks_v2_wire_route": "anthropic-messages",
"normalization_policy": "none",
"registry_label": "Claude Opus 4.7"
},
{
"id": "anthropic-adaptive-xhigh-opus-4-8",
"match_kind": "prefix",
"match_value": "claude-opus-4-8",
"providers": [
"anthropic",
"databricks_v2"
],
"match_priority": 10,
"thinking_mode": "adaptive",
"supported_efforts": [
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "high",
"databricks_v2_wire_route": "anthropic-messages",
"normalization_policy": "none",
"registry_label": "Claude Opus 4.8"
},
{
"id": "anthropic-adaptive-xhigh-opus-5",
"match_kind": "prefix",
"match_value": "claude-opus-5",
"providers": [
"anthropic",
"databricks_v2"
],
"match_priority": 10,
"thinking_mode": "adaptive",
"supported_efforts": [
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "high",
"databricks_v2_wire_route": "anthropic-messages",
"normalization_policy": "none",
"registry_label": "Claude Opus 5"
},
{
"id": "anthropic-adaptive-xhigh-sonnet-5",
"match_kind": "prefix",
"match_value": "claude-sonnet-5",
"providers": [
"anthropic",
"databricks_v2"
],
"match_priority": 10,
"thinking_mode": "adaptive",
"supported_efforts": [
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "high",
"databricks_v2_wire_route": "anthropic-messages",
"normalization_policy": "none",
"registry_label": "Claude Sonnet 5"
},
{
"id": "anthropic-adaptive-xhigh-fable-5",
"match_kind": "prefix",
"match_value": "claude-fable-5",
"providers": [
"anthropic",
"databricks_v2"
],
"match_priority": 10,
"thinking_mode": "adaptive",
"supported_efforts": [
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "high",
"databricks_v2_wire_route": "anthropic-messages",
"normalization_policy": "none",
"registry_label": "Claude Fable 5"
},
{
"id": "anthropic-adaptive-xhigh-mythos-5",
"match_kind": "prefix",
"match_value": "claude-mythos-5",
"providers": [
"anthropic",
"databricks_v2"
],
"match_priority": 10,
"thinking_mode": "adaptive",
"supported_efforts": [
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "high",
"databricks_v2_wire_route": "anthropic-messages",
"normalization_policy": "none",
"registry_label": "Claude Mythos 5"
},
{
"id": "anthropic-adaptive-no-xhigh-opus-4-6",
"match_kind": "prefix",
"match_value": "claude-opus-4-6",
"providers": [
"anthropic",
"databricks_v2"
],
"match_priority": 10,
"thinking_mode": "adaptive",
"supported_efforts": [
"low",
"medium",
"high",
"max"
],
"default_effort": "high",
"databricks_v2_wire_route": "anthropic-messages",
"normalization_policy": "none",
"registry_label": "Claude Opus 4.6"
},
{
"id": "anthropic-adaptive-no-xhigh-sonnet-4-6",
"match_kind": "prefix",
"match_value": "claude-sonnet-4-6",
"providers": [
"anthropic",
"databricks_v2"
],
"match_priority": 10,
"thinking_mode": "adaptive",
"supported_efforts": [
"low",
"medium",
"high",
"max"
],
"default_effort": "high",
"databricks_v2_wire_route": "anthropic-messages",
"normalization_policy": "none",
"registry_label": "Claude Sonnet 4.6"
},
{
"id": "anthropic-adaptive-no-xhigh-mythos-preview",
"match_kind": "prefix",
"match_value": "claude-mythos-preview",
"providers": [
"anthropic",
"databricks_v2"
],
"match_priority": 10,
"thinking_mode": "adaptive",
"supported_efforts": [
"low",
"medium",
"high",
"max"
],
"default_effort": "high",
"databricks_v2_wire_route": "anthropic-messages",
"normalization_policy": "none",
"registry_label": "Claude Mythos Preview"
},
{
"id": "openai-gpt5-pro",
"match_kind": "gpt5-token",
"match_value": "gpt-5-pro",
"match_aliases": [
"gpt5-pro"
],
"providers": [
"openai",
"databricks_v2"
],
"match_priority": 20,
"thinking_mode": "none",
"supported_efforts": [
"high"
],
"default_effort": "high",
"databricks_v2_wire_route": "openai-responses",
"normalization_policy": "openai-standard",
"registry_label": "GPT-5 Pro"
},
{
"id": "openai-gpt5-6",
"match_kind": "gpt5-token",
"match_value": "gpt-5.6",
"match_aliases": [
"gpt5.6",
"gpt-5-6",
"gpt5-6"
],
"providers": [
"openai",
"databricks_v2"
],
"match_priority": 15,
"thinking_mode": "none",
"supported_efforts": [
"none",
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "medium",
"databricks_v2_wire_route": "openai-responses",
"normalization_policy": "openai-standard",
"registry_label": "GPT-5.6"
},
{
"id": "openai-gpt5-5",
"match_kind": "gpt5-token",
"match_value": "gpt-5.5",
"match_aliases": [
"gpt5.5",
"gpt-5-5",
"gpt5-5"
],
"providers": [
"openai",
"databricks_v2"
],
"match_priority": 15,
"thinking_mode": "none",
"supported_efforts": [
"none",
"low",
"medium",
"high",
"xhigh"
],
"default_effort": "medium",
"databricks_v2_wire_route": "openai-responses",
"normalization_policy": "openai-standard",
"registry_label": "GPT-5.5"
},
{
"id": "openai-gpt5-4",
"match_kind": "gpt5-token",
"match_value": "gpt-5.4",
"match_aliases": [
"gpt5.4",
"gpt-5-4",
"gpt5-4"
],
"providers": [
"openai",
"databricks_v2"
],
"match_priority": 15,
"thinking_mode": "none",
"supported_efforts": [
"none",
"low",
"medium",
"high",
"xhigh"
],
"default_effort": "medium",
"databricks_v2_wire_route": "openai-responses",
"normalization_policy": "openai-standard",
"registry_label": "GPT-5.4"
},
{
"id": "openai-gpt5-1",
"match_kind": "gpt5-token",
"match_value": "gpt-5.1",
"match_aliases": [
"gpt5.1",
"gpt-5-1",
"gpt5-1"
],
"providers": [
"openai",
"databricks_v2"
],
"match_priority": 15,
"thinking_mode": "none",
"supported_efforts": [
"none",
"low",
"medium",
"high"
],
"default_effort": "none",
"databricks_v2_wire_route": "openai-responses",
"normalization_policy": "openai-standard",
"registry_label": "GPT-5.1"
},
{
"id": "openai-gpt5-base",
"match_kind": "gpt5-base",
"match_value": "gpt-5",
"match_aliases": [
"gpt5"
],
"providers": [
"openai",
"databricks_v2"
],
"match_priority": 10,
"thinking_mode": "none",
"supported_efforts": [
"minimal",
"low",
"medium",
"high"
],
"default_effort": "medium",
"databricks_v2_wire_route": "openai-responses",
"normalization_policy": "openai-standard",
"registry_label": "GPT-5"
},
{
"id": "dbv2-claude-code-names-segment",
"_comment": "DBv2-only rule: endpoint names containing a Claude code-name segment (opus, sonnet, haiku, mythos, fable, claude) route via Anthropic Messages. This matches goose-opus-5 (segments: goose,opus,5) etc. Effort classification uses conservative defaults because prefix-stripped alias ('opus-5') is not a recognized Claude family.",
"match_kind": "segment",
"match_value": "claude",
"match_aliases": [
"opus",
"sonnet",
"haiku",
"mythos",
"fable"
],
"providers": [
"databricks_v2"
],
"match_priority": 5,
"thinking_mode": "omit-fields",
"supported_efforts": [
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "high",
"databricks_v2_wire_route": "anthropic-messages",
"normalization_policy": "none"
},
{
"id": "dbv2-gpt-code-names-segment",
"_comment": "DBv2-only rule: endpoint names containing a GPT segment prefix (gpt*) route via OpenAI Responses. Handles 'gpt', 'gpt5', 'gpt-5' segments. Priority < individual gpt5 family rules so explicit families take precedence.",
"match_kind": "segment-prefix",
"match_value": "gpt",
"providers": [
"databricks_v2"
],
"match_priority": 5,
"thinking_mode": "none",
"supported_efforts": [
"none",
"minimal",
"low",
"medium",
"high",
"xhigh"
],
"default_effort": "medium",
"databricks_v2_wire_route": "openai-responses",
"normalization_policy": "openai-clamp-max-to-xhigh"
},
{
"id": "dbv2-sol-luna-terra-segment",
"_comment": "DBv2-only rule: sol/luna/terra are OpenAI code names. Route via OpenAI Responses. Must use segment match to avoid matching substrings (consolidated-llama has 'sol' but not as a segment).",
"match_kind": "segment",
"match_value": "sol",
"match_aliases": [
"luna",
"terra"
],
"providers": [
"databricks_v2"
],
"match_priority": 5,
"thinking_mode": "none",
"supported_efforts": [
"none",
"minimal",
"low",
"medium",
"high",
"xhigh"
],
"default_effort": "medium",
"databricks_v2_wire_route": "openai-responses",
"normalization_policy": "openai-clamp-max-to-xhigh"
}
],
"_comment_registry_labels": "All 30 Databricks v2 endpoint-ID to display-name pairs. Represented as [{id,label}] array so duplicate-ID detection is structurally possible. Generated into DATABRICKS_MODEL_NAMES in both Rust and TS.",
"registry_labels": [
{
"id": "databricks-claude-haiku-4-5",
"label": "Claude Haiku 4.5 (latest)"
},
{
"id": "databricks-claude-opus-4-1",
"label": "Claude Opus 4.1 (latest)"
},
{
"id": "databricks-claude-opus-4-5",
"label": "Claude Opus 4.5 (latest)"
},
{
"id": "databricks-claude-opus-4-6",
"label": "Claude Opus 4.6"
},
{
"id": "databricks-claude-opus-4-7",
"label": "Claude Opus 4.7"
},
{
"id": "databricks-claude-sonnet-4",
"label": "Claude Sonnet 4.5"
},
{
"id": "databricks-claude-sonnet-4-5",
"label": "Claude Sonnet 4.5 (latest)"
},
{
"id": "databricks-claude-sonnet-4-6",
"label": "Claude Sonnet 4.6"
},
{
"id": "databricks-gemini-2-5-flash",
"label": "Gemini 2.5 Flash"
},
{
"id": "databricks-gemini-2-5-pro",
"label": "Gemini 2.5 Pro"
},
{
"id": "databricks-gemini-3-1-flash-lite",
"label": "Gemini 3.1 Flash Lite Preview"
},
{
"id": "databricks-gemini-3-1-pro",
"label": "Gemini 3.1 Pro Preview Custom Tools"
},
{
"id": "databricks-gemini-3-flash",
"label": "Gemini 3 Flash Preview"
},
{
"id": "databricks-gemini-3-pro",
"label": "Gemini 3 Pro Preview"
},
{
"id": "databricks-glm-5-2",
"label": "GLM-5.2"
},
{
"id": "databricks-gpt-5",
"label": "GPT-5"
},
{
"id": "databricks-gpt-5-1",
"label": "GPT-5.1"
},
{
"id": "databricks-gpt-5-2",
"label": "GPT-5.2"
},
{
"id": "databricks-gpt-5-4",
"label": "GPT-5.4"
},
{
"id": "databricks-gpt-5-4-mini",
"label": "GPT-5.4 mini"
},
{
"id": "databricks-gpt-5-4-nano",
"label": "GPT-5.4 nano"
},
{
"id": "databricks-gpt-5-5",
"label": "GPT-5.5"
},
{
"id": "databricks-gpt-5-6-luna",
"label": "GPT-5.6 Luna"
},
{
"id": "databricks-gpt-5-6-sol",
"label": "GPT-5.6 Sol"
},
{
"id": "databricks-gpt-5-6-terra",
"label": "GPT-5.6 Terra"
},
{
"id": "databricks-gpt-5-mini",
"label": "GPT-5 Mini"
},
{
"id": "databricks-gpt-5-nano",
"label": "GPT-5 Nano"
},
{
"id": "databricks-gpt-oss-120b",
"label": "GPT OSS 120B"
},
{
"id": "databricks-gpt-oss-20b",
"label": "GPT OSS 20B"
},
{
"id": "databricks-kimi-k2-7-code",
"label": "Kimi K2.7 Code"
}
],
"_comment_databricks_v2_known_models": "Authoritative list of Databricks v2 known model IDs. Mirrors goose DATABRICKS_V2_KNOWN_MODELS at revision 6789d4af (crates/goose-providers/src/databricks_v2.rs:41-42). Generated into DATABRICKS_V2_KNOWN_MODELS in both Rust and TS. Uniqueness enforced by the generator. Opt-in drift check: node scripts/generate-model-capabilities.mjs --check-goose",
"databricks_v2_known_models": [
"databricks-gpt-5-5",
"databricks-claude-opus-4-7"
],
"exact_records": [
{
"provider": "databricks_v2",
"raw_model_id": "databricks-gpt-5-4-mini",
"registry_label": "GPT-5.4 Mini",
"supported_efforts_override": [
"low",
"medium",
"high"
],
"source": "models.dev reasoning_options: low|medium|high (family rule adds none+xhigh \u2014 adopt provider-advertised)",
"_reconciliation": "adopt",
"_reconciliation_note": "models.dev advertises low|medium|high. Family rule (gpt5-4) adds none+xhigh. Provider-advertised wins per plan F1 policy.",
"_reconciliation_doc": "https://models.dev/api.json (retrieved 2026-07-31, SHA-256 d5a4974cd69f19b0f67713acaa6bb3b16e920defdc07ecbdf6b0a936181bb0e0): providers.databricks.models[\"databricks-gpt-5-4-mini\"].reasoning_options=[{\"type\":\"effort\",\"values\":[\"low\",\"medium\",\"high\"]}]"
},
{
"provider": "databricks_v2",
"raw_model_id": "databricks-gpt-5-4-nano",
"registry_label": "GPT-5.4 Nano",
"supported_efforts_override": [
"low",
"medium",
"high"
],
"source": "models.dev reasoning_options: low|medium|high",
"_reconciliation": "adopt",
"_reconciliation_note": "models.dev advertises low|medium|high. Same as gpt-5-4-mini. Adopt.",
"_reconciliation_doc": "https://models.dev/api.json (retrieved 2026-07-31, SHA-256 d5a4974cd69f19b0f67713acaa6bb3b16e920defdc07ecbdf6b0a936181bb0e0): providers.databricks.models[\"databricks-gpt-5-4-nano\"].reasoning_options=[{\"type\":\"effort\",\"values\":[\"low\",\"medium\",\"high\"]}]"
},
{
"provider": "databricks_v2",
"raw_model_id": "databricks-gpt-5-6-sol",
"registry_label": "GPT-5.6 Sol",
"supported_efforts_override": [
"low",
"medium",
"high",
"max"
],
"source": "models.dev reasoning_options: low|medium|high|max (family rule adds none+xhigh \u2014 provider-advertised wins per plan F1)",
"_reconciliation": "adopt",
"_reconciliation_note": "models.dev advertises [low, medium, high, max]. Family rule (gpt5-6) has none+xhigh+max; sol endpoint does not expose none or xhigh. Provider-advertised wins.",
"_reconciliation_doc": "https://models.dev/api.json (retrieved 2026-07-31, SHA-256 d5a4974cd69f19b0f67713acaa6bb3b16e920defdc07ecbdf6b0a936181bb0e0): providers.databricks.models[\"databricks-gpt-5-6-sol\"].reasoning_options=[{\"type\":\"effort\",\"values\":[\"low\",\"medium\",\"high\",\"max\"]}]"
},
{
"provider": "databricks_v2",
"raw_model_id": "databricks-gpt-5-5",
"registry_label": "GPT-5.5",
"source": "models.dev reasoning_options: low|medium|high (family rule adds none+xhigh \u2014 provider-advertised wins per plan F1)",
"_reconciliation": "adopt",
"_reconciliation_note": "models.dev (pinned payload) advertises [low, medium, high]. Family rule (gpt5-5) has none+xhigh; this Databricks endpoint does not expose none or xhigh. Provider-advertised wins.",
"supported_efforts_override": [
"low",
"medium",
"high"
],
"_reconciliation_doc": "https://models.dev/api.json (retrieved 2026-07-31, SHA-256 d5a4974cd69f19b0f67713acaa6bb3b16e920defdc07ecbdf6b0a936181bb0e0): providers.databricks.models[\"databricks-gpt-5-5\"].reasoning_options=[{\"type\":\"effort\",\"values\":[\"low\",\"medium\",\"high\"]}]"
},
{
"provider": "databricks_v2",
"raw_model_id": "databricks-claude-opus-4-7",
"registry_label": "Claude Opus 4.7",
"source": "DATABRICKS_V2_KNOWN_MODELS; family rule anthropic-adaptive-xhigh-opus-4-7 applies",
"_reconciliation": "no-effort-divergence",
"_reconciliation_note": "models.dev advertises reasoning_options=[{\"type\":\"budget_tokens\",\"min\":1024}]. This is a different capability axis (extended thinking token budget), not an effort-level selector. No effort divergence to reconcile \u2014 efforts for this model come from the anthropic family rule (anthropic-adaptive-xhigh-opus-4-7).",
"_reconciliation_doc": "https://models.dev/api.json (retrieved 2026-07-31, SHA-256 d5a4974cd69f19b0f67713acaa6bb3b16e920defdc07ecbdf6b0a936181bb0e0): providers.databricks.models[\"databricks-claude-opus-4-7\"].reasoning_options=[{\"type\":\"budget_tokens\",\"min\":1024}]"
}
],
"provider_fallbacks": {
"anthropic": {
"blank": {
"databricks_v2_wire_route": "not-applicable",
"thinking_mode": "adaptive",
"supported_efforts": [
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "high",
"normalization_policy": "none"
},
"concrete_unknown": {
"databricks_v2_wire_route": "not-applicable",
"thinking_mode": "omit-fields",
"supported_efforts": [
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "high",
"normalization_policy": "none"
}
},
"openai": {
"blank": {
"databricks_v2_wire_route": "not-applicable",
"thinking_mode": "none",
"supported_efforts": [
"none",
"minimal",
"low",
"medium",
"high",
"xhigh"
],
"default_effort": "medium",
"normalization_policy": "openai-clamp-max-to-xhigh"
},
"concrete_unknown": {
"databricks_v2_wire_route": "not-applicable",
"thinking_mode": "none",
"supported_efforts": [
"none",
"minimal",
"low",
"medium",
"high",
"xhigh"
],
"default_effort": "medium",
"normalization_policy": "openai-clamp-max-to-xhigh"
}
},
"databricks_v2": {
"blank": {
"databricks_v2_wire_route": "route-unknown",
"thinking_mode": "none",
"supported_efforts": [
"none",
"minimal",
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "medium",
"normalization_policy": "openai-clamp-max-to-xhigh"
},
"concrete_unknown": {
"databricks_v2_wire_route": "mlflow-chat",
"thinking_mode": "none",
"supported_efforts": [
"none",
"minimal",
"low",
"medium",
"high",
"xhigh"
],
"default_effort": "medium",
"normalization_policy": "openai-clamp-max-to-xhigh"
}
},
"databricks": {
"blank": {
"databricks_v2_wire_route": "not-applicable",
"thinking_mode": "none",
"supported_efforts": [
"none",
"minimal",
"low",
"medium",
"high",
"xhigh"
],
"default_effort": "medium",
"normalization_policy": "openai-clamp-max-to-xhigh"
},
"concrete_unknown": {
"databricks_v2_wire_route": "not-applicable",
"thinking_mode": "none",
"supported_efforts": [
"none",
"minimal",
"low",
"medium",
"high",
"xhigh"
],
"default_effort": "medium",
"normalization_policy": "openai-clamp-max-to-xhigh"
}
},
"openrouter": {
"blank": {
"databricks_v2_wire_route": "not-applicable",
"thinking_mode": "none",
"supported_efforts": [
"none",
"minimal",
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "medium",
"normalization_policy": "none"
},
"concrete_unknown": {
"databricks_v2_wire_route": "not-applicable",
"thinking_mode": "none",
"supported_efforts": [
"none",
"minimal",
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "medium",
"normalization_policy": "none"
}
},
"_default": {
"blank": {
"databricks_v2_wire_route": "not-applicable",
"thinking_mode": "none",
"supported_efforts": [
"none",
"minimal",
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "medium",
"normalization_policy": "none"
},
"concrete_unknown": {
"databricks_v2_wire_route": "not-applicable",
"thinking_mode": "none",
"supported_efforts": [
"none",
"minimal",
"low",
"medium",
"high",
"xhigh",
"max"
],
"default_effort": "medium",
"normalization_policy": "none"
}
}
}
}
+485
View File
@@ -0,0 +1,485 @@
[
{
"_group": "Anthropic exact family rules",
"_note": "All require thinking_mode=manual-budget or adaptive, correct supported_efforts, databricks_v2_wire_route=not-applicable"
},
{
"id": "anthropic-claude-3-family",
"provider": "anthropic",
"raw_model_id": "claude-3-7-sonnet-20250219",
"expect": {
"thinking_mode": "manual-budget",
"supported_efforts": ["low", "medium", "high"],
"default_effort": null,
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "anthropic-claude-opus-4-5",
"provider": "anthropic",
"raw_model_id": "claude-opus-4-5",
"expect": {
"thinking_mode": "manual-budget",
"supported_efforts": ["low", "medium", "high"],
"default_effort": null,
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "anthropic-claude-opus-4-7",
"provider": "anthropic",
"raw_model_id": "claude-opus-4-7",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "anthropic-claude-opus-4-8",
"provider": "anthropic",
"raw_model_id": "claude-opus-4-8",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "anthropic-claude-sonnet-5",
"provider": "anthropic",
"raw_model_id": "claude-sonnet-5-20260101",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "anthropic-claude-fable-5",
"provider": "anthropic",
"raw_model_id": "claude-fable-5",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "anthropic-claude-mythos-5",
"provider": "anthropic",
"raw_model_id": "claude-mythos-5",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "anthropic-claude-opus-4-6",
"provider": "anthropic",
"raw_model_id": "claude-opus-4-6",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "max"],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "anthropic-claude-sonnet-4-6",
"provider": "anthropic",
"raw_model_id": "claude-sonnet-4-6",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "max"],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "anthropic-claude-mythos-preview",
"provider": "anthropic",
"raw_model_id": "claude-mythos-preview",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "max"],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"_group": "Anthropic unknowns and fallbacks"
},
{
"id": "anthropic-unknown-blank",
"provider": "anthropic",
"raw_model_id": "",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "anthropic-unknown-concrete",
"provider": "anthropic",
"raw_model_id": "claude-ultra-9000",
"expect": {
"thinking_mode": "omit-fields",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"_group": "OpenAI exact family rules"
},
{
"id": "openai-gpt5-pro",
"provider": "openai",
"raw_model_id": "gpt-5-pro",
"expect": {
"thinking_mode": "none",
"supported_efforts": ["high"],
"default_effort": "high",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "openai-gpt5.6",
"provider": "openai",
"raw_model_id": "gpt-5.6",
"expect": {
"thinking_mode": "none",
"supported_efforts": ["none", "low", "medium", "high", "xhigh", "max"],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "openai-gpt5-6-dashed",
"provider": "openai",
"raw_model_id": "gpt-5-6",
"expect": {
"thinking_mode": "none",
"supported_efforts": ["none", "low", "medium", "high", "xhigh", "max"],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "openai-gpt5.5",
"provider": "openai",
"raw_model_id": "gpt-5.5",
"expect": {
"thinking_mode": "none",
"supported_efforts": ["none", "low", "medium", "high", "xhigh"],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "openai-gpt5.4",
"provider": "openai",
"raw_model_id": "gpt-5.4",
"expect": {
"thinking_mode": "none",
"supported_efforts": ["none", "low", "medium", "high", "xhigh"],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "openai-gpt5.1",
"provider": "openai",
"raw_model_id": "gpt-5.1",
"expect": {
"thinking_mode": "none",
"supported_efforts": ["none", "low", "medium", "high"],
"default_effort": "none",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "openai-gpt5-base",
"provider": "openai",
"raw_model_id": "gpt-5",
"expect": {
"thinking_mode": "none",
"supported_efforts": ["minimal", "low", "medium", "high"],
"default_effort": "medium",
"databricks_v2_wire_route": "not-applicable"
}
},
{
"_group": "OpenAI adversarial — gpt5 boundary-aware matching (ported from config.rs tests)"
},
{
"id": "openai-gpt5-1106-should-not-match-base",
"provider": "openai",
"raw_model_id": "gpt-5-1106",
"_note": "gpt-5-1106: '-1106' is a 4-digit date segment, NOT a short version (gpt5-base rejects only 1-3 digit suffixes). Must match base table [minimal,low,medium,high], NOT fall through to unknown.",
"expect": {
"supported_efforts": ["minimal", "low", "medium", "high"]
}
},
{
"id": "openai-gpt5-4o-matches-base",
"provider": "openai",
"raw_model_id": "gpt-5-4o",
"_note": "gpt-5-4o: '4o' after '-' is NOT a short numeric suffix (it contains a letter). Must match gpt5-base. Crucially, must NOT match gpt-5.4 (the '4' is followed by 'o', not boundary char).",
"expect": {
"supported_efforts": ["minimal", "low", "medium", "high"]
}
},
{
"id": "openai-gpt5-pro-not-matching-gpt5-base",
"provider": "openai",
"raw_model_id": "gpt-5-pro",
"_note": "gpt-5-pro should hit gpt5-pro rule (priority 20), NOT gpt-5 base.",
"expect": {
"supported_efforts": ["high"],
"default_effort": "high"
}
},
{
"id": "openai-multi-digit-version-gpt5-10",
"provider": "openai",
"raw_model_id": "gpt-5-10",
"_note": "gpt-5-10 — two-digit suffix prevents gpt5-base match. Falls through to unknown.",
"expect": {
"supported_efforts": ["none", "minimal", "low", "medium", "high", "xhigh"]
}
},
{
"id": "openai-gpt5-date-suffix",
"provider": "openai",
"raw_model_id": "gpt-5-20260101",
"_note": "gpt-5-20260101 — long numeric suffix after base: '20260101' is 8 digits, beyond 1-3 digit reject, should hit gpt5-base.",
"expect": {
"supported_efforts": ["minimal", "low", "medium", "high"]
}
},
{
"_group": "DatabricksV2 — segment-based routing (ported from llm.rs tests)"
},
{
"id": "dbv2-gpt5-route-openai-responses",
"provider": "databricks_v2",
"raw_model_id": "gpt-5.5",
"expect": {
"databricks_v2_wire_route": "openai-responses",
"supported_efforts": ["none", "low", "medium", "high", "xhigh"]
}
},
{
"id": "dbv2-claude-route-anthropic-messages",
"provider": "databricks_v2",
"raw_model_id": "claude-opus-4-7",
"expect": {
"databricks_v2_wire_route": "anthropic-messages",
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"]
}
},
{
"id": "dbv2-claude-prefix-stripped",
"provider": "databricks_v2",
"raw_model_id": "databricks-claude-opus-4-7",
"_note": "databricks- prefix stripped → claude-opus-4-7 → Anthropic route",
"expect": {
"databricks_v2_wire_route": "anthropic-messages",
"thinking_mode": "adaptive"
}
},
{
"id": "dbv2-goose-claude-prefix-stripped",
"provider": "databricks_v2",
"raw_model_id": "goose-claude-fable-5",
"_note": "goose- prefix stripped → claude-fable-5 → Anthropic adaptive+xhigh",
"expect": {
"databricks_v2_wire_route": "anthropic-messages",
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"]
}
},
{
"id": "dbv2-team-prefix-stripped",
"provider": "databricks_v2",
"raw_model_id": "team-x-claude-opus-4-7",
"_note": "team-x- prefix stripped → claude-opus-4-7 → Anthropic route",
"expect": {
"databricks_v2_wire_route": "anthropic-messages",
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"]
}
},
{
"id": "dbv2-consolidated-llama-not-sol",
"provider": "databricks_v2",
"raw_model_id": "consolidated-llama",
"_note": "segment test: 'sol' is a SUBSTRING of 'consolidated' — must NOT match DATABRICKS_V2_OPENAI_CODE_NAMES 'sol'. Falls through to mlflow-chat.",
"expect": {
"databricks_v2_wire_route": "mlflow-chat"
}
},
{
"id": "dbv2-terraform-coder-not-terra",
"provider": "databricks_v2",
"raw_model_id": "terraform-coder",
"_note": "segment test: 'terra' is a prefix of 'terraform' — must NOT match 'terra' code name. Falls through to mlflow-chat.",
"expect": {
"databricks_v2_wire_route": "mlflow-chat"
}
},
{
"id": "dbv2-corpus-reranker-not-opus",
"provider": "databricks_v2",
"raw_model_id": "corpus-reranker",
"_note": "segment test: 'opus' is NOT a segment of corpus-reranker (segments: corpus, reranker). mlflow-chat.",
"expect": {
"databricks_v2_wire_route": "mlflow-chat"
}
},
{
"id": "dbv2-octopus-model-not-opus",
"provider": "databricks_v2",
"raw_model_id": "octopus-model",
"_note": "segment test: 'opus' is not a segment of octopus-model (segments: octopus, model). mlflow-chat.",
"expect": {
"databricks_v2_wire_route": "mlflow-chat"
}
},
{
"id": "dbv2-goose-opus-5-is-anthropic",
"provider": "databricks_v2",
"raw_model_id": "goose-opus-5",
"_note": "'opus' IS a named segment of goose-opus-5 (segments: goose, opus, 5). Routes Anthropic. Key test: agrees with llm.rs but disagreed with old config.rs.",
"expect": {
"databricks_v2_wire_route": "anthropic-messages"
}
},
{
"_group": "P2-A resolver-contract vectors (plan v4 §Resolver contract)"
},
{
"id": "resolver-exact-raw-id-hit",
"provider": "databricks_v2",
"raw_model_id": "databricks-gpt-5-4-mini",
"_note": "Exact record exists. Must return exact Databricks override: low|medium|high (not family's none+xhigh).",
"expect": {
"supported_efforts": ["low", "medium", "high"]
}
},
{
"id": "resolver-prefixed-alias-misses-exact",
"provider": "databricks_v2",
"raw_model_id": "team-x-databricks-gpt-5-4-mini",
"_note": "Prefixed alias of an exact ID. Raw exact lookup MUST miss (key is team-x-..., not databricks-...). Falls to family rules (gpt5-4 family → none+xhigh).",
"expect": {
"supported_efforts": ["none", "low", "medium", "high", "xhigh"]
}
},
{
"id": "resolver-cross-provider-misses-exact",
"provider": "openai",
"raw_model_id": "databricks-gpt-5-4-mini",
"_note": "Same raw ID but different provider. Exact record is databricks_v2-scoped; must miss. Falls to openai family rules.",
"expect": {
"databricks_v2_wire_route": "not-applicable"
}
},
{
"id": "resolver-exact-efforts-plus-family-route",
"provider": "databricks_v2",
"raw_model_id": "databricks-gpt-5-6-sol",
"_note": "Exact record with efforts from models.dev (low|medium|high|max — provider-advertised, no none/xhigh). Route materialized from gpt5-6 family rule (openai-responses). Must return both, complete.",
"expect": {
"supported_efforts": ["low", "medium", "high", "max"],
"databricks_v2_wire_route": "openai-responses"
}
},
{
"id": "dbv2-gpt5-5-exact-override",
"provider": "databricks_v2",
"raw_model_id": "databricks-gpt-5-5",
"_note": "Exact record adopts models.dev advertised set [low,medium,high]. Family rule (gpt5-5) has none+xhigh — provider-advertised wins per plan F1.",
"expect": {
"supported_efforts": ["low", "medium", "high"],
"databricks_v2_wire_route": "openai-responses"
}
},
{
"_group": "Blank vs concrete-unknown per provider (P2-B fallback vectors)"
},
{
"id": "dbv2-blank-all7-route-unknown",
"provider": "databricks_v2",
"raw_model_id": "",
"_note": "DBv2 blank: route-unknown, all 7 efforts, default medium.",
"expect": {
"databricks_v2_wire_route": "route-unknown",
"supported_efforts": ["none", "minimal", "low", "medium", "high", "xhigh", "max"],
"default_effort": "medium"
}
},
{
"id": "dbv2-concrete-unknown-mlflow-no-max",
"provider": "databricks_v2",
"raw_model_id": "some-unknown-model-xyz",
"_note": "DBv2 concrete-unknown: mlflow-chat, all-except-max (6 efforts).",
"expect": {
"databricks_v2_wire_route": "mlflow-chat",
"supported_efforts": ["none", "minimal", "low", "medium", "high", "xhigh"]
}
},
{
"id": "openai-blank-all-except-max",
"provider": "openai",
"raw_model_id": "",
"_note": "OpenAI blank: not-applicable route, all-except-max, medium default.",
"expect": {
"databricks_v2_wire_route": "not-applicable",
"supported_efforts": ["none", "minimal", "low", "medium", "high", "xhigh"],
"default_effort": "medium"
}
},
{
"id": "openai-concrete-unknown-all-except-max",
"provider": "openai",
"raw_model_id": "gpt-4o",
"_note": "OpenAI concrete unknown (unverified family): not-applicable route, all-except-max, medium default.",
"expect": {
"databricks_v2_wire_route": "not-applicable",
"supported_efforts": ["none", "minimal", "low", "medium", "high", "xhigh"],
"default_effort": "medium"
}
},
{
"id": "anthropic-blank-adaptive-full",
"provider": "anthropic",
"raw_model_id": "",
"_note": "Anthropic blank: assume adaptive with full support (incl. xhigh).",
"expect": {
"thinking_mode": "adaptive",
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
"default_effort": "high"
}
},
{
"id": "anthropic-concrete-unknown-omit-fields",
"provider": "anthropic",
"raw_model_id": "claude-ultra-9000",
"_note": "Anthropic concrete-unknown: omit-fields (never guess request shape).",
"expect": {
"thinking_mode": "omit-fields"
}
}
]
+83
View File
@@ -0,0 +1,83 @@
#!/usr/bin/env node
/**
* Normative corpus runner — validates the generated TS interpreter against
* scripts/normative-corpus.json.
*
* This is the JS side of the two-interpreter corpus check. The runner imports
* resolveModelCapabilities() directly from the generated TypeScript module via
* Node's --experimental-strip-types flag. The Rust side lives in
* crates/buzz-agent/src/generated_model_capabilities.rs (shared corpus harness).
*
* Usage: node --experimental-strip-types scripts/run-corpus.mjs [--verbose]
* Exits 0 on all pass, 1 on any failure.
*/
import { readFileSync } from "node:fs";
import { join, dirname } from "node:path";
import { fileURLToPath } from "node:url";
const __dirname = dirname(fileURLToPath(import.meta.url));
const repoRoot = join(__dirname, "..");
const VERBOSE = process.argv.includes("--verbose");
// Import the generated TypeScript interpreter directly.
// Node 22+ --experimental-strip-types strips type annotations at load time; no build step needed.
const { resolveModelCapabilities } = await import(
join(repoRoot, "desktop", "src", "features", "agents", "ui", "modelCapabilities.ts")
);
const corpus = JSON.parse(
readFileSync(join(repoRoot, "scripts", "normative-corpus.json"), "utf8"),
);
// ----- Run corpus -----
let passed = 0;
let failed = 0;
for (const entry of corpus) {
// Skip group header entries
if (entry._group) continue;
if (!entry.expect) continue;
// resolveModelCapabilities returns camelCase keys (registryLabel, thinkingMode, etc.)
const result = resolveModelCapabilities(entry.provider, entry.raw_model_id);
const expect = entry.expect;
const failures = [];
for (const [key, expectedVal] of Object.entries(expect)) {
// Corpus uses snake_case; generated TS uses camelCase — convert for lookup.
const camelKey = key.replace(/_([a-z])/g, (_, c) => c.toUpperCase());
const actualVal = camelKey in result ? result[camelKey] : result[key];
if (Array.isArray(expectedVal)) {
// Order-sensitive comparison for effort arrays
const actualArr = Array.isArray(actualVal) ? actualVal : [];
if (JSON.stringify(actualArr) !== JSON.stringify(expectedVal)) {
failures.push(
` ${key}: expected [${expectedVal.join(", ")}] got [${actualArr.join(", ")}]`,
);
}
} else {
if (actualVal !== expectedVal) {
failures.push(` ${key}: expected ${JSON.stringify(expectedVal)} got ${JSON.stringify(actualVal)}`);
}
}
}
if (failures.length === 0) {
passed++;
if (VERBOSE) {
console.log(` PASS ${entry.id}`);
}
} else {
failed++;
console.error(` FAIL ${entry.id} (${entry.provider} / "${entry.raw_model_id}")`);
for (const f of failures) console.error(f);
if (entry._note) console.error(` note: ${entry._note}`);
}
}
console.log(`\nCorpus: ${passed} passed, ${failed} failed`);
process.exit(failed > 0 ? 1 : 0);
+261
View File
@@ -0,0 +1,261 @@
#!/usr/bin/env node
/**
* Per-interpreter mutation evidence runner.
*
* Introduces deliberate resolver faults into the manifest, regenerates artifacts,
* and verifies the shared normative corpus detects every fault in BOTH the generated
* TypeScript interpreter (via --experimental-strip-types import) and the Rust
* interpreter (via cargo test shared-corpus harness). Exits 0 if all mutations are
* killed in both interpreters; exits 1 if any survive.
*
* Usage:
* node --experimental-strip-types scripts/run-mutation-evidence.mjs [--verbose]
*
* This script is non-CI (run manually to generate MUTATION_EVIDENCE.md). It writes
* its findings to stdout in a format suitable for copy-paste into the evidence doc.
*
* Mutations applied (each in isolation, manifest restored after each run):
* M1: Swap anthropic-adaptive-xhigh-opus-4-7 efforts from [low,medium,high,xhigh,max]
* to [low,medium,high] — kills corpus vectors that check xhigh/max.
* M2: Change gpt5-base supported_efforts to include "xhigh" — kills vectors that
* check gpt5-base resolves minimal-only, not xhigh.
* M3: Change openai-gpt5-1 default_effort to "high" instead of "none" — kills
* the gpt5.1 corpus vector that checks default_effort=none.
* M4: Swap databricks_v2_wire_route in dbv2-claude-code-names-segment from
* "anthropic-messages" to "openai-responses" — kills segment-route corpus vectors.
* M5: Remove all three DBv2 segment rules — kills goose-opus-5 and terraform/consolidated
* segment collision vectors.
* M6: Change databricks_v2 concrete_unknown fallback route from "mlflow-chat" to
* "openai-responses" — kills dbv2-concrete-unknown-mlflow-no-max vector.
* M7: Change gpt5-4 supported_efforts to remove "xhigh" — kills
* resolver-prefixed-alias-misses-exact vector.
*/
import { readFileSync, writeFileSync } from "node:fs";
import { join, dirname } from "node:path";
import { fileURLToPath } from "node:url";
import { execSync, spawnSync } from "node:child_process";
const __dirname = dirname(fileURLToPath(import.meta.url));
const repoRoot = join(__dirname, "..");
const VERBOSE = process.argv.includes("--verbose");
const manifestPath = join(repoRoot, "scripts", "model-capabilities.json");
const generatorPath = join(repoRoot, "scripts", "generate-model-capabilities.mjs");
const jsRunnerPath = join(repoRoot, "scripts", "run-corpus.mjs");
const originalManifest = readFileSync(manifestPath, "utf8");
/**
* Run the TS corpus and return:
* { kind: "passed" } — all vectors pass (mutation survived)
* { kind: "killed", output } — nonzero exit AND at least one expectedKiller ID
* appears in stdout/stderr ("FAIL <id>" line)
* { kind: "error", output } — nonzero exit but NO expected corpus output
* (import error, missing file, syntax error, etc.)
*/
function runTsCorpus(expectedKillers) {
let result;
try {
result = spawnSync(
process.execPath,
["--experimental-strip-types", jsRunnerPath],
{ cwd: repoRoot, encoding: "utf8" },
);
} catch (e) {
return { kind: "error", output: `spawn error: ${e.message}` };
}
if (result.status === 0) return { kind: "passed" };
const output = (result.stdout ?? "") + (result.stderr ?? "");
// A genuine corpus kill produces "FAIL <vector-id>" lines.
// An infrastructure failure (import error, syntax error) produces no such lines.
const hasCorpusFailure = expectedKillers.some((id) => output.includes(`FAIL ${id}`));
if (hasCorpusFailure) return { kind: "killed", output };
return { kind: "error", output };
}
/**
* Run the Rust corpus and return:
* { kind: "passed" } — all vectors pass
* { kind: "killed", output } — nonzero exit AND at least one expectedKiller ID
* appears in the panic output ("[<id>]" format)
* { kind: "error", output } — nonzero exit but NO expected corpus output
* (compile error, missing cargo, linker error, etc.)
*/
function runRustCorpus(expectedKillers) {
let result;
try {
result = spawnSync(
"cargo",
[
"test",
"-p", "buzz-agent",
"--",
"generated_model_capabilities::tests::shared_corpus_tests",
"--nocapture",
],
{ cwd: repoRoot, encoding: "utf8", env: { ...process.env, RUST_BACKTRACE: "0" } },
);
} catch (e) {
return { kind: "error", output: `spawn error: ${e.message}` };
}
if (result.status === 0) return { kind: "passed" };
const output = (result.stdout ?? "") + (result.stderr ?? "");
// The Rust corpus runner panics with "[<vector-id>] <axis>: got..." messages.
const hasCorpusFailure = expectedKillers.some((id) => output.includes(`[${id}]`));
if (hasCorpusFailure) return { kind: "killed", output };
return { kind: "error", output };
}
function regen() {
execSync(`"${process.execPath}" "${generatorPath}"`, {
cwd: repoRoot,
stdio: VERBOSE ? "inherit" : "pipe",
});
}
function restore() {
writeFileSync(manifestPath, originalManifest, "utf8");
}
function applyMutation(mutFn) {
const manifest = JSON.parse(originalManifest);
mutFn(manifest);
writeFileSync(manifestPath, JSON.stringify(manifest, null, 2) + "\n", "utf8");
}
const mutations = [
{
id: "M1",
description: "Reduce claude-opus-4-7 supported_efforts to [low,medium,high] (drops xhigh+max)",
expectedKillers: ["anthropic-claude-opus-4-7", "dbv2-claude-prefix-stripped", "dbv2-claude-route-anthropic-messages"],
mutate(manifest) {
const rule = manifest.family_rules.find(r => r.id === "anthropic-adaptive-xhigh-opus-4-7");
rule.supported_efforts = ["low", "medium", "high"];
},
},
{
id: "M2",
description: "Add xhigh to gpt5-base supported_efforts [minimal,low,medium,high,xhigh]",
expectedKillers: ["openai-gpt5-base", "openai-gpt5-1106-should-not-match-base", "openai-gpt5-4o-matches-base", "openai-gpt5-date-suffix"],
mutate(manifest) {
const rule = manifest.family_rules.find(r => r.id === "openai-gpt5-base");
rule.supported_efforts = ["minimal", "low", "medium", "high", "xhigh"];
},
},
{
id: "M3",
description: "Change gpt5-1 default_effort to 'high' instead of 'none'",
expectedKillers: ["openai-gpt5.1"],
mutate(manifest) {
const rule = manifest.family_rules.find(r => r.id === "openai-gpt5-1");
rule.default_effort = "high";
},
},
{
id: "M4",
description: "Swap dbv2-claude-code-names-segment route from anthropic-messages to openai-responses",
expectedKillers: ["dbv2-goose-opus-5-is-anthropic"],
mutate(manifest) {
const rule = manifest.family_rules.find(r => r.id === "dbv2-claude-code-names-segment");
rule.databricks_v2_wire_route = "openai-responses";
},
},
{
id: "M5",
description: "Remove all three DBv2 segment rules (dbv2-claude-code-names-segment, dbv2-gpt-code-names-segment, dbv2-sol-luna-terra-segment)",
expectedKillers: ["dbv2-goose-opus-5-is-anthropic", "dbv2-consolidated-llama-not-sol", "dbv2-terraform-coder-not-terra"],
mutate(manifest) {
manifest.family_rules = manifest.family_rules.filter(
r => !["dbv2-claude-code-names-segment", "dbv2-gpt-code-names-segment", "dbv2-sol-luna-terra-segment"].includes(r.id)
);
},
},
{
id: "M6",
description: "Change databricks_v2 concrete_unknown fallback route from mlflow-chat to openai-responses",
expectedKillers: ["dbv2-concrete-unknown-mlflow-no-max"],
mutate(manifest) {
manifest.provider_fallbacks.databricks_v2.concrete_unknown.databricks_v2_wire_route = "openai-responses";
},
},
{
id: "M7",
description: "Remove xhigh from gpt5-4 supported_efforts [none,low,medium,high]",
expectedKillers: ["resolver-prefixed-alias-misses-exact"],
mutate(manifest) {
const rule = manifest.family_rules.find(r => r.id === "openai-gpt5-4");
rule.supported_efforts = ["none", "low", "medium", "high"];
},
},
];
let allKilled = true;
const results = [];
for (const mut of mutations) {
process.stdout.write(` ${mut.id}: ${mut.description}\n`);
try {
applyMutation(mut.mutate);
regen();
// TS interpreter
process.stdout.write(` TS ... `);
const tsResult = runTsCorpus(mut.expectedKillers);
const tsKilled = tsResult.kind === "killed";
const tsError = tsResult.kind === "error";
if (tsKilled) {
process.stdout.write("killed ✓\n");
} else if (tsError) {
process.stdout.write(`ERROR (infrastructure failure — not a corpus kill)\n`);
if (VERBOSE) process.stdout.write(` ${tsResult.output}\n`);
} else {
process.stdout.write("SURVIVED ✗\n");
if (VERBOSE) process.stdout.write(` ${tsResult.output ?? ""}\n`);
}
// Rust interpreter
process.stdout.write(` Rust... `);
const rustResult = runRustCorpus(mut.expectedKillers);
const rustKilled = rustResult.kind === "killed";
const rustError = rustResult.kind === "error";
if (rustKilled) {
process.stdout.write("killed ✓\n");
} else if (rustError) {
process.stdout.write(`ERROR (infrastructure failure — not a corpus kill)\n`);
if (VERBOSE) process.stdout.write(` ${rustResult.output}\n`);
} else {
process.stdout.write("SURVIVED ✗\n");
if (VERBOSE) process.stdout.write(` ${rustResult.output ?? ""}\n`);
}
const killed = tsKilled && rustKilled;
if (!killed) allKilled = false;
results.push({ ...mut, killed, tsKilled, rustKilled, tsError, rustError });
} catch (e) {
process.stdout.write(` ERROR: ${e.message}\n`);
allKilled = false;
results.push({ ...mut, killed: false, tsKilled: false, rustKilled: false, tsError: true, rustError: true, output: e.message });
} finally {
restore();
regen(); // restore generated files
}
}
console.log("");
const killed = results.filter(r => r.killed).length;
const errored = results.filter(r => r.tsError || r.rustError).length;
console.log(`Mutation results: ${killed}/${results.length} killed (both interpreters)` +
(errored > 0 ? `, ${errored} ERROR (infrastructure failure — see output above)` : ""));
if (!allKilled) {
const hasErrors = results.some(r => r.tsError || r.rustError);
if (hasErrors) {
console.error("ERROR: Infrastructure failures prevented some mutations from being verified as killed.");
console.error(" Run with --verbose to see the full output for ERROR entries.");
}
console.error("ERROR: Some mutations survived — corpus does not kill all resolver faults.");
process.exit(1);
}
console.log("All mutations killed in both TS and Rust interpreters.");
+413
View File
@@ -0,0 +1,413 @@
#!/usr/bin/env node
/**
* Schema-negative validator tests for the manifest generator.
*
* Every validator rule in generate-model-capabilities.mjs must have a
* failing-input test here — if validation is missing, these tests would
* not catch the defect.
*
* Uses Node.js built-in test runner (node --test).
*/
import { test } from "node:test";
import assert from "node:assert/strict";
import { readFileSync, writeFileSync, unlinkSync, mkdirSync, mkdtempSync } from "node:fs";
import { join, dirname } from "node:path";
import { fileURLToPath } from "node:url";
import { tmpdir } from "node:os";
import { execFileSync, spawnSync } from "node:child_process";
const __dirname = dirname(fileURLToPath(import.meta.url));
const repoRoot = join(__dirname, "..");
const manifestPath = join(repoRoot, "scripts", "model-capabilities.json");
const generatorPath = join(repoRoot, "scripts", "generate-model-capabilities.mjs");
/** Load the real manifest so we can mutate copies. */
const BASE_MANIFEST = JSON.parse(readFileSync(manifestPath, "utf8"));
/**
* Run the generator with a mutated manifest, returning { exitCode, stderr, stdout }.
* Writes the mutated manifest to a temp file and overrides the manifest path via env.
*/
function runGeneratorWithManifest(manifestOverride) {
// Write mutated manifest to a temp path
const tmpDir = mkdtempSync(join(tmpdir(), "test-validator-"));
const tmpManifest = join(tmpDir, "model-capabilities.json");
const tmpOutputDir = join(tmpDir, "out");
mkdirSync(tmpOutputDir, { recursive: true });
writeFileSync(tmpManifest, JSON.stringify(manifestOverride));
// Run the generator via node, pointing MANIFEST_PATH env at the temp file
// The generator reads from process.env.MANIFEST_PATH if set (we add this support)
const result = spawnSync(
process.execPath,
[generatorPath, "--manifest-path", tmpManifest, "--output-dir", tmpOutputDir],
{
encoding: "utf8",
env: { ...process.env },
},
);
// Cleanup
try { unlinkSync(tmpManifest); } catch {}
return { exitCode: result.status ?? 1, stderr: result.stderr, stdout: result.stdout };
}
/**
* Assert that the generator REJECTS the given manifest (exits non-zero).
* The optional `expectedMessage` is checked in stderr if provided.
*/
function assertRejects(label, manifest, expectedMessage) {
const { exitCode, stderr, stdout } = runGeneratorWithManifest(manifest);
assert.notEqual(exitCode, 0, `${label}: expected generator to fail but it succeeded.\nstdout: ${stdout}\nstderr: ${stderr}`);
if (expectedMessage) {
const combined = stderr + stdout;
assert.ok(
combined.includes(expectedMessage),
`${label}: expected error message "${expectedMessage}" not found.\nstdout: ${stdout}\nstderr: ${stderr}`,
);
}
}
/** Deep clone the base manifest and apply a mutator function. */
function mutate(fn) {
const clone = JSON.parse(JSON.stringify(BASE_MANIFEST));
fn(clone);
return clone;
}
// ---------------------------------------------------------------------------
// Rule: invalid enum value in family_rule.thinking_mode
// ---------------------------------------------------------------------------
test("schema-negative: invalid thinking_mode in family rule is rejected", () => {
assertRejects(
"invalid thinking_mode",
mutate((m) => {
m.family_rules[0].thinking_mode = "invalid-mode";
}),
"thinking_mode",
);
});
// ---------------------------------------------------------------------------
// Rule: invalid enum value in family_rule.databricks_v2_wire_route
// ---------------------------------------------------------------------------
test("schema-negative: invalid databricks_v2_wire_route in family rule is rejected", () => {
assertRejects(
"invalid databricks_v2_wire_route",
mutate((m) => {
m.family_rules[0].databricks_v2_wire_route = "chat-completions";
}),
"databricks_v2_wire_route",
);
});
// ---------------------------------------------------------------------------
// Rule: invalid enum value in family_rule.supported_efforts[]
// ---------------------------------------------------------------------------
test("schema-negative: invalid effort value in family rule supported_efforts is rejected", () => {
assertRejects(
"invalid supported_efforts value",
mutate((m) => {
m.family_rules[0].supported_efforts = ["low", "ultra-high"];
}),
"supported_efforts",
);
});
// ---------------------------------------------------------------------------
// Rule: empty supported_efforts array in family rule
// ---------------------------------------------------------------------------
test("schema-negative: empty supported_efforts in family rule is rejected", () => {
assertRejects(
"empty supported_efforts",
mutate((m) => {
m.family_rules[0].supported_efforts = [];
}),
"supported_efforts",
);
});
// ---------------------------------------------------------------------------
// Rule: default_effort not in supported_efforts (non-null)
// ---------------------------------------------------------------------------
test("schema-negative: default_effort not in supported_efforts is rejected", () => {
assertRejects(
"default_effort not in supported_efforts",
mutate((m) => {
m.family_rules[0].supported_efforts = ["low", "medium"];
m.family_rules[0].default_effort = "high"; // not in list
}),
"default_effort",
);
});
// ---------------------------------------------------------------------------
// Rule: invalid normalization_policy in family rule
// ---------------------------------------------------------------------------
test("schema-negative: invalid normalization_policy in family rule is rejected", () => {
assertRejects(
"invalid normalization_policy",
mutate((m) => {
m.family_rules[0].normalization_policy = "pass-through-all";
}),
"normalization_policy",
);
});
// ---------------------------------------------------------------------------
// Rule: invalid match_kind in family rule
// ---------------------------------------------------------------------------
test("schema-negative: invalid match_kind in family rule is rejected", () => {
assertRejects(
"invalid match_kind",
mutate((m) => {
m.family_rules[0].match_kind = "regex";
}),
"match_kind",
);
});
// ---------------------------------------------------------------------------
// Rule: duplicate family rule id
// ---------------------------------------------------------------------------
test("schema-negative: duplicate family rule id is rejected", () => {
assertRejects(
"duplicate family rule id",
mutate((m) => {
m.family_rules.push({ ...m.family_rules[0] }); // duplicate id
}),
"duplicate",
);
});
// ---------------------------------------------------------------------------
// Rule: duplicate exact_record (provider, raw_model_id) key
// ---------------------------------------------------------------------------
test("schema-negative: duplicate exact_record key is rejected", () => {
assertRejects(
"duplicate exact_record key",
mutate((m) => {
m.exact_records.push({ ...m.exact_records[0] }); // duplicate
}),
"duplicate",
);
});
// ---------------------------------------------------------------------------
// Rule: exact_record missing provider
// ---------------------------------------------------------------------------
test("schema-negative: exact_record missing provider is rejected", () => {
assertRejects(
"exact_record missing provider",
mutate((m) => {
m.exact_records.push({ raw_model_id: "some-model" });
}),
"provider",
);
});
// ---------------------------------------------------------------------------
// Rule: exact_record missing raw_model_id
// ---------------------------------------------------------------------------
test("schema-negative: exact_record missing raw_model_id is rejected", () => {
assertRejects(
"exact_record missing raw_model_id",
mutate((m) => {
m.exact_records.push({ provider: "databricks_v2" });
}),
"raw_model_id",
);
});
// ---------------------------------------------------------------------------
// Rule: provider fallback record missing blank state
// ---------------------------------------------------------------------------
test("schema-negative: provider fallback missing blank state is rejected", () => {
assertRejects(
"provider fallback missing blank",
mutate((m) => {
delete m.provider_fallbacks.anthropic.blank;
}),
"blank",
);
});
// ---------------------------------------------------------------------------
// Rule: provider fallback record missing concrete_unknown state
// ---------------------------------------------------------------------------
test("schema-negative: provider fallback missing concrete_unknown state is rejected", () => {
assertRejects(
"provider fallback missing concrete_unknown",
mutate((m) => {
delete m.provider_fallbacks.anthropic.concrete_unknown;
}),
"concrete_unknown",
);
});
// ---------------------------------------------------------------------------
// Rule: invalid thinking_mode in provider fallback
// ---------------------------------------------------------------------------
test("schema-negative: invalid thinking_mode in provider fallback is rejected", () => {
assertRejects(
"invalid thinking_mode in fallback",
mutate((m) => {
m.provider_fallbacks.anthropic.blank.thinking_mode = "always-on";
}),
"thinking_mode",
);
});
// ---------------------------------------------------------------------------
// Rule: invalid databricks_v2_wire_route in provider fallback
// ---------------------------------------------------------------------------
test("schema-negative: invalid wire_route in provider fallback is rejected", () => {
assertRejects(
"invalid wire_route in fallback",
mutate((m) => {
m.provider_fallbacks.anthropic.blank.databricks_v2_wire_route = "http-sse";
}),
"databricks_v2_wire_route",
);
});
// ---------------------------------------------------------------------------
// Rule: invalid default_effort in provider fallback (not in supported_efforts)
// ---------------------------------------------------------------------------
test("schema-negative: default_effort not in supported_efforts in fallback is rejected", () => {
assertRejects(
"default_effort not in supported_efforts in fallback",
mutate((m) => {
m.provider_fallbacks.openai.blank.supported_efforts = ["low", "medium"];
m.provider_fallbacks.openai.blank.default_effort = "high"; // not in list
}),
"default_effort",
);
});
// ---------------------------------------------------------------------------
// Rule: family rule missing id
// ---------------------------------------------------------------------------
test("schema-negative: family rule missing id is rejected", () => {
assertRejects(
"family rule missing id",
mutate((m) => {
m.family_rules.push({
match_kind: "prefix",
match_value: "test-",
providers: ["anthropic"],
match_priority: 1,
thinking_mode: "none",
supported_efforts: ["low"],
default_effort: null,
databricks_v2_wire_route: "not-applicable",
normalization_policy: "none",
// id deliberately omitted
});
}),
"id",
);
});
// ---------------------------------------------------------------------------
// Rule: duplicate registry_label IDs
// ---------------------------------------------------------------------------
test("schema-negative: duplicate registry_label ID is rejected", () => {
assertRejects(
"duplicate registry_label ID",
mutate((m) => {
// Array format — duplicate id is structurally detectable
m.registry_labels = [
{ id: "databricks-gpt-5-5", label: "GPT-5.5" },
{ id: "databricks-gpt-5-5", label: "GPT-5.5 duplicate" },
];
}),
"duplicate",
);
});
// ---------------------------------------------------------------------------
// Rule: registry_label entry missing id (empty string)
// ---------------------------------------------------------------------------
test("schema-negative: registry_label entry with empty id is rejected", () => {
assertRejects(
"registry_label empty id",
mutate((m) => {
m.registry_labels = [{ id: "", label: "Some Label" }];
}),
"id",
);
});
// ---------------------------------------------------------------------------
// Rule: registry_label entry with unsafe characters in id
// ---------------------------------------------------------------------------
test("schema-negative: registry_label entry with unsafe id chars is rejected", () => {
assertRejects(
"registry_label unsafe id",
mutate((m) => {
m.registry_labels = [{ id: 'bad"id', label: "Some Label" }];
}),
"unsafe",
);
});
// ---------------------------------------------------------------------------
// Rule: duplicate databricks_v2_known_models IDs
// ---------------------------------------------------------------------------
test("schema-negative: duplicate databricks_v2_known_models ID is rejected", () => {
assertRejects(
"duplicate known model ID",
mutate((m) => {
m.databricks_v2_known_models = ["databricks-gpt-5-5", "databricks-gpt-5-5"];
}),
"duplicate",
);
});
// ---------------------------------------------------------------------------
// Rule: unsafe characters in match_value (family rule)
// ---------------------------------------------------------------------------
test("schema-negative: family rule match_value with unsafe chars is rejected", () => {
assertRejects(
"family rule match_value with backslash",
mutate((m) => {
// Inject a backslash into an existing rule's match_value — would break Rust string literal
const rule = m.family_rules.find((r) => r.id === "anthropic-manual-budget-claude3");
rule.match_value = "claude-3\\evil";
}),
"unsafe",
);
});
// ---------------------------------------------------------------------------
// Rule: unsafe characters in known-model ID
// ---------------------------------------------------------------------------
test("schema-negative: databricks_v2_known_models ID with unsafe chars is rejected", () => {
assertRejects(
"known-model ID with double-quote",
mutate((m) => {
m.databricks_v2_known_models = ['databricks-gpt-5-5', 'bad"id'];
}),
"unsafe",
);
});
// ---------------------------------------------------------------------------
// Rule: unsafe characters in exact_record registry_label
// ---------------------------------------------------------------------------
test("schema-negative: exact_record registry_label with unsafe chars is rejected", () => {
assertRejects(
"exact_record registry_label with backslash",
mutate((m) => {
const rec = m.exact_records.find((r) => r.raw_model_id === "databricks-gpt-5-4-mini");
rec.registry_label = "GPT-5.4 Mini\\injected";
}),
"unsafe",
);
});
console.log("\nSchema-negative validator tests complete.");