mirror of
https://github.com/block/buzz.git
synced 2026-08-18 06:50:31 +02:00
feat(agent): Phase 1 — model-capability manifest, generator, and test oracle (#3821)
## What this does Introduces the model-capability manifest infrastructure (Phase 1 of the Model-Capability Manifest plan v4, Thufir-approved 9/9/9). No consumer cutover — `config.rs`, `llm.rs`, `catalog.rs`, and `buzzAgentConfig.ts` are unchanged. Phase 2 wires them. **Single source of truth** replaces hand-mirrored metadata across four files: ``` scripts/model-capabilities.json → hand-curated manifest scripts/generate-model-capabilities.mjs → emits Rust + TS artifacts crates/buzz-agent/src/generated_model_capabilities.rs desktop/src/features/agents/ui/modelCapabilities.ts ``` ## Resolver contract (plan v4 §Resolver contract) Total function `resolve(provider, raw_model_id) → CapabilityResult`. Three ordered steps: 1. Provider-qualified raw exact lookup — key is `(provider, raw_model_id)`, matched before any prefix stripping. A prefixed alias never inherits an exact record. 2. Provider-scoped ordered family rules — on normalized (prefix-stripped) alias, by `match_priority` desc. 3. Per-axis provider fallback — `blank` vs `concrete_unknown`, per provider. Result is complete — every axis populated, runtime consumers never compose fields. ## Boundaries the manifest does NOT own - Transport for pure OpenAI, legacy Databricks, OpenRouter: `OpenAiApi`/`openai_request()` remain authoritative. `databricks_v2_wire_route` is DBv2-only (all other providers emit `not-applicable`). - Final display labels: `resolveModelLabel()` three-tier precedence unchanged. `registry_label` feeds only the static registry tier. - `llm.rs` scope: only `databricks_v2_route_for_model` (Phase 2). ## Test oracle (three independent layers) 1. Generated full-table coverage — `scripts/generated-model-capabilities-coverage.json`: every manifest entry + provider fallbacks. 2. Hand-authored normative corpus — `scripts/normative-corpus.json` (44 vectors): Anthropic manual-budget/adaptive families, OpenAI gpt-5 adversarial boundary cases, DBv2 segment-routing collision tests, P2-A resolver-contract vectors, P2-B blank/concrete-unknown per provider. Runs against JS resolver (`run-corpus.mjs`) and mirrored in Rust (`generated_model_capabilities_tests.rs`). 3. Schema-negative tests — `scripts/test-manifest-validator.mjs` (17 tests): every validator rule has a failing-input test. ## Reconciliation table `scripts/MODELS_DEV_RECONCILIATION.md` — all models.dev divergences dispositioned. `databricks-gpt-5-4-mini` and `databricks-gpt-5-4-nano` adopt models.dev `[low,medium,high]` (family rule adds `none+xhigh` the endpoint doesn't advertise). ## CI `.github/workflows/model-capability-regen-diff.yml`: triggers on manifest/generator/artifact changes; regenerates and fails if stale; runs JS corpus + schema-negative tests. ## Acceptance criteria (plan v4 Phase 1) - Byte-clean regen: `node scripts/generate-model-capabilities.mjs --check` passes - Rust compiles: `cargo check -p buzz-agent` - TS typechecks: `pnpm tsc --noEmit --strict` - 44/44 normative corpus vectors pass (JS interpreter) - 41/41 Rust corpus tests pass - 17/17 schema-negative tests pass (every validator rule) - Reconciliation table complete with doc citations - No consumer changes (config.rs, llm.rs, catalog.rs, buzzAgentConfig.ts untouched) --------- Signed-off-by: Will Pfleger <pfleger.will@gmail.com> Co-authored-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
This commit is contained in:
co-authored by
npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7
parent
485b36280a
commit
4d47f48143
@@ -0,0 +1,115 @@
|
||||
# models.dev Reasoning Options Reconciliation Table
|
||||
|
||||
**Source queried**: https://models.dev/api.json (2026-07-31)
|
||||
**Payload SHA-256**: `d5a4974cd69f19b0f67713acaa6bb3b16e920defdc07ecbdf6b0a936181bb0e0`
|
||||
**Policy (plan v4 §Behavior policy)**: models.dev `reasoning_options` become exact overrides.
|
||||
Each divergence from the current family rule result is reconciled here: either (a) adopted as an
|
||||
intentional correction or (b) rejected with a curation note.
|
||||
|
||||
**Verbatim source snapshot**: `scripts/catalog-sample-fixture.json` — verbatim `id`, `name`, and
|
||||
nested `reasoning_options` objects captured from the live API without transformation.
|
||||
Re-verify hash: `curl -s https://models.dev/api.json | sha256sum`
|
||||
|
||||
## Divergences
|
||||
|
||||
### `databricks-gpt-5-4-mini`
|
||||
|
||||
| | Current family rule (gpt5-4) | models.dev | Disposition |
|
||||
|---|---|---|---|
|
||||
| `supported_efforts` | `[none, low, medium, high, xhigh]` | `[low, medium, high]` | **ADOPT** |
|
||||
|
||||
**Rationale**: The Databricks AI Gateway v2 endpoint for `databricks-gpt-5-4-mini` explicitly
|
||||
advertises only `[low, medium, high]` in its `reasoning_options`. The family rule's `none` and
|
||||
`xhigh` are derived from the upstream OpenAI GPT-5.4 spec, which this Databricks endpoint does
|
||||
not expose. Provider-advertised wins per plan F1 policy.
|
||||
|
||||
**Source**: [https://models.dev/api.json](https://models.dev/api.json) — retrieved 2026-07-31; `providers.databricks.models["databricks-gpt-5-4-mini"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]`
|
||||
**Snapshot**: `scripts/catalog-sample-fixture.json` key `"databricks-gpt-5-4-mini"`
|
||||
**Test vector**: `resolver-exact-raw-id-hit` in `scripts/normative-corpus.json`
|
||||
|
||||
---
|
||||
|
||||
### `databricks-gpt-5-4-nano`
|
||||
|
||||
| | Current family rule (gpt5-4) | models.dev | Disposition |
|
||||
|---|---|---|---|
|
||||
| `supported_efforts` | `[none, low, medium, high, xhigh]` | `[low, medium, high]` | **ADOPT** |
|
||||
|
||||
**Rationale**: Same as `databricks-gpt-5-4-mini`. The nano variant exposes the same restricted
|
||||
effort set. Provider-advertised wins.
|
||||
|
||||
**Source**: [https://models.dev/api.json](https://models.dev/api.json) — retrieved 2026-07-31; `providers.databricks.models["databricks-gpt-5-4-nano"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]`
|
||||
**Snapshot**: `scripts/catalog-sample-fixture.json` key `"databricks-gpt-5-4-nano"`
|
||||
|
||||
---
|
||||
|
||||
### `databricks-gpt-5-6-sol`
|
||||
|
||||
| | Current family rule (gpt5-6) | models.dev | Disposition |
|
||||
|---|---|---|---|
|
||||
| `supported_efforts` | `[none, low, medium, high, xhigh, max]` | `[low, medium, high, max]` | **ADOPT** |
|
||||
|
||||
**Rationale**: The Databricks AI Gateway v2 endpoint for `databricks-gpt-5-6-sol` advertises only
|
||||
`[low, medium, high, max]` in its `reasoning_options`. The family rule's `none` and `xhigh` are
|
||||
derived from the upstream OpenAI GPT-5.6 spec, which this Databricks endpoint does not expose.
|
||||
Provider-advertised wins per plan F1 policy.
|
||||
|
||||
**Source**: [https://models.dev/api.json](https://models.dev/api.json) — retrieved 2026-07-31; `providers.databricks.models["databricks-gpt-5-6-sol"].reasoning_options = [{"type":"effort","values":["low","medium","high","max"]}]`
|
||||
**Snapshot**: `scripts/catalog-sample-fixture.json` key `"databricks-gpt-5-6-sol"`
|
||||
|
||||
---
|
||||
|
||||
### `databricks-gpt-5-5`
|
||||
|
||||
| | Current family rule (gpt5-5) | models.dev | Disposition |
|
||||
|---|---|---|---|
|
||||
| `supported_efforts` | `[none, low, medium, high, xhigh]` | `[low, medium, high]` | **ADOPT** |
|
||||
|
||||
**Rationale**: The Databricks AI Gateway v2 endpoint for `databricks-gpt-5-5` advertises only
|
||||
`[low, medium, high]` in its `reasoning_options`. The family rule's `none` and `xhigh` are
|
||||
derived from the upstream OpenAI GPT-5.5 spec, which this Databricks endpoint does not expose.
|
||||
Provider-advertised wins per plan F1 policy.
|
||||
|
||||
**Source**: [https://models.dev/api.json](https://models.dev/api.json) — retrieved 2026-07-31; `providers.databricks.models["databricks-gpt-5-5"].reasoning_options = [{"type":"effort","values":["low","medium","high"]}]`
|
||||
**Snapshot**: `scripts/catalog-sample-fixture.json` key `"databricks-gpt-5-5"`
|
||||
|
||||
---
|
||||
|
||||
### `databricks-claude-opus-4-7`
|
||||
|
||||
| | Current family rule (anthropic-adaptive-xhigh-opus-4-7) | models.dev | Disposition |
|
||||
|---|---|---|---|
|
||||
| `reasoning_options` type | effort-based | `budget_tokens` | **NO EFFORT DIVERGENCE** |
|
||||
|
||||
**Rationale**: models.dev advertises `reasoning_options=[{"type":"budget_tokens","min":1024}]` —
|
||||
a different capability axis (extended thinking token budget), not an effort-level selector.
|
||||
There is no effort divergence to reconcile. The effort capabilities for this model come from the
|
||||
`anthropic-adaptive-xhigh-opus-4-7` family rule (Anthropic extended-thinking support table).
|
||||
|
||||
**Source**: [https://models.dev/api.json](https://models.dev/api.json) — retrieved 2026-07-31; `providers.databricks.models["databricks-claude-opus-4-7"].reasoning_options = [{"type":"budget_tokens","min":1024}]`
|
||||
**Snapshot**: `scripts/catalog-sample-fixture.json` key `"databricks-claude-opus-4-7"`
|
||||
|
||||
---
|
||||
|
||||
## Non-divergences (confirmed consistent)
|
||||
|
||||
The following models were checked against models.dev or provider docs and found consistent with
|
||||
the manifest family rules. No exact records needed.
|
||||
|
||||
| Model family | Source | Checked against | Status |
|
||||
|---|---|---|---|
|
||||
| `claude-opus-4-7` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
|
||||
| `claude-opus-4-8` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
|
||||
| `claude-sonnet-5.*` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
|
||||
| `claude-fable-5` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
|
||||
| `claude-mythos-5` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
|
||||
| `claude-opus-4-6` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
|
||||
| `claude-sonnet-4-6` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
|
||||
| `claude-mythos-preview` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
|
||||
| `claude-3*` | [https://platform.claude.com/docs/en/build-with-claude/extended-thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) | Anthropic extended-thinking support table (July 2025) | ✓ Consistent |
|
||||
| `gpt-5-pro` | [https://platform.openai.com/docs/guides/reasoning](https://platform.openai.com/docs/guides/reasoning) | OpenAI reasoning guide (July 2025) | ✓ Consistent |
|
||||
| `gpt-5.6` | [https://platform.openai.com/docs/guides/reasoning](https://platform.openai.com/docs/guides/reasoning) | OpenAI reasoning guide (July 2025) | ✓ Consistent |
|
||||
| `gpt-5.5` | [https://platform.openai.com/docs/guides/reasoning](https://platform.openai.com/docs/guides/reasoning) | OpenAI reasoning guide (July 2025) | ✓ Consistent |
|
||||
| `gpt-5.4` | [https://platform.openai.com/docs/guides/reasoning](https://platform.openai.com/docs/guides/reasoning) | OpenAI reasoning guide (July 2025) | ✓ Consistent |
|
||||
| `gpt-5.1` | [https://platform.openai.com/docs/guides/reasoning](https://platform.openai.com/docs/guides/reasoning) | OpenAI reasoning guide (July 2025) | ✓ Consistent |
|
||||
| `gpt-5` (base) | [https://platform.openai.com/docs/guides/reasoning](https://platform.openai.com/docs/guides/reasoning) | OpenAI reasoning guide (July 2025) | ✓ Consistent |
|
||||
@@ -0,0 +1,103 @@
|
||||
# Model Capabilities Manifest — Schema Reference
|
||||
|
||||
**Source of truth**: `scripts/model-capabilities.json`
|
||||
**Generator**: `scripts/generate-model-capabilities.mjs`
|
||||
**Emitted artifacts**:
|
||||
- `crates/buzz-agent/src/generated_model_capabilities.rs`
|
||||
- `desktop/src/features/agents/ui/modelCapabilities.ts`
|
||||
- `scripts/generated-model-capabilities-coverage.json` (test fixture)
|
||||
|
||||
## Resolver contract (plan v4)
|
||||
|
||||
Resolution is a total function `resolve(provider, raw_model_id) → CapabilityResult`.
|
||||
Three ordered steps:
|
||||
|
||||
1. **Provider-qualified raw exact lookup** — key is `(provider, raw_model_id)`, matched on the
|
||||
RAW ID **before any prefix stripping**. A prefixed alias never inherits an exact record.
|
||||
2. **Provider-scoped ordered family rules** — on the normalized (prefix-stripped) alias.
|
||||
Rules ordered by `match_priority` descending (higher wins). Each rule is tagged with the
|
||||
providers it applies to.
|
||||
3. **Per-axis provider fallback** — `blank` (empty model string) vs `concrete_unknown`
|
||||
(nonblank but unmatched), per provider.
|
||||
|
||||
`CapabilityResult` is complete — every axis is populated. Runtime consumers never compose fields
|
||||
from multiple tiers.
|
||||
|
||||
## Axes (schema fields)
|
||||
|
||||
| Axis | Type | Notes |
|
||||
|------|------|-------|
|
||||
| `registry_label` | `string \| null` | Optional static display label. Feeds `resolveModelLabel()` registry tier only. |
|
||||
| `thinking_mode` | enum | `manual-budget \| adaptive \| omit-fields \| none \| not-applicable` |
|
||||
| `supported_efforts` | `ThinkingEffort[]` | Non-empty. UI effort dropdown options. |
|
||||
| `default_effort` | `ThinkingEffort \| null` | null = "Inherit" (Anthropic manual-budget models). |
|
||||
| `databricks_v2_wire_route` | enum | `openai-responses \| anthropic-messages \| mlflow-chat \| route-unknown \| not-applicable` |
|
||||
| `normalization_policy` | enum | `none \| openai-standard \| openai-clamp-max-to-xhigh` |
|
||||
|
||||
### `thinking_mode` values
|
||||
|
||||
| Value | Meaning |
|
||||
|-------|---------|
|
||||
| `manual-budget` | `thinking:{type:"enabled", budget_tokens}` — claude-3*, claude-opus-4-5 |
|
||||
| `adaptive` | `thinking:{type:"adaptive"}` + `output_config:{effort}` — opus-4-6+, sonnet-4-6+, etc. |
|
||||
| `omit-fields` | Unknown Anthropic model — omit thinking fields rather than guess request shape |
|
||||
| `none` | Non-Anthropic-routed model — thinking fields not applicable |
|
||||
| `not-applicable` | Provider does not use Anthropic thinking API |
|
||||
|
||||
### `databricks_v2_wire_route` values
|
||||
|
||||
Scoped to DBv2 only. All non-DBv2 providers emit `not-applicable`.
|
||||
Transport for pure OpenAI, legacy Databricks, and OpenRouter is selected by `OpenAiApi` /
|
||||
`openai_request()` at runtime.
|
||||
|
||||
| Value | Meaning |
|
||||
|-------|---------|
|
||||
| `openai-responses` | `/ai-gateway/openai/v1/responses` |
|
||||
| `anthropic-messages` | `/ai-gateway/anthropic/v1/messages` |
|
||||
| `mlflow-chat` | `/ai-gateway/mlflow/v1/chat/completions` |
|
||||
| `route-unknown` | DBv2 blank model — route not yet determinable |
|
||||
| `not-applicable` | Not a DBv2 provider |
|
||||
|
||||
## Family rule match kinds
|
||||
|
||||
| Kind | Semantics |
|
||||
|------|-----------|
|
||||
| `exact` | Case-insensitive exact string equality on normalized alias |
|
||||
| `prefix` | Normalized alias starts with match_value |
|
||||
| `gpt5-token` | Boundary-aware token: present at end-of-string or followed by `-` (not digit/letter) |
|
||||
| `gpt5-base` | Like gpt5-token but also rejects `-<1-3 digit>` suffixes (version-number rejection) |
|
||||
| `segment` | Normalized alias contains match_value as a full alphanumeric segment (split on non-alnum) |
|
||||
| `segment-prefix` | Any segment of the normalized alias starts with match_value |
|
||||
|
||||
## Boundaries the manifest does NOT own
|
||||
|
||||
- **Transport/endpoint selection for pure OpenAI, legacy Databricks, OpenRouter**: `OpenAiApi` and
|
||||
`openai_request()` remain authoritative. The `databricks_v2_wire_route` axis is DBv2-only.
|
||||
- **Final display labels**: `resolveModelLabel(discovered_name, registry_label, raw_id)` three-tier
|
||||
precedence is authoritative. The manifest's `registry_label` feeds only the static registry tier.
|
||||
- **`llm.rs` replacement scope**: only `databricks_v2_route_for_model`. Other dispatch paths remain.
|
||||
|
||||
## Reconciliation policy (plan v4 §Behavior policy)
|
||||
|
||||
Not purely behavior-preserving. `models.dev` `reasoning_options` become exact overrides. Each
|
||||
divergence from family rule results is reconciled against provider docs and either:
|
||||
- (a) **adopted** as an intentional correction with its own test + exact record, or
|
||||
- (b) **rejected** with a curation note in the exact record.
|
||||
|
||||
See reconciliation table: `scripts/MODELS_DEV_RECONCILIATION.md`.
|
||||
|
||||
## Adding a new model family
|
||||
|
||||
1. Add a `family_rules` entry with a new unique `id`, appropriate `match_kind`, `providers`,
|
||||
`match_priority`, and all capability axes.
|
||||
2. Run `node scripts/generate-model-capabilities.mjs` to regenerate artifacts.
|
||||
3. CI `model-capability-regen-diff` job verifies byte-clean regeneration.
|
||||
4. The normative corpus (`scripts/normative-corpus.json`) may need new vectors.
|
||||
|
||||
## Adding an exact model override
|
||||
|
||||
1. Add an `exact_records` entry with `provider` + `raw_model_id` (the full raw ID, no prefix
|
||||
stripping). Include a `_reconciliation` note and doc citation.
|
||||
2. Run `node scripts/generate-model-capabilities.mjs` — completeness validator will fail if any
|
||||
axis cannot be resolved.
|
||||
3. Regenerate and commit.
|
||||
@@ -0,0 +1,45 @@
|
||||
# Model-Capability Manifest — Mutation Evidence
|
||||
|
||||
**Interpreter coverage**: both generated interpreters are exercised per mutation fault.
|
||||
- **TypeScript**: `scripts/run-corpus.mjs` imports `resolveModelCapabilities()` from
|
||||
`desktop/src/features/agents/ui/modelCapabilities.ts` via `--experimental-strip-types`.
|
||||
- **Rust**: `cargo test -p buzz-agent -- generated_model_capabilities::tests::shared_corpus_tests`
|
||||
deserializes and executes every vector in `scripts/normative-corpus.json` against
|
||||
`resolve_model_capabilities()`.
|
||||
|
||||
## How to reproduce
|
||||
|
||||
```sh
|
||||
# Runs generator mutations; exercises both TS and Rust interpreters per fault
|
||||
node --experimental-strip-types scripts/run-mutation-evidence.mjs
|
||||
|
||||
# Run interpreters independently:
|
||||
node --experimental-strip-types scripts/run-corpus.mjs
|
||||
cargo test -p buzz-agent -- generated_model_capabilities::tests::shared_corpus_tests
|
||||
```
|
||||
|
||||
## Mutation run results (both interpreters)
|
||||
|
||||
All 7 mutations applied in isolation; manifest restored after each run.
|
||||
Each mutation must be detected (killed) by **both** interpreters for it to count as covered.
|
||||
|
||||
| ID | Mutation | Expected killer(s) | TS | Rust |
|
||||
|----|----------|--------------------|----|------|
|
||||
| M1 | Reduce `claude-opus-4-7` `supported_efforts` to `[low,medium,high]` (drops xhigh+max) | `anthropic-claude-opus-4-7`, `dbv2-claude-prefix-stripped`, `dbv2-claude-route-anthropic-messages` | **killed ✓** | **killed ✓** |
|
||||
| M2 | Add `xhigh` to `gpt5-base` `supported_efforts` | `openai-gpt5-base`, `openai-gpt5-1106-should-not-match-base`, `openai-gpt5-4o-matches-base`, `openai-gpt5-date-suffix` | **killed ✓** | **killed ✓** |
|
||||
| M3 | Change `gpt5-1` `default_effort` to `"high"` instead of `"none"` | `openai-gpt5.1` | **killed ✓** | **killed ✓** |
|
||||
| M4 | Swap `dbv2-claude-code-names-segment` route from `anthropic-messages` to `openai-responses` | `dbv2-goose-opus-5-is-anthropic` | **killed ✓** | **killed ✓** |
|
||||
| M5 | Remove all three DBv2 segment rules | `dbv2-goose-opus-5-is-anthropic`, `dbv2-consolidated-llama-not-sol`, `dbv2-terraform-coder-not-terra` | **killed ✓** | **killed ✓** |
|
||||
| M6 | Change `databricks_v2` concrete-unknown fallback route from `mlflow-chat` to `openai-responses` | `dbv2-concrete-unknown-mlflow-no-max` | **killed ✓** | **killed ✓** |
|
||||
| M7 | Remove `xhigh` from `gpt5-4` `supported_efforts` | `resolver-prefixed-alias-misses-exact` | **killed ✓** | **killed ✓** |
|
||||
|
||||
**Summary: 7/7 mutations killed in both TS and Rust interpreters.**
|
||||
|
||||
## Coverage gaps
|
||||
|
||||
- Provider fallback mutations for `anthropic`, `openai`, `databricks`, `openrouter`, and
|
||||
`_default` are not individually mutated. These are covered by explicit fallback vectors
|
||||
in the corpus for `anthropic`, `openai`, and `databricks_v2`.
|
||||
- Rust mutations are run by recompiling the mutated generated file per fault (via `cargo
|
||||
test` after `node generate-model-capabilities.mjs`). Compile time is acceptable for
|
||||
offline mutation runs; CI only runs the already-compiled shared corpus harness.
|
||||
@@ -0,0 +1,134 @@
|
||||
{
|
||||
"_comment": "Verbatim models.dev snapshot for differential harness (plan v4 §Oracle). Contains exact records captured from the live API for exact-override entries. Verbatim: name and reasoning_options are reproduced without transformation.",
|
||||
"_source_url": "https://models.dev/api.json",
|
||||
"_retrieval_date": "2026-07-31",
|
||||
"_payload_sha256": "d5a4974cd69f19b0f67713acaa6bb3b16e920defdc07ecbdf6b0a936181bb0e0",
|
||||
"_retrieval_note": "Full payload SHA-256 computed over the raw response body of GET https://models.dev/api.json (no transforms). Re-verify: curl -s https://models.dev/api.json | sha256sum",
|
||||
"_models_dev_records": {
|
||||
"databricks-gpt-5-5": {
|
||||
"id": "databricks-gpt-5-5",
|
||||
"name": "GPT-5.5",
|
||||
"reasoning_options": [
|
||||
{
|
||||
"type": "effort",
|
||||
"values": [
|
||||
"low",
|
||||
"medium",
|
||||
"high"
|
||||
]
|
||||
}
|
||||
]
|
||||
},
|
||||
"databricks-gpt-5-4-mini": {
|
||||
"id": "databricks-gpt-5-4-mini",
|
||||
"name": "GPT-5.4 mini",
|
||||
"reasoning_options": [
|
||||
{
|
||||
"type": "effort",
|
||||
"values": [
|
||||
"low",
|
||||
"medium",
|
||||
"high"
|
||||
]
|
||||
}
|
||||
]
|
||||
},
|
||||
"databricks-gpt-5-4-nano": {
|
||||
"id": "databricks-gpt-5-4-nano",
|
||||
"name": "GPT-5.4 nano",
|
||||
"reasoning_options": [
|
||||
{
|
||||
"type": "effort",
|
||||
"values": [
|
||||
"low",
|
||||
"medium",
|
||||
"high"
|
||||
]
|
||||
}
|
||||
]
|
||||
},
|
||||
"databricks-gpt-5-6-sol": {
|
||||
"id": "databricks-gpt-5-6-sol",
|
||||
"name": "GPT-5.6 Sol",
|
||||
"reasoning_options": [
|
||||
{
|
||||
"type": "effort",
|
||||
"values": [
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"max"
|
||||
]
|
||||
}
|
||||
]
|
||||
},
|
||||
"databricks-claude-opus-4-7": {
|
||||
"id": "databricks-claude-opus-4-7",
|
||||
"name": "Claude Opus 4.7",
|
||||
"reasoning_options": [
|
||||
{
|
||||
"type": "budget_tokens",
|
||||
"min": 1024
|
||||
}
|
||||
]
|
||||
}
|
||||
},
|
||||
"endpoints": [
|
||||
{
|
||||
"name": "databricks-gpt-5-5",
|
||||
"note": "DATABRICKS_V2_KNOWN_MODELS entry; gpt5-5 family; openai-responses route"
|
||||
},
|
||||
{
|
||||
"name": "databricks-gpt-5-4-mini",
|
||||
"note": "exact record; models.dev override: low|medium|high (not family rule none+xhigh)"
|
||||
},
|
||||
{
|
||||
"name": "databricks-gpt-5-4-nano",
|
||||
"note": "exact record; models.dev override: low|medium|high"
|
||||
},
|
||||
{
|
||||
"name": "databricks-gpt-5-6-sol",
|
||||
"note": "exact record; models.dev source: low|medium|high|max (adopted as-is)"
|
||||
},
|
||||
{
|
||||
"name": "databricks-claude-opus-4-7",
|
||||
"note": "DATABRICKS_V2_KNOWN_MODELS entry; anthropic adaptive xhigh-capable; anthropic-messages route"
|
||||
},
|
||||
{
|
||||
"name": "goose-claude-fable-5",
|
||||
"note": "goose- prefix stripped; claude-fable-5 → anthropic adaptive xhigh-capable; anthropic-messages"
|
||||
},
|
||||
{
|
||||
"name": "goose-claude-sonnet-5-20260101",
|
||||
"note": "goose- prefix stripped; claude-sonnet-5 family; anthropic adaptive xhigh-capable"
|
||||
},
|
||||
{
|
||||
"name": "goose-opus-5",
|
||||
"note": "'opus' segment → anthropic-messages route; effort: fallback (prefix-stripped alias 'opus-5' not recognized Claude family)"
|
||||
},
|
||||
{
|
||||
"name": "consolidated-llama",
|
||||
"note": "segment test: 'sol' is substring of 'consolidated', NOT a segment → mlflow-chat"
|
||||
},
|
||||
{
|
||||
"name": "terraform-coder",
|
||||
"note": "segment test: 'terra' is prefix of 'terraform', NOT a segment → mlflow-chat"
|
||||
},
|
||||
{
|
||||
"name": "corpus-reranker",
|
||||
"note": "segment test: 'opus' is NOT a segment of 'corpus-reranker' → mlflow-chat"
|
||||
},
|
||||
{
|
||||
"name": "octopus-model",
|
||||
"note": "segment test: 'opus' is NOT a segment of 'octopus-model' → mlflow-chat"
|
||||
},
|
||||
{
|
||||
"name": "llama-3-70b",
|
||||
"note": "concrete non-Claude non-GPT → mlflow-chat; effort: all-except-max"
|
||||
},
|
||||
{
|
||||
"name": "",
|
||||
"note": "blank model → route-unknown; all 7 efforts; default medium"
|
||||
}
|
||||
]
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,862 @@
|
||||
{
|
||||
"$schema": "./model-capabilities-schema.json",
|
||||
"_comment": "Hand-curated model capability manifest. Edit here; run scripts/generate-model-capabilities.mjs to regenerate artifacts.",
|
||||
"_generated_by": "scripts/generate-model-capabilities.mjs",
|
||||
"_sources": {
|
||||
"models_dev": "https://models.dev/api.json (retrieved 2026-07-31, SHA-256 d5a4974cd69f19b0f67713acaa6bb3b16e920defdc07ecbdf6b0a936181bb0e0)",
|
||||
"anthropic_thinking": "https://platform.claude.com/docs/en/build-with-claude/extended-thinking (July 2025)",
|
||||
"anthropic_effort": "https://platform.claude.com/docs/en/build-with-claude/effort (July 2025)",
|
||||
"openai_reasoning": "https://platform.openai.com/docs/guides/reasoning (July 2025)",
|
||||
"goose_known_models": "goose revision 6789d4af (crates/goose-providers/src/databricks_v2.rs:41-42) \u2014 two IDs: databricks-gpt-5-5, databricks-claude-opus-4-7"
|
||||
},
|
||||
"family_tokens": [
|
||||
"claude-",
|
||||
"gpt-"
|
||||
],
|
||||
"family_rules": [
|
||||
{
|
||||
"id": "anthropic-manual-budget-claude3",
|
||||
"match_kind": "prefix",
|
||||
"match_value": "claude-3",
|
||||
"providers": [
|
||||
"anthropic",
|
||||
"databricks_v2"
|
||||
],
|
||||
"match_priority": 10,
|
||||
"thinking_mode": "manual-budget",
|
||||
"supported_efforts": [
|
||||
"low",
|
||||
"medium",
|
||||
"high"
|
||||
],
|
||||
"default_effort": null,
|
||||
"databricks_v2_wire_route": "anthropic-messages",
|
||||
"normalization_policy": "none"
|
||||
},
|
||||
{
|
||||
"id": "anthropic-manual-budget-opus-4-5",
|
||||
"match_kind": "exact",
|
||||
"match_value": "claude-opus-4-5",
|
||||
"providers": [
|
||||
"anthropic",
|
||||
"databricks_v2"
|
||||
],
|
||||
"match_priority": 10,
|
||||
"thinking_mode": "manual-budget",
|
||||
"supported_efforts": [
|
||||
"low",
|
||||
"medium",
|
||||
"high"
|
||||
],
|
||||
"default_effort": null,
|
||||
"databricks_v2_wire_route": "anthropic-messages",
|
||||
"normalization_policy": "none",
|
||||
"registry_label": "Claude Opus 4.5"
|
||||
},
|
||||
{
|
||||
"id": "anthropic-adaptive-xhigh-opus-4-7",
|
||||
"match_kind": "prefix",
|
||||
"match_value": "claude-opus-4-7",
|
||||
"providers": [
|
||||
"anthropic",
|
||||
"databricks_v2"
|
||||
],
|
||||
"match_priority": 10,
|
||||
"thinking_mode": "adaptive",
|
||||
"supported_efforts": [
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"xhigh",
|
||||
"max"
|
||||
],
|
||||
"default_effort": "high",
|
||||
"databricks_v2_wire_route": "anthropic-messages",
|
||||
"normalization_policy": "none",
|
||||
"registry_label": "Claude Opus 4.7"
|
||||
},
|
||||
{
|
||||
"id": "anthropic-adaptive-xhigh-opus-4-8",
|
||||
"match_kind": "prefix",
|
||||
"match_value": "claude-opus-4-8",
|
||||
"providers": [
|
||||
"anthropic",
|
||||
"databricks_v2"
|
||||
],
|
||||
"match_priority": 10,
|
||||
"thinking_mode": "adaptive",
|
||||
"supported_efforts": [
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"xhigh",
|
||||
"max"
|
||||
],
|
||||
"default_effort": "high",
|
||||
"databricks_v2_wire_route": "anthropic-messages",
|
||||
"normalization_policy": "none",
|
||||
"registry_label": "Claude Opus 4.8"
|
||||
},
|
||||
{
|
||||
"id": "anthropic-adaptive-xhigh-opus-5",
|
||||
"match_kind": "prefix",
|
||||
"match_value": "claude-opus-5",
|
||||
"providers": [
|
||||
"anthropic",
|
||||
"databricks_v2"
|
||||
],
|
||||
"match_priority": 10,
|
||||
"thinking_mode": "adaptive",
|
||||
"supported_efforts": [
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"xhigh",
|
||||
"max"
|
||||
],
|
||||
"default_effort": "high",
|
||||
"databricks_v2_wire_route": "anthropic-messages",
|
||||
"normalization_policy": "none",
|
||||
"registry_label": "Claude Opus 5"
|
||||
},
|
||||
{
|
||||
"id": "anthropic-adaptive-xhigh-sonnet-5",
|
||||
"match_kind": "prefix",
|
||||
"match_value": "claude-sonnet-5",
|
||||
"providers": [
|
||||
"anthropic",
|
||||
"databricks_v2"
|
||||
],
|
||||
"match_priority": 10,
|
||||
"thinking_mode": "adaptive",
|
||||
"supported_efforts": [
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"xhigh",
|
||||
"max"
|
||||
],
|
||||
"default_effort": "high",
|
||||
"databricks_v2_wire_route": "anthropic-messages",
|
||||
"normalization_policy": "none",
|
||||
"registry_label": "Claude Sonnet 5"
|
||||
},
|
||||
{
|
||||
"id": "anthropic-adaptive-xhigh-fable-5",
|
||||
"match_kind": "prefix",
|
||||
"match_value": "claude-fable-5",
|
||||
"providers": [
|
||||
"anthropic",
|
||||
"databricks_v2"
|
||||
],
|
||||
"match_priority": 10,
|
||||
"thinking_mode": "adaptive",
|
||||
"supported_efforts": [
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"xhigh",
|
||||
"max"
|
||||
],
|
||||
"default_effort": "high",
|
||||
"databricks_v2_wire_route": "anthropic-messages",
|
||||
"normalization_policy": "none",
|
||||
"registry_label": "Claude Fable 5"
|
||||
},
|
||||
{
|
||||
"id": "anthropic-adaptive-xhigh-mythos-5",
|
||||
"match_kind": "prefix",
|
||||
"match_value": "claude-mythos-5",
|
||||
"providers": [
|
||||
"anthropic",
|
||||
"databricks_v2"
|
||||
],
|
||||
"match_priority": 10,
|
||||
"thinking_mode": "adaptive",
|
||||
"supported_efforts": [
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"xhigh",
|
||||
"max"
|
||||
],
|
||||
"default_effort": "high",
|
||||
"databricks_v2_wire_route": "anthropic-messages",
|
||||
"normalization_policy": "none",
|
||||
"registry_label": "Claude Mythos 5"
|
||||
},
|
||||
{
|
||||
"id": "anthropic-adaptive-no-xhigh-opus-4-6",
|
||||
"match_kind": "prefix",
|
||||
"match_value": "claude-opus-4-6",
|
||||
"providers": [
|
||||
"anthropic",
|
||||
"databricks_v2"
|
||||
],
|
||||
"match_priority": 10,
|
||||
"thinking_mode": "adaptive",
|
||||
"supported_efforts": [
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"max"
|
||||
],
|
||||
"default_effort": "high",
|
||||
"databricks_v2_wire_route": "anthropic-messages",
|
||||
"normalization_policy": "none",
|
||||
"registry_label": "Claude Opus 4.6"
|
||||
},
|
||||
{
|
||||
"id": "anthropic-adaptive-no-xhigh-sonnet-4-6",
|
||||
"match_kind": "prefix",
|
||||
"match_value": "claude-sonnet-4-6",
|
||||
"providers": [
|
||||
"anthropic",
|
||||
"databricks_v2"
|
||||
],
|
||||
"match_priority": 10,
|
||||
"thinking_mode": "adaptive",
|
||||
"supported_efforts": [
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"max"
|
||||
],
|
||||
"default_effort": "high",
|
||||
"databricks_v2_wire_route": "anthropic-messages",
|
||||
"normalization_policy": "none",
|
||||
"registry_label": "Claude Sonnet 4.6"
|
||||
},
|
||||
{
|
||||
"id": "anthropic-adaptive-no-xhigh-mythos-preview",
|
||||
"match_kind": "prefix",
|
||||
"match_value": "claude-mythos-preview",
|
||||
"providers": [
|
||||
"anthropic",
|
||||
"databricks_v2"
|
||||
],
|
||||
"match_priority": 10,
|
||||
"thinking_mode": "adaptive",
|
||||
"supported_efforts": [
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"max"
|
||||
],
|
||||
"default_effort": "high",
|
||||
"databricks_v2_wire_route": "anthropic-messages",
|
||||
"normalization_policy": "none",
|
||||
"registry_label": "Claude Mythos Preview"
|
||||
},
|
||||
{
|
||||
"id": "openai-gpt5-pro",
|
||||
"match_kind": "gpt5-token",
|
||||
"match_value": "gpt-5-pro",
|
||||
"match_aliases": [
|
||||
"gpt5-pro"
|
||||
],
|
||||
"providers": [
|
||||
"openai",
|
||||
"databricks_v2"
|
||||
],
|
||||
"match_priority": 20,
|
||||
"thinking_mode": "none",
|
||||
"supported_efforts": [
|
||||
"high"
|
||||
],
|
||||
"default_effort": "high",
|
||||
"databricks_v2_wire_route": "openai-responses",
|
||||
"normalization_policy": "openai-standard",
|
||||
"registry_label": "GPT-5 Pro"
|
||||
},
|
||||
{
|
||||
"id": "openai-gpt5-6",
|
||||
"match_kind": "gpt5-token",
|
||||
"match_value": "gpt-5.6",
|
||||
"match_aliases": [
|
||||
"gpt5.6",
|
||||
"gpt-5-6",
|
||||
"gpt5-6"
|
||||
],
|
||||
"providers": [
|
||||
"openai",
|
||||
"databricks_v2"
|
||||
],
|
||||
"match_priority": 15,
|
||||
"thinking_mode": "none",
|
||||
"supported_efforts": [
|
||||
"none",
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"xhigh",
|
||||
"max"
|
||||
],
|
||||
"default_effort": "medium",
|
||||
"databricks_v2_wire_route": "openai-responses",
|
||||
"normalization_policy": "openai-standard",
|
||||
"registry_label": "GPT-5.6"
|
||||
},
|
||||
{
|
||||
"id": "openai-gpt5-5",
|
||||
"match_kind": "gpt5-token",
|
||||
"match_value": "gpt-5.5",
|
||||
"match_aliases": [
|
||||
"gpt5.5",
|
||||
"gpt-5-5",
|
||||
"gpt5-5"
|
||||
],
|
||||
"providers": [
|
||||
"openai",
|
||||
"databricks_v2"
|
||||
],
|
||||
"match_priority": 15,
|
||||
"thinking_mode": "none",
|
||||
"supported_efforts": [
|
||||
"none",
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"xhigh"
|
||||
],
|
||||
"default_effort": "medium",
|
||||
"databricks_v2_wire_route": "openai-responses",
|
||||
"normalization_policy": "openai-standard",
|
||||
"registry_label": "GPT-5.5"
|
||||
},
|
||||
{
|
||||
"id": "openai-gpt5-4",
|
||||
"match_kind": "gpt5-token",
|
||||
"match_value": "gpt-5.4",
|
||||
"match_aliases": [
|
||||
"gpt5.4",
|
||||
"gpt-5-4",
|
||||
"gpt5-4"
|
||||
],
|
||||
"providers": [
|
||||
"openai",
|
||||
"databricks_v2"
|
||||
],
|
||||
"match_priority": 15,
|
||||
"thinking_mode": "none",
|
||||
"supported_efforts": [
|
||||
"none",
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"xhigh"
|
||||
],
|
||||
"default_effort": "medium",
|
||||
"databricks_v2_wire_route": "openai-responses",
|
||||
"normalization_policy": "openai-standard",
|
||||
"registry_label": "GPT-5.4"
|
||||
},
|
||||
{
|
||||
"id": "openai-gpt5-1",
|
||||
"match_kind": "gpt5-token",
|
||||
"match_value": "gpt-5.1",
|
||||
"match_aliases": [
|
||||
"gpt5.1",
|
||||
"gpt-5-1",
|
||||
"gpt5-1"
|
||||
],
|
||||
"providers": [
|
||||
"openai",
|
||||
"databricks_v2"
|
||||
],
|
||||
"match_priority": 15,
|
||||
"thinking_mode": "none",
|
||||
"supported_efforts": [
|
||||
"none",
|
||||
"low",
|
||||
"medium",
|
||||
"high"
|
||||
],
|
||||
"default_effort": "none",
|
||||
"databricks_v2_wire_route": "openai-responses",
|
||||
"normalization_policy": "openai-standard",
|
||||
"registry_label": "GPT-5.1"
|
||||
},
|
||||
{
|
||||
"id": "openai-gpt5-base",
|
||||
"match_kind": "gpt5-base",
|
||||
"match_value": "gpt-5",
|
||||
"match_aliases": [
|
||||
"gpt5"
|
||||
],
|
||||
"providers": [
|
||||
"openai",
|
||||
"databricks_v2"
|
||||
],
|
||||
"match_priority": 10,
|
||||
"thinking_mode": "none",
|
||||
"supported_efforts": [
|
||||
"minimal",
|
||||
"low",
|
||||
"medium",
|
||||
"high"
|
||||
],
|
||||
"default_effort": "medium",
|
||||
"databricks_v2_wire_route": "openai-responses",
|
||||
"normalization_policy": "openai-standard",
|
||||
"registry_label": "GPT-5"
|
||||
},
|
||||
{
|
||||
"id": "dbv2-claude-code-names-segment",
|
||||
"_comment": "DBv2-only rule: endpoint names containing a Claude code-name segment (opus, sonnet, haiku, mythos, fable, claude) route via Anthropic Messages. This matches goose-opus-5 (segments: goose,opus,5) etc. Effort classification uses conservative defaults because prefix-stripped alias ('opus-5') is not a recognized Claude family.",
|
||||
"match_kind": "segment",
|
||||
"match_value": "claude",
|
||||
"match_aliases": [
|
||||
"opus",
|
||||
"sonnet",
|
||||
"haiku",
|
||||
"mythos",
|
||||
"fable"
|
||||
],
|
||||
"providers": [
|
||||
"databricks_v2"
|
||||
],
|
||||
"match_priority": 5,
|
||||
"thinking_mode": "omit-fields",
|
||||
"supported_efforts": [
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"xhigh",
|
||||
"max"
|
||||
],
|
||||
"default_effort": "high",
|
||||
"databricks_v2_wire_route": "anthropic-messages",
|
||||
"normalization_policy": "none"
|
||||
},
|
||||
{
|
||||
"id": "dbv2-gpt-code-names-segment",
|
||||
"_comment": "DBv2-only rule: endpoint names containing a GPT segment prefix (gpt*) route via OpenAI Responses. Handles 'gpt', 'gpt5', 'gpt-5' segments. Priority < individual gpt5 family rules so explicit families take precedence.",
|
||||
"match_kind": "segment-prefix",
|
||||
"match_value": "gpt",
|
||||
"providers": [
|
||||
"databricks_v2"
|
||||
],
|
||||
"match_priority": 5,
|
||||
"thinking_mode": "none",
|
||||
"supported_efforts": [
|
||||
"none",
|
||||
"minimal",
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"xhigh"
|
||||
],
|
||||
"default_effort": "medium",
|
||||
"databricks_v2_wire_route": "openai-responses",
|
||||
"normalization_policy": "openai-clamp-max-to-xhigh"
|
||||
},
|
||||
{
|
||||
"id": "dbv2-sol-luna-terra-segment",
|
||||
"_comment": "DBv2-only rule: sol/luna/terra are OpenAI code names. Route via OpenAI Responses. Must use segment match to avoid matching substrings (consolidated-llama has 'sol' but not as a segment).",
|
||||
"match_kind": "segment",
|
||||
"match_value": "sol",
|
||||
"match_aliases": [
|
||||
"luna",
|
||||
"terra"
|
||||
],
|
||||
"providers": [
|
||||
"databricks_v2"
|
||||
],
|
||||
"match_priority": 5,
|
||||
"thinking_mode": "none",
|
||||
"supported_efforts": [
|
||||
"none",
|
||||
"minimal",
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"xhigh"
|
||||
],
|
||||
"default_effort": "medium",
|
||||
"databricks_v2_wire_route": "openai-responses",
|
||||
"normalization_policy": "openai-clamp-max-to-xhigh"
|
||||
}
|
||||
],
|
||||
"_comment_registry_labels": "All 30 Databricks v2 endpoint-ID to display-name pairs. Represented as [{id,label}] array so duplicate-ID detection is structurally possible. Generated into DATABRICKS_MODEL_NAMES in both Rust and TS.",
|
||||
"registry_labels": [
|
||||
{
|
||||
"id": "databricks-claude-haiku-4-5",
|
||||
"label": "Claude Haiku 4.5 (latest)"
|
||||
},
|
||||
{
|
||||
"id": "databricks-claude-opus-4-1",
|
||||
"label": "Claude Opus 4.1 (latest)"
|
||||
},
|
||||
{
|
||||
"id": "databricks-claude-opus-4-5",
|
||||
"label": "Claude Opus 4.5 (latest)"
|
||||
},
|
||||
{
|
||||
"id": "databricks-claude-opus-4-6",
|
||||
"label": "Claude Opus 4.6"
|
||||
},
|
||||
{
|
||||
"id": "databricks-claude-opus-4-7",
|
||||
"label": "Claude Opus 4.7"
|
||||
},
|
||||
{
|
||||
"id": "databricks-claude-sonnet-4",
|
||||
"label": "Claude Sonnet 4.5"
|
||||
},
|
||||
{
|
||||
"id": "databricks-claude-sonnet-4-5",
|
||||
"label": "Claude Sonnet 4.5 (latest)"
|
||||
},
|
||||
{
|
||||
"id": "databricks-claude-sonnet-4-6",
|
||||
"label": "Claude Sonnet 4.6"
|
||||
},
|
||||
{
|
||||
"id": "databricks-gemini-2-5-flash",
|
||||
"label": "Gemini 2.5 Flash"
|
||||
},
|
||||
{
|
||||
"id": "databricks-gemini-2-5-pro",
|
||||
"label": "Gemini 2.5 Pro"
|
||||
},
|
||||
{
|
||||
"id": "databricks-gemini-3-1-flash-lite",
|
||||
"label": "Gemini 3.1 Flash Lite Preview"
|
||||
},
|
||||
{
|
||||
"id": "databricks-gemini-3-1-pro",
|
||||
"label": "Gemini 3.1 Pro Preview Custom Tools"
|
||||
},
|
||||
{
|
||||
"id": "databricks-gemini-3-flash",
|
||||
"label": "Gemini 3 Flash Preview"
|
||||
},
|
||||
{
|
||||
"id": "databricks-gemini-3-pro",
|
||||
"label": "Gemini 3 Pro Preview"
|
||||
},
|
||||
{
|
||||
"id": "databricks-glm-5-2",
|
||||
"label": "GLM-5.2"
|
||||
},
|
||||
{
|
||||
"id": "databricks-gpt-5",
|
||||
"label": "GPT-5"
|
||||
},
|
||||
{
|
||||
"id": "databricks-gpt-5-1",
|
||||
"label": "GPT-5.1"
|
||||
},
|
||||
{
|
||||
"id": "databricks-gpt-5-2",
|
||||
"label": "GPT-5.2"
|
||||
},
|
||||
{
|
||||
"id": "databricks-gpt-5-4",
|
||||
"label": "GPT-5.4"
|
||||
},
|
||||
{
|
||||
"id": "databricks-gpt-5-4-mini",
|
||||
"label": "GPT-5.4 mini"
|
||||
},
|
||||
{
|
||||
"id": "databricks-gpt-5-4-nano",
|
||||
"label": "GPT-5.4 nano"
|
||||
},
|
||||
{
|
||||
"id": "databricks-gpt-5-5",
|
||||
"label": "GPT-5.5"
|
||||
},
|
||||
{
|
||||
"id": "databricks-gpt-5-6-luna",
|
||||
"label": "GPT-5.6 Luna"
|
||||
},
|
||||
{
|
||||
"id": "databricks-gpt-5-6-sol",
|
||||
"label": "GPT-5.6 Sol"
|
||||
},
|
||||
{
|
||||
"id": "databricks-gpt-5-6-terra",
|
||||
"label": "GPT-5.6 Terra"
|
||||
},
|
||||
{
|
||||
"id": "databricks-gpt-5-mini",
|
||||
"label": "GPT-5 Mini"
|
||||
},
|
||||
{
|
||||
"id": "databricks-gpt-5-nano",
|
||||
"label": "GPT-5 Nano"
|
||||
},
|
||||
{
|
||||
"id": "databricks-gpt-oss-120b",
|
||||
"label": "GPT OSS 120B"
|
||||
},
|
||||
{
|
||||
"id": "databricks-gpt-oss-20b",
|
||||
"label": "GPT OSS 20B"
|
||||
},
|
||||
{
|
||||
"id": "databricks-kimi-k2-7-code",
|
||||
"label": "Kimi K2.7 Code"
|
||||
}
|
||||
],
|
||||
"_comment_databricks_v2_known_models": "Authoritative list of Databricks v2 known model IDs. Mirrors goose DATABRICKS_V2_KNOWN_MODELS at revision 6789d4af (crates/goose-providers/src/databricks_v2.rs:41-42). Generated into DATABRICKS_V2_KNOWN_MODELS in both Rust and TS. Uniqueness enforced by the generator. Opt-in drift check: node scripts/generate-model-capabilities.mjs --check-goose",
|
||||
"databricks_v2_known_models": [
|
||||
"databricks-gpt-5-5",
|
||||
"databricks-claude-opus-4-7"
|
||||
],
|
||||
"exact_records": [
|
||||
{
|
||||
"provider": "databricks_v2",
|
||||
"raw_model_id": "databricks-gpt-5-4-mini",
|
||||
"registry_label": "GPT-5.4 Mini",
|
||||
"supported_efforts_override": [
|
||||
"low",
|
||||
"medium",
|
||||
"high"
|
||||
],
|
||||
"source": "models.dev reasoning_options: low|medium|high (family rule adds none+xhigh \u2014 adopt provider-advertised)",
|
||||
"_reconciliation": "adopt",
|
||||
"_reconciliation_note": "models.dev advertises low|medium|high. Family rule (gpt5-4) adds none+xhigh. Provider-advertised wins per plan F1 policy.",
|
||||
"_reconciliation_doc": "https://models.dev/api.json (retrieved 2026-07-31, SHA-256 d5a4974cd69f19b0f67713acaa6bb3b16e920defdc07ecbdf6b0a936181bb0e0): providers.databricks.models[\"databricks-gpt-5-4-mini\"].reasoning_options=[{\"type\":\"effort\",\"values\":[\"low\",\"medium\",\"high\"]}]"
|
||||
},
|
||||
{
|
||||
"provider": "databricks_v2",
|
||||
"raw_model_id": "databricks-gpt-5-4-nano",
|
||||
"registry_label": "GPT-5.4 Nano",
|
||||
"supported_efforts_override": [
|
||||
"low",
|
||||
"medium",
|
||||
"high"
|
||||
],
|
||||
"source": "models.dev reasoning_options: low|medium|high",
|
||||
"_reconciliation": "adopt",
|
||||
"_reconciliation_note": "models.dev advertises low|medium|high. Same as gpt-5-4-mini. Adopt.",
|
||||
"_reconciliation_doc": "https://models.dev/api.json (retrieved 2026-07-31, SHA-256 d5a4974cd69f19b0f67713acaa6bb3b16e920defdc07ecbdf6b0a936181bb0e0): providers.databricks.models[\"databricks-gpt-5-4-nano\"].reasoning_options=[{\"type\":\"effort\",\"values\":[\"low\",\"medium\",\"high\"]}]"
|
||||
},
|
||||
{
|
||||
"provider": "databricks_v2",
|
||||
"raw_model_id": "databricks-gpt-5-6-sol",
|
||||
"registry_label": "GPT-5.6 Sol",
|
||||
"supported_efforts_override": [
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"max"
|
||||
],
|
||||
"source": "models.dev reasoning_options: low|medium|high|max (family rule adds none+xhigh \u2014 provider-advertised wins per plan F1)",
|
||||
"_reconciliation": "adopt",
|
||||
"_reconciliation_note": "models.dev advertises [low, medium, high, max]. Family rule (gpt5-6) has none+xhigh+max; sol endpoint does not expose none or xhigh. Provider-advertised wins.",
|
||||
"_reconciliation_doc": "https://models.dev/api.json (retrieved 2026-07-31, SHA-256 d5a4974cd69f19b0f67713acaa6bb3b16e920defdc07ecbdf6b0a936181bb0e0): providers.databricks.models[\"databricks-gpt-5-6-sol\"].reasoning_options=[{\"type\":\"effort\",\"values\":[\"low\",\"medium\",\"high\",\"max\"]}]"
|
||||
},
|
||||
{
|
||||
"provider": "databricks_v2",
|
||||
"raw_model_id": "databricks-gpt-5-5",
|
||||
"registry_label": "GPT-5.5",
|
||||
"source": "models.dev reasoning_options: low|medium|high (family rule adds none+xhigh \u2014 provider-advertised wins per plan F1)",
|
||||
"_reconciliation": "adopt",
|
||||
"_reconciliation_note": "models.dev (pinned payload) advertises [low, medium, high]. Family rule (gpt5-5) has none+xhigh; this Databricks endpoint does not expose none or xhigh. Provider-advertised wins.",
|
||||
"supported_efforts_override": [
|
||||
"low",
|
||||
"medium",
|
||||
"high"
|
||||
],
|
||||
"_reconciliation_doc": "https://models.dev/api.json (retrieved 2026-07-31, SHA-256 d5a4974cd69f19b0f67713acaa6bb3b16e920defdc07ecbdf6b0a936181bb0e0): providers.databricks.models[\"databricks-gpt-5-5\"].reasoning_options=[{\"type\":\"effort\",\"values\":[\"low\",\"medium\",\"high\"]}]"
|
||||
},
|
||||
{
|
||||
"provider": "databricks_v2",
|
||||
"raw_model_id": "databricks-claude-opus-4-7",
|
||||
"registry_label": "Claude Opus 4.7",
|
||||
"source": "DATABRICKS_V2_KNOWN_MODELS; family rule anthropic-adaptive-xhigh-opus-4-7 applies",
|
||||
"_reconciliation": "no-effort-divergence",
|
||||
"_reconciliation_note": "models.dev advertises reasoning_options=[{\"type\":\"budget_tokens\",\"min\":1024}]. This is a different capability axis (extended thinking token budget), not an effort-level selector. No effort divergence to reconcile \u2014 efforts for this model come from the anthropic family rule (anthropic-adaptive-xhigh-opus-4-7).",
|
||||
"_reconciliation_doc": "https://models.dev/api.json (retrieved 2026-07-31, SHA-256 d5a4974cd69f19b0f67713acaa6bb3b16e920defdc07ecbdf6b0a936181bb0e0): providers.databricks.models[\"databricks-claude-opus-4-7\"].reasoning_options=[{\"type\":\"budget_tokens\",\"min\":1024}]"
|
||||
}
|
||||
],
|
||||
"provider_fallbacks": {
|
||||
"anthropic": {
|
||||
"blank": {
|
||||
"databricks_v2_wire_route": "not-applicable",
|
||||
"thinking_mode": "adaptive",
|
||||
"supported_efforts": [
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"xhigh",
|
||||
"max"
|
||||
],
|
||||
"default_effort": "high",
|
||||
"normalization_policy": "none"
|
||||
},
|
||||
"concrete_unknown": {
|
||||
"databricks_v2_wire_route": "not-applicable",
|
||||
"thinking_mode": "omit-fields",
|
||||
"supported_efforts": [
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"xhigh",
|
||||
"max"
|
||||
],
|
||||
"default_effort": "high",
|
||||
"normalization_policy": "none"
|
||||
}
|
||||
},
|
||||
"openai": {
|
||||
"blank": {
|
||||
"databricks_v2_wire_route": "not-applicable",
|
||||
"thinking_mode": "none",
|
||||
"supported_efforts": [
|
||||
"none",
|
||||
"minimal",
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"xhigh"
|
||||
],
|
||||
"default_effort": "medium",
|
||||
"normalization_policy": "openai-clamp-max-to-xhigh"
|
||||
},
|
||||
"concrete_unknown": {
|
||||
"databricks_v2_wire_route": "not-applicable",
|
||||
"thinking_mode": "none",
|
||||
"supported_efforts": [
|
||||
"none",
|
||||
"minimal",
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"xhigh"
|
||||
],
|
||||
"default_effort": "medium",
|
||||
"normalization_policy": "openai-clamp-max-to-xhigh"
|
||||
}
|
||||
},
|
||||
"databricks_v2": {
|
||||
"blank": {
|
||||
"databricks_v2_wire_route": "route-unknown",
|
||||
"thinking_mode": "none",
|
||||
"supported_efforts": [
|
||||
"none",
|
||||
"minimal",
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"xhigh",
|
||||
"max"
|
||||
],
|
||||
"default_effort": "medium",
|
||||
"normalization_policy": "openai-clamp-max-to-xhigh"
|
||||
},
|
||||
"concrete_unknown": {
|
||||
"databricks_v2_wire_route": "mlflow-chat",
|
||||
"thinking_mode": "none",
|
||||
"supported_efforts": [
|
||||
"none",
|
||||
"minimal",
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"xhigh"
|
||||
],
|
||||
"default_effort": "medium",
|
||||
"normalization_policy": "openai-clamp-max-to-xhigh"
|
||||
}
|
||||
},
|
||||
"databricks": {
|
||||
"blank": {
|
||||
"databricks_v2_wire_route": "not-applicable",
|
||||
"thinking_mode": "none",
|
||||
"supported_efforts": [
|
||||
"none",
|
||||
"minimal",
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"xhigh"
|
||||
],
|
||||
"default_effort": "medium",
|
||||
"normalization_policy": "openai-clamp-max-to-xhigh"
|
||||
},
|
||||
"concrete_unknown": {
|
||||
"databricks_v2_wire_route": "not-applicable",
|
||||
"thinking_mode": "none",
|
||||
"supported_efforts": [
|
||||
"none",
|
||||
"minimal",
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"xhigh"
|
||||
],
|
||||
"default_effort": "medium",
|
||||
"normalization_policy": "openai-clamp-max-to-xhigh"
|
||||
}
|
||||
},
|
||||
"openrouter": {
|
||||
"blank": {
|
||||
"databricks_v2_wire_route": "not-applicable",
|
||||
"thinking_mode": "none",
|
||||
"supported_efforts": [
|
||||
"none",
|
||||
"minimal",
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"xhigh",
|
||||
"max"
|
||||
],
|
||||
"default_effort": "medium",
|
||||
"normalization_policy": "none"
|
||||
},
|
||||
"concrete_unknown": {
|
||||
"databricks_v2_wire_route": "not-applicable",
|
||||
"thinking_mode": "none",
|
||||
"supported_efforts": [
|
||||
"none",
|
||||
"minimal",
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"xhigh",
|
||||
"max"
|
||||
],
|
||||
"default_effort": "medium",
|
||||
"normalization_policy": "none"
|
||||
}
|
||||
},
|
||||
"_default": {
|
||||
"blank": {
|
||||
"databricks_v2_wire_route": "not-applicable",
|
||||
"thinking_mode": "none",
|
||||
"supported_efforts": [
|
||||
"none",
|
||||
"minimal",
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"xhigh",
|
||||
"max"
|
||||
],
|
||||
"default_effort": "medium",
|
||||
"normalization_policy": "none"
|
||||
},
|
||||
"concrete_unknown": {
|
||||
"databricks_v2_wire_route": "not-applicable",
|
||||
"thinking_mode": "none",
|
||||
"supported_efforts": [
|
||||
"none",
|
||||
"minimal",
|
||||
"low",
|
||||
"medium",
|
||||
"high",
|
||||
"xhigh",
|
||||
"max"
|
||||
],
|
||||
"default_effort": "medium",
|
||||
"normalization_policy": "none"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,485 @@
|
||||
[
|
||||
{
|
||||
"_group": "Anthropic exact family rules",
|
||||
"_note": "All require thinking_mode=manual-budget or adaptive, correct supported_efforts, databricks_v2_wire_route=not-applicable"
|
||||
},
|
||||
{
|
||||
"id": "anthropic-claude-3-family",
|
||||
"provider": "anthropic",
|
||||
"raw_model_id": "claude-3-7-sonnet-20250219",
|
||||
"expect": {
|
||||
"thinking_mode": "manual-budget",
|
||||
"supported_efforts": ["low", "medium", "high"],
|
||||
"default_effort": null,
|
||||
"databricks_v2_wire_route": "not-applicable"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "anthropic-claude-opus-4-5",
|
||||
"provider": "anthropic",
|
||||
"raw_model_id": "claude-opus-4-5",
|
||||
"expect": {
|
||||
"thinking_mode": "manual-budget",
|
||||
"supported_efforts": ["low", "medium", "high"],
|
||||
"default_effort": null,
|
||||
"databricks_v2_wire_route": "not-applicable"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "anthropic-claude-opus-4-7",
|
||||
"provider": "anthropic",
|
||||
"raw_model_id": "claude-opus-4-7",
|
||||
"expect": {
|
||||
"thinking_mode": "adaptive",
|
||||
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
|
||||
"default_effort": "high",
|
||||
"databricks_v2_wire_route": "not-applicable"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "anthropic-claude-opus-4-8",
|
||||
"provider": "anthropic",
|
||||
"raw_model_id": "claude-opus-4-8",
|
||||
"expect": {
|
||||
"thinking_mode": "adaptive",
|
||||
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
|
||||
"default_effort": "high",
|
||||
"databricks_v2_wire_route": "not-applicable"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "anthropic-claude-sonnet-5",
|
||||
"provider": "anthropic",
|
||||
"raw_model_id": "claude-sonnet-5-20260101",
|
||||
"expect": {
|
||||
"thinking_mode": "adaptive",
|
||||
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
|
||||
"default_effort": "high",
|
||||
"databricks_v2_wire_route": "not-applicable"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "anthropic-claude-fable-5",
|
||||
"provider": "anthropic",
|
||||
"raw_model_id": "claude-fable-5",
|
||||
"expect": {
|
||||
"thinking_mode": "adaptive",
|
||||
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
|
||||
"default_effort": "high",
|
||||
"databricks_v2_wire_route": "not-applicable"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "anthropic-claude-mythos-5",
|
||||
"provider": "anthropic",
|
||||
"raw_model_id": "claude-mythos-5",
|
||||
"expect": {
|
||||
"thinking_mode": "adaptive",
|
||||
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
|
||||
"default_effort": "high",
|
||||
"databricks_v2_wire_route": "not-applicable"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "anthropic-claude-opus-4-6",
|
||||
"provider": "anthropic",
|
||||
"raw_model_id": "claude-opus-4-6",
|
||||
"expect": {
|
||||
"thinking_mode": "adaptive",
|
||||
"supported_efforts": ["low", "medium", "high", "max"],
|
||||
"default_effort": "high",
|
||||
"databricks_v2_wire_route": "not-applicable"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "anthropic-claude-sonnet-4-6",
|
||||
"provider": "anthropic",
|
||||
"raw_model_id": "claude-sonnet-4-6",
|
||||
"expect": {
|
||||
"thinking_mode": "adaptive",
|
||||
"supported_efforts": ["low", "medium", "high", "max"],
|
||||
"default_effort": "high",
|
||||
"databricks_v2_wire_route": "not-applicable"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "anthropic-claude-mythos-preview",
|
||||
"provider": "anthropic",
|
||||
"raw_model_id": "claude-mythos-preview",
|
||||
"expect": {
|
||||
"thinking_mode": "adaptive",
|
||||
"supported_efforts": ["low", "medium", "high", "max"],
|
||||
"default_effort": "high",
|
||||
"databricks_v2_wire_route": "not-applicable"
|
||||
}
|
||||
},
|
||||
{
|
||||
"_group": "Anthropic unknowns and fallbacks"
|
||||
},
|
||||
{
|
||||
"id": "anthropic-unknown-blank",
|
||||
"provider": "anthropic",
|
||||
"raw_model_id": "",
|
||||
"expect": {
|
||||
"thinking_mode": "adaptive",
|
||||
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
|
||||
"default_effort": "high",
|
||||
"databricks_v2_wire_route": "not-applicable"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "anthropic-unknown-concrete",
|
||||
"provider": "anthropic",
|
||||
"raw_model_id": "claude-ultra-9000",
|
||||
"expect": {
|
||||
"thinking_mode": "omit-fields",
|
||||
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
|
||||
"default_effort": "high",
|
||||
"databricks_v2_wire_route": "not-applicable"
|
||||
}
|
||||
},
|
||||
{
|
||||
"_group": "OpenAI exact family rules"
|
||||
},
|
||||
{
|
||||
"id": "openai-gpt5-pro",
|
||||
"provider": "openai",
|
||||
"raw_model_id": "gpt-5-pro",
|
||||
"expect": {
|
||||
"thinking_mode": "none",
|
||||
"supported_efforts": ["high"],
|
||||
"default_effort": "high",
|
||||
"databricks_v2_wire_route": "not-applicable"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "openai-gpt5.6",
|
||||
"provider": "openai",
|
||||
"raw_model_id": "gpt-5.6",
|
||||
"expect": {
|
||||
"thinking_mode": "none",
|
||||
"supported_efforts": ["none", "low", "medium", "high", "xhigh", "max"],
|
||||
"default_effort": "medium",
|
||||
"databricks_v2_wire_route": "not-applicable"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "openai-gpt5-6-dashed",
|
||||
"provider": "openai",
|
||||
"raw_model_id": "gpt-5-6",
|
||||
"expect": {
|
||||
"thinking_mode": "none",
|
||||
"supported_efforts": ["none", "low", "medium", "high", "xhigh", "max"],
|
||||
"default_effort": "medium",
|
||||
"databricks_v2_wire_route": "not-applicable"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "openai-gpt5.5",
|
||||
"provider": "openai",
|
||||
"raw_model_id": "gpt-5.5",
|
||||
"expect": {
|
||||
"thinking_mode": "none",
|
||||
"supported_efforts": ["none", "low", "medium", "high", "xhigh"],
|
||||
"default_effort": "medium",
|
||||
"databricks_v2_wire_route": "not-applicable"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "openai-gpt5.4",
|
||||
"provider": "openai",
|
||||
"raw_model_id": "gpt-5.4",
|
||||
"expect": {
|
||||
"thinking_mode": "none",
|
||||
"supported_efforts": ["none", "low", "medium", "high", "xhigh"],
|
||||
"default_effort": "medium",
|
||||
"databricks_v2_wire_route": "not-applicable"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "openai-gpt5.1",
|
||||
"provider": "openai",
|
||||
"raw_model_id": "gpt-5.1",
|
||||
"expect": {
|
||||
"thinking_mode": "none",
|
||||
"supported_efforts": ["none", "low", "medium", "high"],
|
||||
"default_effort": "none",
|
||||
"databricks_v2_wire_route": "not-applicable"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "openai-gpt5-base",
|
||||
"provider": "openai",
|
||||
"raw_model_id": "gpt-5",
|
||||
"expect": {
|
||||
"thinking_mode": "none",
|
||||
"supported_efforts": ["minimal", "low", "medium", "high"],
|
||||
"default_effort": "medium",
|
||||
"databricks_v2_wire_route": "not-applicable"
|
||||
}
|
||||
},
|
||||
{
|
||||
"_group": "OpenAI adversarial — gpt5 boundary-aware matching (ported from config.rs tests)"
|
||||
},
|
||||
{
|
||||
"id": "openai-gpt5-1106-should-not-match-base",
|
||||
"provider": "openai",
|
||||
"raw_model_id": "gpt-5-1106",
|
||||
"_note": "gpt-5-1106: '-1106' is a 4-digit date segment, NOT a short version (gpt5-base rejects only 1-3 digit suffixes). Must match base table [minimal,low,medium,high], NOT fall through to unknown.",
|
||||
"expect": {
|
||||
"supported_efforts": ["minimal", "low", "medium", "high"]
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "openai-gpt5-4o-matches-base",
|
||||
"provider": "openai",
|
||||
"raw_model_id": "gpt-5-4o",
|
||||
"_note": "gpt-5-4o: '4o' after '-' is NOT a short numeric suffix (it contains a letter). Must match gpt5-base. Crucially, must NOT match gpt-5.4 (the '4' is followed by 'o', not boundary char).",
|
||||
"expect": {
|
||||
"supported_efforts": ["minimal", "low", "medium", "high"]
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "openai-gpt5-pro-not-matching-gpt5-base",
|
||||
"provider": "openai",
|
||||
"raw_model_id": "gpt-5-pro",
|
||||
"_note": "gpt-5-pro should hit gpt5-pro rule (priority 20), NOT gpt-5 base.",
|
||||
"expect": {
|
||||
"supported_efforts": ["high"],
|
||||
"default_effort": "high"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "openai-multi-digit-version-gpt5-10",
|
||||
"provider": "openai",
|
||||
"raw_model_id": "gpt-5-10",
|
||||
"_note": "gpt-5-10 — two-digit suffix prevents gpt5-base match. Falls through to unknown.",
|
||||
"expect": {
|
||||
"supported_efforts": ["none", "minimal", "low", "medium", "high", "xhigh"]
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "openai-gpt5-date-suffix",
|
||||
"provider": "openai",
|
||||
"raw_model_id": "gpt-5-20260101",
|
||||
"_note": "gpt-5-20260101 — long numeric suffix after base: '20260101' is 8 digits, beyond 1-3 digit reject, should hit gpt5-base.",
|
||||
"expect": {
|
||||
"supported_efforts": ["minimal", "low", "medium", "high"]
|
||||
}
|
||||
},
|
||||
{
|
||||
"_group": "DatabricksV2 — segment-based routing (ported from llm.rs tests)"
|
||||
},
|
||||
{
|
||||
"id": "dbv2-gpt5-route-openai-responses",
|
||||
"provider": "databricks_v2",
|
||||
"raw_model_id": "gpt-5.5",
|
||||
"expect": {
|
||||
"databricks_v2_wire_route": "openai-responses",
|
||||
"supported_efforts": ["none", "low", "medium", "high", "xhigh"]
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "dbv2-claude-route-anthropic-messages",
|
||||
"provider": "databricks_v2",
|
||||
"raw_model_id": "claude-opus-4-7",
|
||||
"expect": {
|
||||
"databricks_v2_wire_route": "anthropic-messages",
|
||||
"thinking_mode": "adaptive",
|
||||
"supported_efforts": ["low", "medium", "high", "xhigh", "max"]
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "dbv2-claude-prefix-stripped",
|
||||
"provider": "databricks_v2",
|
||||
"raw_model_id": "databricks-claude-opus-4-7",
|
||||
"_note": "databricks- prefix stripped → claude-opus-4-7 → Anthropic route",
|
||||
"expect": {
|
||||
"databricks_v2_wire_route": "anthropic-messages",
|
||||
"thinking_mode": "adaptive"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "dbv2-goose-claude-prefix-stripped",
|
||||
"provider": "databricks_v2",
|
||||
"raw_model_id": "goose-claude-fable-5",
|
||||
"_note": "goose- prefix stripped → claude-fable-5 → Anthropic adaptive+xhigh",
|
||||
"expect": {
|
||||
"databricks_v2_wire_route": "anthropic-messages",
|
||||
"thinking_mode": "adaptive",
|
||||
"supported_efforts": ["low", "medium", "high", "xhigh", "max"]
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "dbv2-team-prefix-stripped",
|
||||
"provider": "databricks_v2",
|
||||
"raw_model_id": "team-x-claude-opus-4-7",
|
||||
"_note": "team-x- prefix stripped → claude-opus-4-7 → Anthropic route",
|
||||
"expect": {
|
||||
"databricks_v2_wire_route": "anthropic-messages",
|
||||
"thinking_mode": "adaptive",
|
||||
"supported_efforts": ["low", "medium", "high", "xhigh", "max"]
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "dbv2-consolidated-llama-not-sol",
|
||||
"provider": "databricks_v2",
|
||||
"raw_model_id": "consolidated-llama",
|
||||
"_note": "segment test: 'sol' is a SUBSTRING of 'consolidated' — must NOT match DATABRICKS_V2_OPENAI_CODE_NAMES 'sol'. Falls through to mlflow-chat.",
|
||||
"expect": {
|
||||
"databricks_v2_wire_route": "mlflow-chat"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "dbv2-terraform-coder-not-terra",
|
||||
"provider": "databricks_v2",
|
||||
"raw_model_id": "terraform-coder",
|
||||
"_note": "segment test: 'terra' is a prefix of 'terraform' — must NOT match 'terra' code name. Falls through to mlflow-chat.",
|
||||
"expect": {
|
||||
"databricks_v2_wire_route": "mlflow-chat"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "dbv2-corpus-reranker-not-opus",
|
||||
"provider": "databricks_v2",
|
||||
"raw_model_id": "corpus-reranker",
|
||||
"_note": "segment test: 'opus' is NOT a segment of corpus-reranker (segments: corpus, reranker). mlflow-chat.",
|
||||
"expect": {
|
||||
"databricks_v2_wire_route": "mlflow-chat"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "dbv2-octopus-model-not-opus",
|
||||
"provider": "databricks_v2",
|
||||
"raw_model_id": "octopus-model",
|
||||
"_note": "segment test: 'opus' is not a segment of octopus-model (segments: octopus, model). mlflow-chat.",
|
||||
"expect": {
|
||||
"databricks_v2_wire_route": "mlflow-chat"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "dbv2-goose-opus-5-is-anthropic",
|
||||
"provider": "databricks_v2",
|
||||
"raw_model_id": "goose-opus-5",
|
||||
"_note": "'opus' IS a named segment of goose-opus-5 (segments: goose, opus, 5). Routes Anthropic. Key test: agrees with llm.rs but disagreed with old config.rs.",
|
||||
"expect": {
|
||||
"databricks_v2_wire_route": "anthropic-messages"
|
||||
}
|
||||
},
|
||||
{
|
||||
"_group": "P2-A resolver-contract vectors (plan v4 §Resolver contract)"
|
||||
},
|
||||
{
|
||||
"id": "resolver-exact-raw-id-hit",
|
||||
"provider": "databricks_v2",
|
||||
"raw_model_id": "databricks-gpt-5-4-mini",
|
||||
"_note": "Exact record exists. Must return exact Databricks override: low|medium|high (not family's none+xhigh).",
|
||||
"expect": {
|
||||
"supported_efforts": ["low", "medium", "high"]
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "resolver-prefixed-alias-misses-exact",
|
||||
"provider": "databricks_v2",
|
||||
"raw_model_id": "team-x-databricks-gpt-5-4-mini",
|
||||
"_note": "Prefixed alias of an exact ID. Raw exact lookup MUST miss (key is team-x-..., not databricks-...). Falls to family rules (gpt5-4 family → none+xhigh).",
|
||||
"expect": {
|
||||
"supported_efforts": ["none", "low", "medium", "high", "xhigh"]
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "resolver-cross-provider-misses-exact",
|
||||
"provider": "openai",
|
||||
"raw_model_id": "databricks-gpt-5-4-mini",
|
||||
"_note": "Same raw ID but different provider. Exact record is databricks_v2-scoped; must miss. Falls to openai family rules.",
|
||||
"expect": {
|
||||
"databricks_v2_wire_route": "not-applicable"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "resolver-exact-efforts-plus-family-route",
|
||||
"provider": "databricks_v2",
|
||||
"raw_model_id": "databricks-gpt-5-6-sol",
|
||||
"_note": "Exact record with efforts from models.dev (low|medium|high|max — provider-advertised, no none/xhigh). Route materialized from gpt5-6 family rule (openai-responses). Must return both, complete.",
|
||||
"expect": {
|
||||
"supported_efforts": ["low", "medium", "high", "max"],
|
||||
"databricks_v2_wire_route": "openai-responses"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "dbv2-gpt5-5-exact-override",
|
||||
"provider": "databricks_v2",
|
||||
"raw_model_id": "databricks-gpt-5-5",
|
||||
"_note": "Exact record adopts models.dev advertised set [low,medium,high]. Family rule (gpt5-5) has none+xhigh — provider-advertised wins per plan F1.",
|
||||
"expect": {
|
||||
"supported_efforts": ["low", "medium", "high"],
|
||||
"databricks_v2_wire_route": "openai-responses"
|
||||
}
|
||||
},
|
||||
{
|
||||
"_group": "Blank vs concrete-unknown per provider (P2-B fallback vectors)"
|
||||
},
|
||||
{
|
||||
"id": "dbv2-blank-all7-route-unknown",
|
||||
"provider": "databricks_v2",
|
||||
"raw_model_id": "",
|
||||
"_note": "DBv2 blank: route-unknown, all 7 efforts, default medium.",
|
||||
"expect": {
|
||||
"databricks_v2_wire_route": "route-unknown",
|
||||
"supported_efforts": ["none", "minimal", "low", "medium", "high", "xhigh", "max"],
|
||||
"default_effort": "medium"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "dbv2-concrete-unknown-mlflow-no-max",
|
||||
"provider": "databricks_v2",
|
||||
"raw_model_id": "some-unknown-model-xyz",
|
||||
"_note": "DBv2 concrete-unknown: mlflow-chat, all-except-max (6 efforts).",
|
||||
"expect": {
|
||||
"databricks_v2_wire_route": "mlflow-chat",
|
||||
"supported_efforts": ["none", "minimal", "low", "medium", "high", "xhigh"]
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "openai-blank-all-except-max",
|
||||
"provider": "openai",
|
||||
"raw_model_id": "",
|
||||
"_note": "OpenAI blank: not-applicable route, all-except-max, medium default.",
|
||||
"expect": {
|
||||
"databricks_v2_wire_route": "not-applicable",
|
||||
"supported_efforts": ["none", "minimal", "low", "medium", "high", "xhigh"],
|
||||
"default_effort": "medium"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "openai-concrete-unknown-all-except-max",
|
||||
"provider": "openai",
|
||||
"raw_model_id": "gpt-4o",
|
||||
"_note": "OpenAI concrete unknown (unverified family): not-applicable route, all-except-max, medium default.",
|
||||
"expect": {
|
||||
"databricks_v2_wire_route": "not-applicable",
|
||||
"supported_efforts": ["none", "minimal", "low", "medium", "high", "xhigh"],
|
||||
"default_effort": "medium"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "anthropic-blank-adaptive-full",
|
||||
"provider": "anthropic",
|
||||
"raw_model_id": "",
|
||||
"_note": "Anthropic blank: assume adaptive with full support (incl. xhigh).",
|
||||
"expect": {
|
||||
"thinking_mode": "adaptive",
|
||||
"supported_efforts": ["low", "medium", "high", "xhigh", "max"],
|
||||
"default_effort": "high"
|
||||
}
|
||||
},
|
||||
{
|
||||
"id": "anthropic-concrete-unknown-omit-fields",
|
||||
"provider": "anthropic",
|
||||
"raw_model_id": "claude-ultra-9000",
|
||||
"_note": "Anthropic concrete-unknown: omit-fields (never guess request shape).",
|
||||
"expect": {
|
||||
"thinking_mode": "omit-fields"
|
||||
}
|
||||
}
|
||||
]
|
||||
@@ -0,0 +1,83 @@
|
||||
#!/usr/bin/env node
|
||||
/**
|
||||
* Normative corpus runner — validates the generated TS interpreter against
|
||||
* scripts/normative-corpus.json.
|
||||
*
|
||||
* This is the JS side of the two-interpreter corpus check. The runner imports
|
||||
* resolveModelCapabilities() directly from the generated TypeScript module via
|
||||
* Node's --experimental-strip-types flag. The Rust side lives in
|
||||
* crates/buzz-agent/src/generated_model_capabilities.rs (shared corpus harness).
|
||||
*
|
||||
* Usage: node --experimental-strip-types scripts/run-corpus.mjs [--verbose]
|
||||
* Exits 0 on all pass, 1 on any failure.
|
||||
*/
|
||||
|
||||
import { readFileSync } from "node:fs";
|
||||
import { join, dirname } from "node:path";
|
||||
import { fileURLToPath } from "node:url";
|
||||
|
||||
const __dirname = dirname(fileURLToPath(import.meta.url));
|
||||
const repoRoot = join(__dirname, "..");
|
||||
const VERBOSE = process.argv.includes("--verbose");
|
||||
|
||||
// Import the generated TypeScript interpreter directly.
|
||||
// Node 22+ --experimental-strip-types strips type annotations at load time; no build step needed.
|
||||
const { resolveModelCapabilities } = await import(
|
||||
join(repoRoot, "desktop", "src", "features", "agents", "ui", "modelCapabilities.ts")
|
||||
);
|
||||
|
||||
const corpus = JSON.parse(
|
||||
readFileSync(join(repoRoot, "scripts", "normative-corpus.json"), "utf8"),
|
||||
);
|
||||
|
||||
// ----- Run corpus -----
|
||||
|
||||
let passed = 0;
|
||||
let failed = 0;
|
||||
|
||||
for (const entry of corpus) {
|
||||
// Skip group header entries
|
||||
if (entry._group) continue;
|
||||
if (!entry.expect) continue;
|
||||
|
||||
// resolveModelCapabilities returns camelCase keys (registryLabel, thinkingMode, etc.)
|
||||
const result = resolveModelCapabilities(entry.provider, entry.raw_model_id);
|
||||
const expect = entry.expect;
|
||||
|
||||
const failures = [];
|
||||
|
||||
for (const [key, expectedVal] of Object.entries(expect)) {
|
||||
// Corpus uses snake_case; generated TS uses camelCase — convert for lookup.
|
||||
const camelKey = key.replace(/_([a-z])/g, (_, c) => c.toUpperCase());
|
||||
const actualVal = camelKey in result ? result[camelKey] : result[key];
|
||||
|
||||
if (Array.isArray(expectedVal)) {
|
||||
// Order-sensitive comparison for effort arrays
|
||||
const actualArr = Array.isArray(actualVal) ? actualVal : [];
|
||||
if (JSON.stringify(actualArr) !== JSON.stringify(expectedVal)) {
|
||||
failures.push(
|
||||
` ${key}: expected [${expectedVal.join(", ")}] got [${actualArr.join(", ")}]`,
|
||||
);
|
||||
}
|
||||
} else {
|
||||
if (actualVal !== expectedVal) {
|
||||
failures.push(` ${key}: expected ${JSON.stringify(expectedVal)} got ${JSON.stringify(actualVal)}`);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
if (failures.length === 0) {
|
||||
passed++;
|
||||
if (VERBOSE) {
|
||||
console.log(` PASS ${entry.id}`);
|
||||
}
|
||||
} else {
|
||||
failed++;
|
||||
console.error(` FAIL ${entry.id} (${entry.provider} / "${entry.raw_model_id}")`);
|
||||
for (const f of failures) console.error(f);
|
||||
if (entry._note) console.error(` note: ${entry._note}`);
|
||||
}
|
||||
}
|
||||
|
||||
console.log(`\nCorpus: ${passed} passed, ${failed} failed`);
|
||||
process.exit(failed > 0 ? 1 : 0);
|
||||
Executable
+261
@@ -0,0 +1,261 @@
|
||||
#!/usr/bin/env node
|
||||
/**
|
||||
* Per-interpreter mutation evidence runner.
|
||||
*
|
||||
* Introduces deliberate resolver faults into the manifest, regenerates artifacts,
|
||||
* and verifies the shared normative corpus detects every fault in BOTH the generated
|
||||
* TypeScript interpreter (via --experimental-strip-types import) and the Rust
|
||||
* interpreter (via cargo test shared-corpus harness). Exits 0 if all mutations are
|
||||
* killed in both interpreters; exits 1 if any survive.
|
||||
*
|
||||
* Usage:
|
||||
* node --experimental-strip-types scripts/run-mutation-evidence.mjs [--verbose]
|
||||
*
|
||||
* This script is non-CI (run manually to generate MUTATION_EVIDENCE.md). It writes
|
||||
* its findings to stdout in a format suitable for copy-paste into the evidence doc.
|
||||
*
|
||||
* Mutations applied (each in isolation, manifest restored after each run):
|
||||
* M1: Swap anthropic-adaptive-xhigh-opus-4-7 efforts from [low,medium,high,xhigh,max]
|
||||
* to [low,medium,high] — kills corpus vectors that check xhigh/max.
|
||||
* M2: Change gpt5-base supported_efforts to include "xhigh" — kills vectors that
|
||||
* check gpt5-base resolves minimal-only, not xhigh.
|
||||
* M3: Change openai-gpt5-1 default_effort to "high" instead of "none" — kills
|
||||
* the gpt5.1 corpus vector that checks default_effort=none.
|
||||
* M4: Swap databricks_v2_wire_route in dbv2-claude-code-names-segment from
|
||||
* "anthropic-messages" to "openai-responses" — kills segment-route corpus vectors.
|
||||
* M5: Remove all three DBv2 segment rules — kills goose-opus-5 and terraform/consolidated
|
||||
* segment collision vectors.
|
||||
* M6: Change databricks_v2 concrete_unknown fallback route from "mlflow-chat" to
|
||||
* "openai-responses" — kills dbv2-concrete-unknown-mlflow-no-max vector.
|
||||
* M7: Change gpt5-4 supported_efforts to remove "xhigh" — kills
|
||||
* resolver-prefixed-alias-misses-exact vector.
|
||||
*/
|
||||
|
||||
import { readFileSync, writeFileSync } from "node:fs";
|
||||
import { join, dirname } from "node:path";
|
||||
import { fileURLToPath } from "node:url";
|
||||
import { execSync, spawnSync } from "node:child_process";
|
||||
|
||||
const __dirname = dirname(fileURLToPath(import.meta.url));
|
||||
const repoRoot = join(__dirname, "..");
|
||||
const VERBOSE = process.argv.includes("--verbose");
|
||||
|
||||
const manifestPath = join(repoRoot, "scripts", "model-capabilities.json");
|
||||
const generatorPath = join(repoRoot, "scripts", "generate-model-capabilities.mjs");
|
||||
const jsRunnerPath = join(repoRoot, "scripts", "run-corpus.mjs");
|
||||
|
||||
const originalManifest = readFileSync(manifestPath, "utf8");
|
||||
|
||||
/**
|
||||
* Run the TS corpus and return:
|
||||
* { kind: "passed" } — all vectors pass (mutation survived)
|
||||
* { kind: "killed", output } — nonzero exit AND at least one expectedKiller ID
|
||||
* appears in stdout/stderr ("FAIL <id>" line)
|
||||
* { kind: "error", output } — nonzero exit but NO expected corpus output
|
||||
* (import error, missing file, syntax error, etc.)
|
||||
*/
|
||||
function runTsCorpus(expectedKillers) {
|
||||
let result;
|
||||
try {
|
||||
result = spawnSync(
|
||||
process.execPath,
|
||||
["--experimental-strip-types", jsRunnerPath],
|
||||
{ cwd: repoRoot, encoding: "utf8" },
|
||||
);
|
||||
} catch (e) {
|
||||
return { kind: "error", output: `spawn error: ${e.message}` };
|
||||
}
|
||||
if (result.status === 0) return { kind: "passed" };
|
||||
const output = (result.stdout ?? "") + (result.stderr ?? "");
|
||||
// A genuine corpus kill produces "FAIL <vector-id>" lines.
|
||||
// An infrastructure failure (import error, syntax error) produces no such lines.
|
||||
const hasCorpusFailure = expectedKillers.some((id) => output.includes(`FAIL ${id}`));
|
||||
if (hasCorpusFailure) return { kind: "killed", output };
|
||||
return { kind: "error", output };
|
||||
}
|
||||
|
||||
/**
|
||||
* Run the Rust corpus and return:
|
||||
* { kind: "passed" } — all vectors pass
|
||||
* { kind: "killed", output } — nonzero exit AND at least one expectedKiller ID
|
||||
* appears in the panic output ("[<id>]" format)
|
||||
* { kind: "error", output } — nonzero exit but NO expected corpus output
|
||||
* (compile error, missing cargo, linker error, etc.)
|
||||
*/
|
||||
function runRustCorpus(expectedKillers) {
|
||||
let result;
|
||||
try {
|
||||
result = spawnSync(
|
||||
"cargo",
|
||||
[
|
||||
"test",
|
||||
"-p", "buzz-agent",
|
||||
"--",
|
||||
"generated_model_capabilities::tests::shared_corpus_tests",
|
||||
"--nocapture",
|
||||
],
|
||||
{ cwd: repoRoot, encoding: "utf8", env: { ...process.env, RUST_BACKTRACE: "0" } },
|
||||
);
|
||||
} catch (e) {
|
||||
return { kind: "error", output: `spawn error: ${e.message}` };
|
||||
}
|
||||
if (result.status === 0) return { kind: "passed" };
|
||||
const output = (result.stdout ?? "") + (result.stderr ?? "");
|
||||
// The Rust corpus runner panics with "[<vector-id>] <axis>: got..." messages.
|
||||
const hasCorpusFailure = expectedKillers.some((id) => output.includes(`[${id}]`));
|
||||
if (hasCorpusFailure) return { kind: "killed", output };
|
||||
return { kind: "error", output };
|
||||
}
|
||||
|
||||
function regen() {
|
||||
execSync(`"${process.execPath}" "${generatorPath}"`, {
|
||||
cwd: repoRoot,
|
||||
stdio: VERBOSE ? "inherit" : "pipe",
|
||||
});
|
||||
}
|
||||
|
||||
function restore() {
|
||||
writeFileSync(manifestPath, originalManifest, "utf8");
|
||||
}
|
||||
|
||||
function applyMutation(mutFn) {
|
||||
const manifest = JSON.parse(originalManifest);
|
||||
mutFn(manifest);
|
||||
writeFileSync(manifestPath, JSON.stringify(manifest, null, 2) + "\n", "utf8");
|
||||
}
|
||||
|
||||
const mutations = [
|
||||
{
|
||||
id: "M1",
|
||||
description: "Reduce claude-opus-4-7 supported_efforts to [low,medium,high] (drops xhigh+max)",
|
||||
expectedKillers: ["anthropic-claude-opus-4-7", "dbv2-claude-prefix-stripped", "dbv2-claude-route-anthropic-messages"],
|
||||
mutate(manifest) {
|
||||
const rule = manifest.family_rules.find(r => r.id === "anthropic-adaptive-xhigh-opus-4-7");
|
||||
rule.supported_efforts = ["low", "medium", "high"];
|
||||
},
|
||||
},
|
||||
{
|
||||
id: "M2",
|
||||
description: "Add xhigh to gpt5-base supported_efforts [minimal,low,medium,high,xhigh]",
|
||||
expectedKillers: ["openai-gpt5-base", "openai-gpt5-1106-should-not-match-base", "openai-gpt5-4o-matches-base", "openai-gpt5-date-suffix"],
|
||||
mutate(manifest) {
|
||||
const rule = manifest.family_rules.find(r => r.id === "openai-gpt5-base");
|
||||
rule.supported_efforts = ["minimal", "low", "medium", "high", "xhigh"];
|
||||
},
|
||||
},
|
||||
{
|
||||
id: "M3",
|
||||
description: "Change gpt5-1 default_effort to 'high' instead of 'none'",
|
||||
expectedKillers: ["openai-gpt5.1"],
|
||||
mutate(manifest) {
|
||||
const rule = manifest.family_rules.find(r => r.id === "openai-gpt5-1");
|
||||
rule.default_effort = "high";
|
||||
},
|
||||
},
|
||||
{
|
||||
id: "M4",
|
||||
description: "Swap dbv2-claude-code-names-segment route from anthropic-messages to openai-responses",
|
||||
expectedKillers: ["dbv2-goose-opus-5-is-anthropic"],
|
||||
mutate(manifest) {
|
||||
const rule = manifest.family_rules.find(r => r.id === "dbv2-claude-code-names-segment");
|
||||
rule.databricks_v2_wire_route = "openai-responses";
|
||||
},
|
||||
},
|
||||
{
|
||||
id: "M5",
|
||||
description: "Remove all three DBv2 segment rules (dbv2-claude-code-names-segment, dbv2-gpt-code-names-segment, dbv2-sol-luna-terra-segment)",
|
||||
expectedKillers: ["dbv2-goose-opus-5-is-anthropic", "dbv2-consolidated-llama-not-sol", "dbv2-terraform-coder-not-terra"],
|
||||
mutate(manifest) {
|
||||
manifest.family_rules = manifest.family_rules.filter(
|
||||
r => !["dbv2-claude-code-names-segment", "dbv2-gpt-code-names-segment", "dbv2-sol-luna-terra-segment"].includes(r.id)
|
||||
);
|
||||
},
|
||||
},
|
||||
{
|
||||
id: "M6",
|
||||
description: "Change databricks_v2 concrete_unknown fallback route from mlflow-chat to openai-responses",
|
||||
expectedKillers: ["dbv2-concrete-unknown-mlflow-no-max"],
|
||||
mutate(manifest) {
|
||||
manifest.provider_fallbacks.databricks_v2.concrete_unknown.databricks_v2_wire_route = "openai-responses";
|
||||
},
|
||||
},
|
||||
{
|
||||
id: "M7",
|
||||
description: "Remove xhigh from gpt5-4 supported_efforts [none,low,medium,high]",
|
||||
expectedKillers: ["resolver-prefixed-alias-misses-exact"],
|
||||
mutate(manifest) {
|
||||
const rule = manifest.family_rules.find(r => r.id === "openai-gpt5-4");
|
||||
rule.supported_efforts = ["none", "low", "medium", "high"];
|
||||
},
|
||||
},
|
||||
];
|
||||
|
||||
let allKilled = true;
|
||||
const results = [];
|
||||
|
||||
for (const mut of mutations) {
|
||||
process.stdout.write(` ${mut.id}: ${mut.description}\n`);
|
||||
try {
|
||||
applyMutation(mut.mutate);
|
||||
regen();
|
||||
|
||||
// TS interpreter
|
||||
process.stdout.write(` TS ... `);
|
||||
const tsResult = runTsCorpus(mut.expectedKillers);
|
||||
const tsKilled = tsResult.kind === "killed";
|
||||
const tsError = tsResult.kind === "error";
|
||||
if (tsKilled) {
|
||||
process.stdout.write("killed ✓\n");
|
||||
} else if (tsError) {
|
||||
process.stdout.write(`ERROR (infrastructure failure — not a corpus kill)\n`);
|
||||
if (VERBOSE) process.stdout.write(` ${tsResult.output}\n`);
|
||||
} else {
|
||||
process.stdout.write("SURVIVED ✗\n");
|
||||
if (VERBOSE) process.stdout.write(` ${tsResult.output ?? ""}\n`);
|
||||
}
|
||||
|
||||
// Rust interpreter
|
||||
process.stdout.write(` Rust... `);
|
||||
const rustResult = runRustCorpus(mut.expectedKillers);
|
||||
const rustKilled = rustResult.kind === "killed";
|
||||
const rustError = rustResult.kind === "error";
|
||||
if (rustKilled) {
|
||||
process.stdout.write("killed ✓\n");
|
||||
} else if (rustError) {
|
||||
process.stdout.write(`ERROR (infrastructure failure — not a corpus kill)\n`);
|
||||
if (VERBOSE) process.stdout.write(` ${rustResult.output}\n`);
|
||||
} else {
|
||||
process.stdout.write("SURVIVED ✗\n");
|
||||
if (VERBOSE) process.stdout.write(` ${rustResult.output ?? ""}\n`);
|
||||
}
|
||||
|
||||
const killed = tsKilled && rustKilled;
|
||||
if (!killed) allKilled = false;
|
||||
results.push({ ...mut, killed, tsKilled, rustKilled, tsError, rustError });
|
||||
} catch (e) {
|
||||
process.stdout.write(` ERROR: ${e.message}\n`);
|
||||
allKilled = false;
|
||||
results.push({ ...mut, killed: false, tsKilled: false, rustKilled: false, tsError: true, rustError: true, output: e.message });
|
||||
} finally {
|
||||
restore();
|
||||
regen(); // restore generated files
|
||||
}
|
||||
}
|
||||
|
||||
console.log("");
|
||||
const killed = results.filter(r => r.killed).length;
|
||||
const errored = results.filter(r => r.tsError || r.rustError).length;
|
||||
console.log(`Mutation results: ${killed}/${results.length} killed (both interpreters)` +
|
||||
(errored > 0 ? `, ${errored} ERROR (infrastructure failure — see output above)` : ""));
|
||||
|
||||
if (!allKilled) {
|
||||
const hasErrors = results.some(r => r.tsError || r.rustError);
|
||||
if (hasErrors) {
|
||||
console.error("ERROR: Infrastructure failures prevented some mutations from being verified as killed.");
|
||||
console.error(" Run with --verbose to see the full output for ERROR entries.");
|
||||
}
|
||||
console.error("ERROR: Some mutations survived — corpus does not kill all resolver faults.");
|
||||
process.exit(1);
|
||||
}
|
||||
|
||||
console.log("All mutations killed in both TS and Rust interpreters.");
|
||||
@@ -0,0 +1,413 @@
|
||||
#!/usr/bin/env node
|
||||
/**
|
||||
* Schema-negative validator tests for the manifest generator.
|
||||
*
|
||||
* Every validator rule in generate-model-capabilities.mjs must have a
|
||||
* failing-input test here — if validation is missing, these tests would
|
||||
* not catch the defect.
|
||||
*
|
||||
* Uses Node.js built-in test runner (node --test).
|
||||
*/
|
||||
|
||||
import { test } from "node:test";
|
||||
import assert from "node:assert/strict";
|
||||
import { readFileSync, writeFileSync, unlinkSync, mkdirSync, mkdtempSync } from "node:fs";
|
||||
import { join, dirname } from "node:path";
|
||||
import { fileURLToPath } from "node:url";
|
||||
import { tmpdir } from "node:os";
|
||||
import { execFileSync, spawnSync } from "node:child_process";
|
||||
|
||||
const __dirname = dirname(fileURLToPath(import.meta.url));
|
||||
const repoRoot = join(__dirname, "..");
|
||||
const manifestPath = join(repoRoot, "scripts", "model-capabilities.json");
|
||||
const generatorPath = join(repoRoot, "scripts", "generate-model-capabilities.mjs");
|
||||
|
||||
/** Load the real manifest so we can mutate copies. */
|
||||
const BASE_MANIFEST = JSON.parse(readFileSync(manifestPath, "utf8"));
|
||||
|
||||
/**
|
||||
* Run the generator with a mutated manifest, returning { exitCode, stderr, stdout }.
|
||||
* Writes the mutated manifest to a temp file and overrides the manifest path via env.
|
||||
*/
|
||||
function runGeneratorWithManifest(manifestOverride) {
|
||||
// Write mutated manifest to a temp path
|
||||
const tmpDir = mkdtempSync(join(tmpdir(), "test-validator-"));
|
||||
const tmpManifest = join(tmpDir, "model-capabilities.json");
|
||||
const tmpOutputDir = join(tmpDir, "out");
|
||||
mkdirSync(tmpOutputDir, { recursive: true });
|
||||
|
||||
writeFileSync(tmpManifest, JSON.stringify(manifestOverride));
|
||||
|
||||
// Run the generator via node, pointing MANIFEST_PATH env at the temp file
|
||||
// The generator reads from process.env.MANIFEST_PATH if set (we add this support)
|
||||
const result = spawnSync(
|
||||
process.execPath,
|
||||
[generatorPath, "--manifest-path", tmpManifest, "--output-dir", tmpOutputDir],
|
||||
{
|
||||
encoding: "utf8",
|
||||
env: { ...process.env },
|
||||
},
|
||||
);
|
||||
|
||||
// Cleanup
|
||||
try { unlinkSync(tmpManifest); } catch {}
|
||||
|
||||
return { exitCode: result.status ?? 1, stderr: result.stderr, stdout: result.stdout };
|
||||
}
|
||||
|
||||
/**
|
||||
* Assert that the generator REJECTS the given manifest (exits non-zero).
|
||||
* The optional `expectedMessage` is checked in stderr if provided.
|
||||
*/
|
||||
function assertRejects(label, manifest, expectedMessage) {
|
||||
const { exitCode, stderr, stdout } = runGeneratorWithManifest(manifest);
|
||||
assert.notEqual(exitCode, 0, `${label}: expected generator to fail but it succeeded.\nstdout: ${stdout}\nstderr: ${stderr}`);
|
||||
if (expectedMessage) {
|
||||
const combined = stderr + stdout;
|
||||
assert.ok(
|
||||
combined.includes(expectedMessage),
|
||||
`${label}: expected error message "${expectedMessage}" not found.\nstdout: ${stdout}\nstderr: ${stderr}`,
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
/** Deep clone the base manifest and apply a mutator function. */
|
||||
function mutate(fn) {
|
||||
const clone = JSON.parse(JSON.stringify(BASE_MANIFEST));
|
||||
fn(clone);
|
||||
return clone;
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Rule: invalid enum value in family_rule.thinking_mode
|
||||
// ---------------------------------------------------------------------------
|
||||
test("schema-negative: invalid thinking_mode in family rule is rejected", () => {
|
||||
assertRejects(
|
||||
"invalid thinking_mode",
|
||||
mutate((m) => {
|
||||
m.family_rules[0].thinking_mode = "invalid-mode";
|
||||
}),
|
||||
"thinking_mode",
|
||||
);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Rule: invalid enum value in family_rule.databricks_v2_wire_route
|
||||
// ---------------------------------------------------------------------------
|
||||
test("schema-negative: invalid databricks_v2_wire_route in family rule is rejected", () => {
|
||||
assertRejects(
|
||||
"invalid databricks_v2_wire_route",
|
||||
mutate((m) => {
|
||||
m.family_rules[0].databricks_v2_wire_route = "chat-completions";
|
||||
}),
|
||||
"databricks_v2_wire_route",
|
||||
);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Rule: invalid enum value in family_rule.supported_efforts[]
|
||||
// ---------------------------------------------------------------------------
|
||||
test("schema-negative: invalid effort value in family rule supported_efforts is rejected", () => {
|
||||
assertRejects(
|
||||
"invalid supported_efforts value",
|
||||
mutate((m) => {
|
||||
m.family_rules[0].supported_efforts = ["low", "ultra-high"];
|
||||
}),
|
||||
"supported_efforts",
|
||||
);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Rule: empty supported_efforts array in family rule
|
||||
// ---------------------------------------------------------------------------
|
||||
test("schema-negative: empty supported_efforts in family rule is rejected", () => {
|
||||
assertRejects(
|
||||
"empty supported_efforts",
|
||||
mutate((m) => {
|
||||
m.family_rules[0].supported_efforts = [];
|
||||
}),
|
||||
"supported_efforts",
|
||||
);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Rule: default_effort not in supported_efforts (non-null)
|
||||
// ---------------------------------------------------------------------------
|
||||
test("schema-negative: default_effort not in supported_efforts is rejected", () => {
|
||||
assertRejects(
|
||||
"default_effort not in supported_efforts",
|
||||
mutate((m) => {
|
||||
m.family_rules[0].supported_efforts = ["low", "medium"];
|
||||
m.family_rules[0].default_effort = "high"; // not in list
|
||||
}),
|
||||
"default_effort",
|
||||
);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Rule: invalid normalization_policy in family rule
|
||||
// ---------------------------------------------------------------------------
|
||||
test("schema-negative: invalid normalization_policy in family rule is rejected", () => {
|
||||
assertRejects(
|
||||
"invalid normalization_policy",
|
||||
mutate((m) => {
|
||||
m.family_rules[0].normalization_policy = "pass-through-all";
|
||||
}),
|
||||
"normalization_policy",
|
||||
);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Rule: invalid match_kind in family rule
|
||||
// ---------------------------------------------------------------------------
|
||||
test("schema-negative: invalid match_kind in family rule is rejected", () => {
|
||||
assertRejects(
|
||||
"invalid match_kind",
|
||||
mutate((m) => {
|
||||
m.family_rules[0].match_kind = "regex";
|
||||
}),
|
||||
"match_kind",
|
||||
);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Rule: duplicate family rule id
|
||||
// ---------------------------------------------------------------------------
|
||||
test("schema-negative: duplicate family rule id is rejected", () => {
|
||||
assertRejects(
|
||||
"duplicate family rule id",
|
||||
mutate((m) => {
|
||||
m.family_rules.push({ ...m.family_rules[0] }); // duplicate id
|
||||
}),
|
||||
"duplicate",
|
||||
);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Rule: duplicate exact_record (provider, raw_model_id) key
|
||||
// ---------------------------------------------------------------------------
|
||||
test("schema-negative: duplicate exact_record key is rejected", () => {
|
||||
assertRejects(
|
||||
"duplicate exact_record key",
|
||||
mutate((m) => {
|
||||
m.exact_records.push({ ...m.exact_records[0] }); // duplicate
|
||||
}),
|
||||
"duplicate",
|
||||
);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Rule: exact_record missing provider
|
||||
// ---------------------------------------------------------------------------
|
||||
test("schema-negative: exact_record missing provider is rejected", () => {
|
||||
assertRejects(
|
||||
"exact_record missing provider",
|
||||
mutate((m) => {
|
||||
m.exact_records.push({ raw_model_id: "some-model" });
|
||||
}),
|
||||
"provider",
|
||||
);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Rule: exact_record missing raw_model_id
|
||||
// ---------------------------------------------------------------------------
|
||||
test("schema-negative: exact_record missing raw_model_id is rejected", () => {
|
||||
assertRejects(
|
||||
"exact_record missing raw_model_id",
|
||||
mutate((m) => {
|
||||
m.exact_records.push({ provider: "databricks_v2" });
|
||||
}),
|
||||
"raw_model_id",
|
||||
);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Rule: provider fallback record missing blank state
|
||||
// ---------------------------------------------------------------------------
|
||||
test("schema-negative: provider fallback missing blank state is rejected", () => {
|
||||
assertRejects(
|
||||
"provider fallback missing blank",
|
||||
mutate((m) => {
|
||||
delete m.provider_fallbacks.anthropic.blank;
|
||||
}),
|
||||
"blank",
|
||||
);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Rule: provider fallback record missing concrete_unknown state
|
||||
// ---------------------------------------------------------------------------
|
||||
test("schema-negative: provider fallback missing concrete_unknown state is rejected", () => {
|
||||
assertRejects(
|
||||
"provider fallback missing concrete_unknown",
|
||||
mutate((m) => {
|
||||
delete m.provider_fallbacks.anthropic.concrete_unknown;
|
||||
}),
|
||||
"concrete_unknown",
|
||||
);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Rule: invalid thinking_mode in provider fallback
|
||||
// ---------------------------------------------------------------------------
|
||||
test("schema-negative: invalid thinking_mode in provider fallback is rejected", () => {
|
||||
assertRejects(
|
||||
"invalid thinking_mode in fallback",
|
||||
mutate((m) => {
|
||||
m.provider_fallbacks.anthropic.blank.thinking_mode = "always-on";
|
||||
}),
|
||||
"thinking_mode",
|
||||
);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Rule: invalid databricks_v2_wire_route in provider fallback
|
||||
// ---------------------------------------------------------------------------
|
||||
test("schema-negative: invalid wire_route in provider fallback is rejected", () => {
|
||||
assertRejects(
|
||||
"invalid wire_route in fallback",
|
||||
mutate((m) => {
|
||||
m.provider_fallbacks.anthropic.blank.databricks_v2_wire_route = "http-sse";
|
||||
}),
|
||||
"databricks_v2_wire_route",
|
||||
);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Rule: invalid default_effort in provider fallback (not in supported_efforts)
|
||||
// ---------------------------------------------------------------------------
|
||||
test("schema-negative: default_effort not in supported_efforts in fallback is rejected", () => {
|
||||
assertRejects(
|
||||
"default_effort not in supported_efforts in fallback",
|
||||
mutate((m) => {
|
||||
m.provider_fallbacks.openai.blank.supported_efforts = ["low", "medium"];
|
||||
m.provider_fallbacks.openai.blank.default_effort = "high"; // not in list
|
||||
}),
|
||||
"default_effort",
|
||||
);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Rule: family rule missing id
|
||||
// ---------------------------------------------------------------------------
|
||||
test("schema-negative: family rule missing id is rejected", () => {
|
||||
assertRejects(
|
||||
"family rule missing id",
|
||||
mutate((m) => {
|
||||
m.family_rules.push({
|
||||
match_kind: "prefix",
|
||||
match_value: "test-",
|
||||
providers: ["anthropic"],
|
||||
match_priority: 1,
|
||||
thinking_mode: "none",
|
||||
supported_efforts: ["low"],
|
||||
default_effort: null,
|
||||
databricks_v2_wire_route: "not-applicable",
|
||||
normalization_policy: "none",
|
||||
// id deliberately omitted
|
||||
});
|
||||
}),
|
||||
"id",
|
||||
);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Rule: duplicate registry_label IDs
|
||||
// ---------------------------------------------------------------------------
|
||||
test("schema-negative: duplicate registry_label ID is rejected", () => {
|
||||
assertRejects(
|
||||
"duplicate registry_label ID",
|
||||
mutate((m) => {
|
||||
// Array format — duplicate id is structurally detectable
|
||||
m.registry_labels = [
|
||||
{ id: "databricks-gpt-5-5", label: "GPT-5.5" },
|
||||
{ id: "databricks-gpt-5-5", label: "GPT-5.5 duplicate" },
|
||||
];
|
||||
}),
|
||||
"duplicate",
|
||||
);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Rule: registry_label entry missing id (empty string)
|
||||
// ---------------------------------------------------------------------------
|
||||
test("schema-negative: registry_label entry with empty id is rejected", () => {
|
||||
assertRejects(
|
||||
"registry_label empty id",
|
||||
mutate((m) => {
|
||||
m.registry_labels = [{ id: "", label: "Some Label" }];
|
||||
}),
|
||||
"id",
|
||||
);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Rule: registry_label entry with unsafe characters in id
|
||||
// ---------------------------------------------------------------------------
|
||||
test("schema-negative: registry_label entry with unsafe id chars is rejected", () => {
|
||||
assertRejects(
|
||||
"registry_label unsafe id",
|
||||
mutate((m) => {
|
||||
m.registry_labels = [{ id: 'bad"id', label: "Some Label" }];
|
||||
}),
|
||||
"unsafe",
|
||||
);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Rule: duplicate databricks_v2_known_models IDs
|
||||
// ---------------------------------------------------------------------------
|
||||
test("schema-negative: duplicate databricks_v2_known_models ID is rejected", () => {
|
||||
assertRejects(
|
||||
"duplicate known model ID",
|
||||
mutate((m) => {
|
||||
m.databricks_v2_known_models = ["databricks-gpt-5-5", "databricks-gpt-5-5"];
|
||||
}),
|
||||
"duplicate",
|
||||
);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Rule: unsafe characters in match_value (family rule)
|
||||
// ---------------------------------------------------------------------------
|
||||
test("schema-negative: family rule match_value with unsafe chars is rejected", () => {
|
||||
assertRejects(
|
||||
"family rule match_value with backslash",
|
||||
mutate((m) => {
|
||||
// Inject a backslash into an existing rule's match_value — would break Rust string literal
|
||||
const rule = m.family_rules.find((r) => r.id === "anthropic-manual-budget-claude3");
|
||||
rule.match_value = "claude-3\\evil";
|
||||
}),
|
||||
"unsafe",
|
||||
);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Rule: unsafe characters in known-model ID
|
||||
// ---------------------------------------------------------------------------
|
||||
test("schema-negative: databricks_v2_known_models ID with unsafe chars is rejected", () => {
|
||||
assertRejects(
|
||||
"known-model ID with double-quote",
|
||||
mutate((m) => {
|
||||
m.databricks_v2_known_models = ['databricks-gpt-5-5', 'bad"id'];
|
||||
}),
|
||||
"unsafe",
|
||||
);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Rule: unsafe characters in exact_record registry_label
|
||||
// ---------------------------------------------------------------------------
|
||||
test("schema-negative: exact_record registry_label with unsafe chars is rejected", () => {
|
||||
assertRejects(
|
||||
"exact_record registry_label with backslash",
|
||||
mutate((m) => {
|
||||
const rec = m.exact_records.find((r) => r.raw_model_id === "databricks-gpt-5-4-mini");
|
||||
rec.registry_label = "GPT-5.4 Mini\\injected";
|
||||
}),
|
||||
"unsafe",
|
||||
);
|
||||
});
|
||||
|
||||
console.log("\nSchema-negative validator tests complete.");
|
||||
Reference in New Issue
Block a user