Commit Graph
4 Commits
Author SHA1 Message Date
npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7andWill Pfleger 6301594dc6 fix(models): close round-2 gaps: 6-axis Rust harness + gpt-neox honesty
Co-authored-by: Will Pfleger <pfleger.will@gmail.com>
Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
2026-08-04 15:29:14 -04:00
npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7andWill Pfleger 8ba3814854 fix(models): address Kalvin-review: gpt-segment/left-boundary/luna-terra/validator
CRITICAL: Fix Rust gpt5-base short-version guard to mirror TS regex semantics.
The old 4-char window check diverged from TS /^\d{1,3}(?:[^a-z\d]|$)/i for
inputs like gpt-5-10-preview (10 digits). Replaced with find(non-digit) approach
that exactly matches the TS digit-run semantics. Add left-boundary guards (start
or preceded by '-'/'.') to both gpt5_token_matches_rs and gpt5_base_matches_rs
to prevent sgpt-model-class false-positives.

IMPORTANT: Add exact_records validation to generator — enum checks for all
optional override axes, non-empty override check, canonical-order enforcement,
default_effort-in-override check, match_priority non-negative-int check
(injected into source comments, treated as injection surface). Add 9 schema-
negative tests in test-manifest-validator.mjs.

IMPORTANT: Restrict DBv2 gpt segment rule from starts_with("gpt") to exact
segment match ("gpt" or "gpt5"). Raise priority 5->6 to restore old dual-marker
OpenAI-before-Claude contract. Add left-boundary guards to token helpers. Add
collision-negative corpus vectors: gptoss-model, gptj-6b, gpt-neox,
customgpt-5-5-endpoint.

MINOR: Restore .trim() on model string before resolveModelCapabilities call in
buzzAgentConfig.ts. Make exact-record lookup case-insensitive (lowercase keys at
build + lookup). Extend run-corpus.mjs to compare all 6 axes including
normalization_policy and registry_label. Add positive corpus vectors: opus-5
rule, dbv2 gpt-segment, sol/luna/terra, gpt-5-4-nano exact record, openrouter/
unknown/legacy-databricks fallbacks, case-insensitive lookup vectors. Fix Rust
corpus test harness provider lowercasing to handle vectors like provider=OpenAI.

HYGIENE: Fix "Generated 3 files" -> "Generated 2 files" message. Remove dead
$schema pointer from manifest. Add module doc to generated_model_capabilities_tests.rs
noting hand-maintained status. Switch PROVIDER_ALIASES from plain object to Map
in formatAgentModelLabel.ts and run-corpus.mjs. Align CI node-version to 24.
Tighten python SAFE_NAME_RE from * to + to reject empty display names.

luna/terra: models.dev confirms databricks-gpt-5-6-luna and -terra advertise
[low,medium,high] (differs from sol's [low,medium,high,max]). Added two exact
records with supported_efforts=[low,medium,high], default_effort=medium.

Co-authored-by: Will Pfleger <pfleger.will@gmail.com>
Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
2026-08-04 13:15:11 -04:00
cc00060ea3 feat(agent): Phase 2 — wire Rust and TS consumers to generated model-capabilities module (#3958)
## What

Phase 2 consumer cutover targeting the
`duncan/databricks-model-label-registry` umbrella branch. Wires
`crates/**` and `desktop/**` consumers to the generated capability
module introduced in Phase 1 (#3821), while keeping old and new paths
both live for differential testing. Phase 3 removes the old paths.

## Commits (boundary-separated)

### feat(agent): Phase 2a — wire Rust consumers to generated capability
module (`crates/**`, `scripts/**`)

- `catalog.rs`: `DATABRICKS_V2_KNOWN_MODELS` re-exported from the
generated module — single source of truth.
- `llm.rs`: `databricks_v2_route_for_model` delegates to
`resolve_model_capabilities("databricks_v2", model)`. Old segment-based
classifier preserved as `#[cfg(test)] _old_*` for the differential
harness. New `databricks_v2_route_differential_old_vs_new` test confirms
100% agreement on all 20 route vectors.
- `config.rs`: new `effort_table_fixture_differential_old_vs_new` test
runs `resolve_model_capabilities` over the 36-entry
`effortTable.fixture.json` and asserts old/new agree modulo a doc-cited
allowlist (4 F1 corrections).
- `scripts/run-differential.mjs`: JS differential harness over
effortTable fixture + normative corpus + catalog-sample fixture. 85
checks, 0 unexpected divergences (5 allowlisted: 4 F1 corrections +
goose-opus-5 anthropic route correction).
- `scripts/MODELS_DEV_RECONCILIATION.md`: deferred MINOR from Phase 1 —
8 trailing-double-space line breaks replaced with `<br>`.

### feat(desktop): Phase 2b — cut TS consumers to generated
model-capabilities module (`desktop/**`)

- `buzzAgentConfig.ts`: adds
`getProviderEffortConfigFromManifest(provider, model?)` — thin wrapper
over `resolveModelCapabilities()` from `modelCapabilities.ts`. Maps
`supportedEfforts → validValues` and `defaultEffort → defaultValue`
(null preserved for manual-budget/Inherit). Old
`getProviderEffortConfig()` and all hand-tables stay live for the
differential harness; Phase 3 retires them.
- `formatAgentModelLabel.ts`: registry-label lookup re-pointed from
hand-maintained `databricksModelNames.ts` import to generated
`DATABRICKS_MODEL_NAMES` exported from `modelCapabilities.ts`. Same Map
shape, identical contents, behavior unchanged.

### fix(scripts): add ts-esm-loader and fix allowlist coverage in
run-differential (`scripts/**`)

- `scripts/ts-esm-loader.mjs`: minimal ESM custom loader that resolves
extensionless relative TS imports. Required because Phase 2b's
`buzzAgentConfig.ts` imports `modelCapabilities` without `.ts` extension
— which Node's `--experimental-strip-types` runner cannot resolve
without a hook.
- `scripts/run-differential.mjs`: shebang updated to self-bootstrap with
the loader; fixes the `totalAllowlisted` counter (was declared but never
incremented — always printed `0 allowlisted`). Replaced with per-axis
hit tracking: reports exercised slot count (`N/total`) in summary; fails
with `STALE_ALLOWLIST` if any declared entry fires zero divergences,
preventing stale entries from silently masking future regressions.

## Verification

- `cargo test -p buzz-agent --lib`: 426/426
- Corpus: 45/45 · schema-negative: 24/24 · `--check` byte-clean
- Differential: 85 checks, 0 unexpected divergences, 6/6 allowlist slots
exercised
- Desktop: 3847/3847 · typecheck clean · biome clean
- Mobile: 1019 pass, 1 skipped — same 5 flaky tests in
`mobile/test/features/channels/` that reproduce at the umbrella base;
zero mobile files in this branch range
- `git diff --check`: clean

## What Remains (Phase 3)

Remove old hand-maintained paths: `_old_*` functions in
`llm.rs`/`config.rs`, old `getProviderEffortConfig` tables in
`buzzAgentConfig.ts`, old `databricksModelNames.ts` import in
`formatAgentModelLabel.ts`, old `databricks_model_names.rs` module.

---------

Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
Signed-off-by: npub1g8493u0xfsjrvflg4n08ezd7vec99mnwzlv0qgwpr9d7gvjwhuzqx59rhw <41ea58f1e64c243627e8acde7c89be667052ee6e17d8f021c1195be4324ebf04@buzz.block.builderlab.xyz>
Signed-off-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
Co-authored-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
Co-authored-by: npub1g8493u0xfsjrvflg4n08ezd7vec99mnwzlv0qgwpr9d7gvjwhuzqx59rhw <41ea58f1e64c243627e8acde7c89be667052ee6e17d8f021c1195be4324ebf04@buzz.block.builderlab.xyz>
2026-08-03 14:58:34 -04:00
4d47f48143 feat(agent): Phase 1 — model-capability manifest, generator, and test oracle (#3821)
## What this does

Introduces the model-capability manifest infrastructure (Phase 1 of the
Model-Capability Manifest plan v4, Thufir-approved 9/9/9). No consumer
cutover — `config.rs`, `llm.rs`, `catalog.rs`, and `buzzAgentConfig.ts`
are unchanged. Phase 2 wires them.

**Single source of truth** replaces hand-mirrored metadata across four
files:

```
scripts/model-capabilities.json         → hand-curated manifest
scripts/generate-model-capabilities.mjs → emits Rust + TS artifacts
crates/buzz-agent/src/generated_model_capabilities.rs
desktop/src/features/agents/ui/modelCapabilities.ts
```

## Resolver contract (plan v4 §Resolver contract)

Total function `resolve(provider, raw_model_id) → CapabilityResult`.
Three ordered steps:

1. Provider-qualified raw exact lookup — key is `(provider,
raw_model_id)`, matched before any prefix stripping. A prefixed alias
never inherits an exact record.
2. Provider-scoped ordered family rules — on normalized
(prefix-stripped) alias, by `match_priority` desc.
3. Per-axis provider fallback — `blank` vs `concrete_unknown`, per
provider.

Result is complete — every axis populated, runtime consumers never
compose fields.

## Boundaries the manifest does NOT own

- Transport for pure OpenAI, legacy Databricks, OpenRouter:
`OpenAiApi`/`openai_request()` remain authoritative.
`databricks_v2_wire_route` is DBv2-only (all other providers emit
`not-applicable`).
- Final display labels: `resolveModelLabel()` three-tier precedence
unchanged. `registry_label` feeds only the static registry tier.
- `llm.rs` scope: only `databricks_v2_route_for_model` (Phase 2).

## Test oracle (three independent layers)

1. Generated full-table coverage —
`scripts/generated-model-capabilities-coverage.json`: every manifest
entry + provider fallbacks.
2. Hand-authored normative corpus — `scripts/normative-corpus.json` (44
vectors): Anthropic manual-budget/adaptive families, OpenAI gpt-5
adversarial boundary cases, DBv2 segment-routing collision tests, P2-A
resolver-contract vectors, P2-B blank/concrete-unknown per provider.
Runs against JS resolver (`run-corpus.mjs`) and mirrored in Rust
(`generated_model_capabilities_tests.rs`).
3. Schema-negative tests — `scripts/test-manifest-validator.mjs` (17
tests): every validator rule has a failing-input test.

## Reconciliation table

`scripts/MODELS_DEV_RECONCILIATION.md` — all models.dev divergences
dispositioned. `databricks-gpt-5-4-mini` and `databricks-gpt-5-4-nano`
adopt models.dev `[low,medium,high]` (family rule adds `none+xhigh` the
endpoint doesn't advertise).

## CI

`.github/workflows/model-capability-regen-diff.yml`: triggers on
manifest/generator/artifact changes; regenerates and fails if stale;
runs JS corpus + schema-negative tests.

## Acceptance criteria (plan v4 Phase 1)

- Byte-clean regen: `node scripts/generate-model-capabilities.mjs
--check` passes
- Rust compiles: `cargo check -p buzz-agent`
- TS typechecks: `pnpm tsc --noEmit --strict`
- 44/44 normative corpus vectors pass (JS interpreter)
- 41/41 Rust corpus tests pass
- 17/17 schema-negative tests pass (every validator rule)
- Reconciliation table complete with doc citations
- No consumer changes (config.rs, llm.rs, catalog.rs, buzzAgentConfig.ts
untouched)

---------

Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
Co-authored-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
2026-07-31 12:39:37 -04:00