mirror of
https://github.com/block/buzz.git
synced 2026-08-18 06:50:31 +02:00
6301594dc621c033ea10cd3b083b06f32ba9259e
4
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6301594dc6 |
fix(models): close round-2 gaps: 6-axis Rust harness + gpt-neox honesty
Co-authored-by: Will Pfleger <pfleger.will@gmail.com> Signed-off-by: Will Pfleger <pfleger.will@gmail.com> |
||
|
|
8ba3814854 |
fix(models): address Kalvin-review: gpt-segment/left-boundary/luna-terra/validator
CRITICAL: Fix Rust gpt5-base short-version guard to mirror TS regex semantics.
The old 4-char window check diverged from TS /^\d{1,3}(?:[^a-z\d]|$)/i for
inputs like gpt-5-10-preview (10 digits). Replaced with find(non-digit) approach
that exactly matches the TS digit-run semantics. Add left-boundary guards (start
or preceded by '-'/'.') to both gpt5_token_matches_rs and gpt5_base_matches_rs
to prevent sgpt-model-class false-positives.
IMPORTANT: Add exact_records validation to generator — enum checks for all
optional override axes, non-empty override check, canonical-order enforcement,
default_effort-in-override check, match_priority non-negative-int check
(injected into source comments, treated as injection surface). Add 9 schema-
negative tests in test-manifest-validator.mjs.
IMPORTANT: Restrict DBv2 gpt segment rule from starts_with("gpt") to exact
segment match ("gpt" or "gpt5"). Raise priority 5->6 to restore old dual-marker
OpenAI-before-Claude contract. Add left-boundary guards to token helpers. Add
collision-negative corpus vectors: gptoss-model, gptj-6b, gpt-neox,
customgpt-5-5-endpoint.
MINOR: Restore .trim() on model string before resolveModelCapabilities call in
buzzAgentConfig.ts. Make exact-record lookup case-insensitive (lowercase keys at
build + lookup). Extend run-corpus.mjs to compare all 6 axes including
normalization_policy and registry_label. Add positive corpus vectors: opus-5
rule, dbv2 gpt-segment, sol/luna/terra, gpt-5-4-nano exact record, openrouter/
unknown/legacy-databricks fallbacks, case-insensitive lookup vectors. Fix Rust
corpus test harness provider lowercasing to handle vectors like provider=OpenAI.
HYGIENE: Fix "Generated 3 files" -> "Generated 2 files" message. Remove dead
$schema pointer from manifest. Add module doc to generated_model_capabilities_tests.rs
noting hand-maintained status. Switch PROVIDER_ALIASES from plain object to Map
in formatAgentModelLabel.ts and run-corpus.mjs. Align CI node-version to 24.
Tighten python SAFE_NAME_RE from * to + to reject empty display names.
luna/terra: models.dev confirms databricks-gpt-5-6-luna and -terra advertise
[low,medium,high] (differs from sol's [low,medium,high,max]). Added two exact
records with supported_efforts=[low,medium,high], default_effort=medium.
Co-authored-by: Will Pfleger <pfleger.will@gmail.com>
Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
|
||
|
|
cc00060ea3 |
feat(agent): Phase 2 — wire Rust and TS consumers to generated model-capabilities module (#3958)
## What Phase 2 consumer cutover targeting the `duncan/databricks-model-label-registry` umbrella branch. Wires `crates/**` and `desktop/**` consumers to the generated capability module introduced in Phase 1 (#3821), while keeping old and new paths both live for differential testing. Phase 3 removes the old paths. ## Commits (boundary-separated) ### feat(agent): Phase 2a — wire Rust consumers to generated capability module (`crates/**`, `scripts/**`) - `catalog.rs`: `DATABRICKS_V2_KNOWN_MODELS` re-exported from the generated module — single source of truth. - `llm.rs`: `databricks_v2_route_for_model` delegates to `resolve_model_capabilities("databricks_v2", model)`. Old segment-based classifier preserved as `#[cfg(test)] _old_*` for the differential harness. New `databricks_v2_route_differential_old_vs_new` test confirms 100% agreement on all 20 route vectors. - `config.rs`: new `effort_table_fixture_differential_old_vs_new` test runs `resolve_model_capabilities` over the 36-entry `effortTable.fixture.json` and asserts old/new agree modulo a doc-cited allowlist (4 F1 corrections). - `scripts/run-differential.mjs`: JS differential harness over effortTable fixture + normative corpus + catalog-sample fixture. 85 checks, 0 unexpected divergences (5 allowlisted: 4 F1 corrections + goose-opus-5 anthropic route correction). - `scripts/MODELS_DEV_RECONCILIATION.md`: deferred MINOR from Phase 1 — 8 trailing-double-space line breaks replaced with `<br>`. ### feat(desktop): Phase 2b — cut TS consumers to generated model-capabilities module (`desktop/**`) - `buzzAgentConfig.ts`: adds `getProviderEffortConfigFromManifest(provider, model?)` — thin wrapper over `resolveModelCapabilities()` from `modelCapabilities.ts`. Maps `supportedEfforts → validValues` and `defaultEffort → defaultValue` (null preserved for manual-budget/Inherit). Old `getProviderEffortConfig()` and all hand-tables stay live for the differential harness; Phase 3 retires them. - `formatAgentModelLabel.ts`: registry-label lookup re-pointed from hand-maintained `databricksModelNames.ts` import to generated `DATABRICKS_MODEL_NAMES` exported from `modelCapabilities.ts`. Same Map shape, identical contents, behavior unchanged. ### fix(scripts): add ts-esm-loader and fix allowlist coverage in run-differential (`scripts/**`) - `scripts/ts-esm-loader.mjs`: minimal ESM custom loader that resolves extensionless relative TS imports. Required because Phase 2b's `buzzAgentConfig.ts` imports `modelCapabilities` without `.ts` extension — which Node's `--experimental-strip-types` runner cannot resolve without a hook. - `scripts/run-differential.mjs`: shebang updated to self-bootstrap with the loader; fixes the `totalAllowlisted` counter (was declared but never incremented — always printed `0 allowlisted`). Replaced with per-axis hit tracking: reports exercised slot count (`N/total`) in summary; fails with `STALE_ALLOWLIST` if any declared entry fires zero divergences, preventing stale entries from silently masking future regressions. ## Verification - `cargo test -p buzz-agent --lib`: 426/426 - Corpus: 45/45 · schema-negative: 24/24 · `--check` byte-clean - Differential: 85 checks, 0 unexpected divergences, 6/6 allowlist slots exercised - Desktop: 3847/3847 · typecheck clean · biome clean - Mobile: 1019 pass, 1 skipped — same 5 flaky tests in `mobile/test/features/channels/` that reproduce at the umbrella base; zero mobile files in this branch range - `git diff --check`: clean ## What Remains (Phase 3) Remove old hand-maintained paths: `_old_*` functions in `llm.rs`/`config.rs`, old `getProviderEffortConfig` tables in `buzzAgentConfig.ts`, old `databricksModelNames.ts` import in `formatAgentModelLabel.ts`, old `databricks_model_names.rs` module. --------- Signed-off-by: Will Pfleger <pfleger.will@gmail.com> Signed-off-by: npub1g8493u0xfsjrvflg4n08ezd7vec99mnwzlv0qgwpr9d7gvjwhuzqx59rhw <41ea58f1e64c243627e8acde7c89be667052ee6e17d8f021c1195be4324ebf04@buzz.block.builderlab.xyz> Signed-off-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz> Co-authored-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz> Co-authored-by: npub1g8493u0xfsjrvflg4n08ezd7vec99mnwzlv0qgwpr9d7gvjwhuzqx59rhw <41ea58f1e64c243627e8acde7c89be667052ee6e17d8f021c1195be4324ebf04@buzz.block.builderlab.xyz> |
||
|
|
4d47f48143 |
feat(agent): Phase 1 — model-capability manifest, generator, and test oracle (#3821)
## What this does Introduces the model-capability manifest infrastructure (Phase 1 of the Model-Capability Manifest plan v4, Thufir-approved 9/9/9). No consumer cutover — `config.rs`, `llm.rs`, `catalog.rs`, and `buzzAgentConfig.ts` are unchanged. Phase 2 wires them. **Single source of truth** replaces hand-mirrored metadata across four files: ``` scripts/model-capabilities.json → hand-curated manifest scripts/generate-model-capabilities.mjs → emits Rust + TS artifacts crates/buzz-agent/src/generated_model_capabilities.rs desktop/src/features/agents/ui/modelCapabilities.ts ``` ## Resolver contract (plan v4 §Resolver contract) Total function `resolve(provider, raw_model_id) → CapabilityResult`. Three ordered steps: 1. Provider-qualified raw exact lookup — key is `(provider, raw_model_id)`, matched before any prefix stripping. A prefixed alias never inherits an exact record. 2. Provider-scoped ordered family rules — on normalized (prefix-stripped) alias, by `match_priority` desc. 3. Per-axis provider fallback — `blank` vs `concrete_unknown`, per provider. Result is complete — every axis populated, runtime consumers never compose fields. ## Boundaries the manifest does NOT own - Transport for pure OpenAI, legacy Databricks, OpenRouter: `OpenAiApi`/`openai_request()` remain authoritative. `databricks_v2_wire_route` is DBv2-only (all other providers emit `not-applicable`). - Final display labels: `resolveModelLabel()` three-tier precedence unchanged. `registry_label` feeds only the static registry tier. - `llm.rs` scope: only `databricks_v2_route_for_model` (Phase 2). ## Test oracle (three independent layers) 1. Generated full-table coverage — `scripts/generated-model-capabilities-coverage.json`: every manifest entry + provider fallbacks. 2. Hand-authored normative corpus — `scripts/normative-corpus.json` (44 vectors): Anthropic manual-budget/adaptive families, OpenAI gpt-5 adversarial boundary cases, DBv2 segment-routing collision tests, P2-A resolver-contract vectors, P2-B blank/concrete-unknown per provider. Runs against JS resolver (`run-corpus.mjs`) and mirrored in Rust (`generated_model_capabilities_tests.rs`). 3. Schema-negative tests — `scripts/test-manifest-validator.mjs` (17 tests): every validator rule has a failing-input test. ## Reconciliation table `scripts/MODELS_DEV_RECONCILIATION.md` — all models.dev divergences dispositioned. `databricks-gpt-5-4-mini` and `databricks-gpt-5-4-nano` adopt models.dev `[low,medium,high]` (family rule adds `none+xhigh` the endpoint doesn't advertise). ## CI `.github/workflows/model-capability-regen-diff.yml`: triggers on manifest/generator/artifact changes; regenerates and fails if stale; runs JS corpus + schema-negative tests. ## Acceptance criteria (plan v4 Phase 1) - Byte-clean regen: `node scripts/generate-model-capabilities.mjs --check` passes - Rust compiles: `cargo check -p buzz-agent` - TS typechecks: `pnpm tsc --noEmit --strict` - 44/44 normative corpus vectors pass (JS interpreter) - 41/41 Rust corpus tests pass - 17/17 schema-negative tests pass (every validator rule) - Reconciliation table complete with doc citations - No consumer changes (config.rs, llm.rs, catalog.rs, buzzAgentConfig.ts untouched) --------- Signed-off-by: Will Pfleger <pfleger.will@gmail.com> Co-authored-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz> |