mirror of
https://github.com/block/buzz.git
synced 2026-08-18 06:50:31 +02:00
## What Phase 2 consumer cutover targeting the `duncan/databricks-model-label-registry` umbrella branch. Wires `crates/**` and `desktop/**` consumers to the generated capability module introduced in Phase 1 (#3821), while keeping old and new paths both live for differential testing. Phase 3 removes the old paths. ## Commits (boundary-separated) ### feat(agent): Phase 2a — wire Rust consumers to generated capability module (`crates/**`, `scripts/**`) - `catalog.rs`: `DATABRICKS_V2_KNOWN_MODELS` re-exported from the generated module — single source of truth. - `llm.rs`: `databricks_v2_route_for_model` delegates to `resolve_model_capabilities("databricks_v2", model)`. Old segment-based classifier preserved as `#[cfg(test)] _old_*` for the differential harness. New `databricks_v2_route_differential_old_vs_new` test confirms 100% agreement on all 20 route vectors. - `config.rs`: new `effort_table_fixture_differential_old_vs_new` test runs `resolve_model_capabilities` over the 36-entry `effortTable.fixture.json` and asserts old/new agree modulo a doc-cited allowlist (4 F1 corrections). - `scripts/run-differential.mjs`: JS differential harness over effortTable fixture + normative corpus + catalog-sample fixture. 85 checks, 0 unexpected divergences (5 allowlisted: 4 F1 corrections + goose-opus-5 anthropic route correction). - `scripts/MODELS_DEV_RECONCILIATION.md`: deferred MINOR from Phase 1 — 8 trailing-double-space line breaks replaced with `<br>`. ### feat(desktop): Phase 2b — cut TS consumers to generated model-capabilities module (`desktop/**`) - `buzzAgentConfig.ts`: adds `getProviderEffortConfigFromManifest(provider, model?)` — thin wrapper over `resolveModelCapabilities()` from `modelCapabilities.ts`. Maps `supportedEfforts → validValues` and `defaultEffort → defaultValue` (null preserved for manual-budget/Inherit). Old `getProviderEffortConfig()` and all hand-tables stay live for the differential harness; Phase 3 retires them. - `formatAgentModelLabel.ts`: registry-label lookup re-pointed from hand-maintained `databricksModelNames.ts` import to generated `DATABRICKS_MODEL_NAMES` exported from `modelCapabilities.ts`. Same Map shape, identical contents, behavior unchanged. ### fix(scripts): add ts-esm-loader and fix allowlist coverage in run-differential (`scripts/**`) - `scripts/ts-esm-loader.mjs`: minimal ESM custom loader that resolves extensionless relative TS imports. Required because Phase 2b's `buzzAgentConfig.ts` imports `modelCapabilities` without `.ts` extension — which Node's `--experimental-strip-types` runner cannot resolve without a hook. - `scripts/run-differential.mjs`: shebang updated to self-bootstrap with the loader; fixes the `totalAllowlisted` counter (was declared but never incremented — always printed `0 allowlisted`). Replaced with per-axis hit tracking: reports exercised slot count (`N/total`) in summary; fails with `STALE_ALLOWLIST` if any declared entry fires zero divergences, preventing stale entries from silently masking future regressions. ## Verification - `cargo test -p buzz-agent --lib`: 426/426 - Corpus: 45/45 · schema-negative: 24/24 · `--check` byte-clean - Differential: 85 checks, 0 unexpected divergences, 6/6 allowlist slots exercised - Desktop: 3847/3847 · typecheck clean · biome clean - Mobile: 1019 pass, 1 skipped — same 5 flaky tests in `mobile/test/features/channels/` that reproduce at the umbrella base; zero mobile files in this branch range - `git diff --check`: clean ## What Remains (Phase 3) Remove old hand-maintained paths: `_old_*` functions in `llm.rs`/`config.rs`, old `getProviderEffortConfig` tables in `buzzAgentConfig.ts`, old `databricksModelNames.ts` import in `formatAgentModelLabel.ts`, old `databricks_model_names.rs` module. --------- Signed-off-by: Will Pfleger <pfleger.will@gmail.com> Signed-off-by: npub1g8493u0xfsjrvflg4n08ezd7vec99mnwzlv0qgwpr9d7gvjwhuzqx59rhw <41ea58f1e64c243627e8acde7c89be667052ee6e17d8f021c1195be4324ebf04@buzz.block.builderlab.xyz> Signed-off-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz> Co-authored-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz> Co-authored-by: npub1g8493u0xfsjrvflg4n08ezd7vec99mnwzlv0qgwpr9d7gvjwhuzqx59rhw <41ea58f1e64c243627e8acde7c89be667052ee6e17d8f021c1195be4324ebf04@buzz.block.builderlab.xyz>
240 lines
8.3 KiB
JavaScript
Executable File
240 lines
8.3 KiB
JavaScript
Executable File
#!/usr/bin/env node
|
|
/**
|
|
* Phase-2 differential harness — compare old buzzAgentConfig.ts effort logic with
|
|
* the new generated modelCapabilities.ts interpreter over:
|
|
* 1. The 36-entry effortTable.fixture.json (cross-boundary Rust/TS fixture)
|
|
* 2. The 45-vector normative corpus (scripts/normative-corpus.json)
|
|
* 3. The catalog-sample fixture (scripts/catalog-sample-fixture.json)
|
|
*
|
|
* Equality is required except for entries in the committed allowlist of intentional
|
|
* F1 corrections (models.dev provider-capability reconciliations).
|
|
*
|
|
* Usage: node --experimental-strip-types scripts/run-differential.mjs [--verbose]
|
|
* Exits 0 on all-pass (modulo allowlist), 1 on unexpected divergence or unexercised allowlist entry.
|
|
*/
|
|
|
|
import { readFileSync } from "node:fs";
|
|
import { join, dirname } from "node:path";
|
|
import { fileURLToPath } from "node:url";
|
|
|
|
const __dirname = dirname(fileURLToPath(import.meta.url));
|
|
const repoRoot = join(__dirname, "..");
|
|
const VERBOSE = process.argv.includes("--verbose");
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Import both interpreters
|
|
// ---------------------------------------------------------------------------
|
|
|
|
// NEW: generated capability module
|
|
const { resolveModelCapabilities: resolveNew } = await import(
|
|
join(repoRoot, "desktop", "src", "features", "agents", "ui", "modelCapabilities.ts")
|
|
);
|
|
|
|
// OLD: buzzAgentConfig.ts effort config
|
|
const { getProviderEffortConfig_oldHandTable: getOldEffortConfig } = await import(
|
|
join(repoRoot, "desktop", "src", "features", "agents", "ui", "buzzAgentConfig.ts")
|
|
);
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Intentional corrections allowlist (Phase 1 F1 reconciliations)
|
|
// Each entry: { provider, raw_model_id, reason }
|
|
// ---------------------------------------------------------------------------
|
|
const ALLOWLIST = [
|
|
{
|
|
provider: "databricks_v2",
|
|
raw_model_id: "databricks-gpt-5-5",
|
|
axes: ["supported_efforts"],
|
|
reason: "Phase 1 ADOPT: models.dev d5a4974c advertises [low,medium,high]; old returns [none,low,medium,high,xhigh]",
|
|
},
|
|
{
|
|
provider: "databricks_v2",
|
|
raw_model_id: "databricks-gpt-5-4-mini",
|
|
axes: ["supported_efforts"],
|
|
reason: "Phase 1 ADOPT: models.dev advertises [low,medium,high]; old returns [none,low,medium,high,xhigh]",
|
|
},
|
|
{
|
|
provider: "databricks_v2",
|
|
raw_model_id: "databricks-gpt-5-4-nano",
|
|
axes: ["supported_efforts"],
|
|
reason: "Phase 1 ADOPT: models.dev advertises [low,medium,high]; old returns [none,low,medium,high,xhigh]",
|
|
},
|
|
{
|
|
provider: "databricks_v2",
|
|
raw_model_id: "databricks-gpt-5-6-sol",
|
|
axes: ["supported_efforts"],
|
|
reason: "Phase 1 ADOPT: models.dev advertises [low,medium,high,max]; old returns [none,low,medium,high,xhigh,max]",
|
|
},
|
|
{
|
|
provider: "databricks_v2",
|
|
raw_model_id: "goose-opus-5",
|
|
axes: ["supported_efforts", "default_effort"],
|
|
reason: "Phase 1 correction: 'opus' is a named DBv2 segment → anthropic-messages route; old config.rs disagreed with llm.rs (corpus note dbv2-goose-opus-5-is-anthropic). Generated adopts anthropic adaptive-xhigh capabilities consistent with the wire route.",
|
|
},
|
|
];
|
|
|
|
// Track which allowlist entries are actually exercised (suppressed a divergence).
|
|
// Keyed as "provider:raw_model_id:axis".
|
|
const allowlistHits = new Set();
|
|
|
|
function isAllowlisted(provider, rawModelId, axis) {
|
|
const entry = ALLOWLIST.find(
|
|
(e) =>
|
|
e.provider === provider &&
|
|
e.raw_model_id === rawModelId &&
|
|
e.axes.includes(axis),
|
|
);
|
|
if (entry) {
|
|
allowlistHits.add(`${provider}:${rawModelId}:${axis}`);
|
|
return true;
|
|
}
|
|
return false;
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Comparison helpers
|
|
// ---------------------------------------------------------------------------
|
|
|
|
/**
|
|
* Compare effort axes from both interpreters for one (provider, model) pair.
|
|
* Returns array of divergence objects.
|
|
*/
|
|
function compareEffortAxes(provider, model) {
|
|
const newResult = resolveNew(provider, model);
|
|
const oldResult = getOldEffortConfig(provider, model);
|
|
|
|
const divergences = [];
|
|
|
|
// supported_efforts
|
|
const newEfforts = newResult.supportedEfforts ?? [];
|
|
const oldEfforts = oldResult?.validValues ?? [];
|
|
if (JSON.stringify(newEfforts) !== JSON.stringify(oldEfforts)) {
|
|
if (!isAllowlisted(provider, model, "supported_efforts")) {
|
|
divergences.push({
|
|
axis: "supported_efforts",
|
|
old: oldEfforts,
|
|
new: newEfforts,
|
|
});
|
|
}
|
|
}
|
|
|
|
// default_effort
|
|
const newDefault = newResult.defaultEffort ?? null;
|
|
const oldDefault = oldResult?.defaultValue ?? null;
|
|
if (newDefault !== oldDefault) {
|
|
if (!isAllowlisted(provider, model, "default_effort")) {
|
|
divergences.push({
|
|
axis: "default_effort",
|
|
old: oldDefault,
|
|
new: newDefault,
|
|
});
|
|
}
|
|
}
|
|
|
|
return divergences;
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Test suites
|
|
// ---------------------------------------------------------------------------
|
|
|
|
let totalChecks = 0;
|
|
let totalDivergences = 0;
|
|
|
|
function runCheck(label, provider, model) {
|
|
totalChecks++;
|
|
const divs = compareEffortAxes(provider, model);
|
|
if (divs.length > 0) {
|
|
totalDivergences += divs.length;
|
|
for (const d of divs) {
|
|
console.error(
|
|
`DIVERGE [${label}] provider=${provider} model=${model} axis=${d.axis}\n` +
|
|
` old: ${JSON.stringify(d.old)}\n` +
|
|
` new: ${JSON.stringify(d.new)}`,
|
|
);
|
|
}
|
|
} else if (VERBOSE) {
|
|
console.log(`OK [${label}] provider=${provider} model=${model}`);
|
|
}
|
|
}
|
|
|
|
// 1. effortTable.fixture.json
|
|
console.log("--- effortTable.fixture.json ---");
|
|
const fixture = JSON.parse(
|
|
readFileSync(
|
|
join(repoRoot, "desktop", "src", "features", "agents", "ui", "effortTable.fixture.json"),
|
|
"utf8",
|
|
),
|
|
);
|
|
for (const entry of fixture) {
|
|
if (!entry.provider) continue;
|
|
runCheck("fixture", entry.provider, entry.model ?? "");
|
|
}
|
|
|
|
// 2. normative-corpus.json (effort axes only)
|
|
console.log("--- normative-corpus.json ---");
|
|
const corpus = JSON.parse(
|
|
readFileSync(join(repoRoot, "scripts", "normative-corpus.json"), "utf8"),
|
|
);
|
|
for (const entry of corpus) {
|
|
if (entry._group) continue;
|
|
if (!entry.provider || !entry.expect) continue;
|
|
if (!entry.expect.supported_efforts && !entry.expect.default_effort) continue;
|
|
runCheck("corpus", entry.provider, entry.raw_model_id ?? "");
|
|
}
|
|
|
|
// 3. catalog-sample-fixture.json (exact records from pinned models.dev payload)
|
|
console.log("--- catalog-sample-fixture.json ---");
|
|
const catalogFixture = JSON.parse(
|
|
readFileSync(join(repoRoot, "scripts", "catalog-sample-fixture.json"), "utf8"),
|
|
);
|
|
for (const ep of catalogFixture.endpoints ?? []) {
|
|
if (!ep.name) continue;
|
|
// All catalog endpoints are databricks_v2 provider
|
|
runCheck("catalog-sample", "databricks_v2", ep.name);
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Summary
|
|
// ---------------------------------------------------------------------------
|
|
|
|
// Count total allowlist axis slots expected to be hit
|
|
const totalAllowlistSlots = ALLOWLIST.reduce((n, e) => n + e.axes.length, 0);
|
|
const allowlistHitCount = allowlistHits.size;
|
|
|
|
// Detect stale allowlist entries (declared but never actually suppressed a divergence)
|
|
const staleEntries = [];
|
|
for (const entry of ALLOWLIST) {
|
|
for (const axis of entry.axes) {
|
|
const key = `${entry.provider}:${entry.raw_model_id}:${axis}`;
|
|
if (!allowlistHits.has(key)) {
|
|
staleEntries.push({ ...entry, axis });
|
|
}
|
|
}
|
|
}
|
|
|
|
console.log(
|
|
`\nDifferential: ${totalChecks} checks, ${totalDivergences} unexpected divergences, ${allowlistHitCount}/${totalAllowlistSlots} allowlist slots exercised`,
|
|
);
|
|
|
|
if (staleEntries.length > 0) {
|
|
for (const e of staleEntries) {
|
|
console.error(
|
|
`STALE_ALLOWLIST provider=${e.provider} model=${e.raw_model_id} axis=${e.axis} — entry never fired; remove or update it`,
|
|
);
|
|
}
|
|
}
|
|
|
|
if (totalDivergences > 0) {
|
|
console.error(
|
|
`FAIL: ${totalDivergences} unexpected divergence(s) — see output above`,
|
|
);
|
|
process.exit(1);
|
|
} else if (staleEntries.length > 0) {
|
|
console.error(
|
|
`FAIL: ${staleEntries.length} stale allowlist entry(ies) — entries that never suppress a divergence mask future regressions`,
|
|
);
|
|
process.exit(1);
|
|
} else {
|
|
console.log("PASS: old and new effort logic agree on all non-allowlisted entries");
|
|
}
|