mirror of
https://github.com/block/buzz.git
synced 2026-08-18 06:50:31 +02:00
Uses MeshLLM built-in `mesh` collective intelligence / Mixture of Agents when Buzz Auto sees two or more distinct physical models. With zero or one model, Auto remains ordinary `auto`. This is an alternative way to improve tool responses, accuracy, and resistance to hallucination when a high-latency distributed mesh contains diverse models. Models and members may come and go: collective routing enables only after stable capacity, drops on confirmed contraction, and can recover later. Mesh-specific failures retry once through ordinary Auto. This update also pins MeshLLM to a v0.73.1-compatible backport of [MeshLLM #1074](https://github.com/Mesh-LLM/mesh-llm/pull/1074), so client-only Buzz nodes cannot enter model election or download a remote provider model. Buzz preserves the selected local sharing model and switches an existing client to sharing across a controlled app restart, retaining one runtime and one `:9337` / `:3131` pair per machine. Validation: - Full local `just ci` passes on the cleaned branch. - MeshLLM host-runtime suite: 1,568 passed, 0 failed; strict Clippy passes. - Buzz desktop Tauri suite with `mesh-llm`: 1,721 passed, 0 failed; strict feature Clippy passes. - Playwright covers client-to-share using the saved local model and no destructive stop. - Packaged two-machine testing proved single-model routing, dual-model collective routing, tool-markup fallback, and runtime reuse. - Packaged client-only recheck routed a real Mini Buzz turn through M5 while Mini stayed `is_client=true`, `is_host=false`, hosted no models, and created no Gemma cache. Builds on the recovery work merged in #2823; this PR does not duplicate it. --------- Signed-off-by: Michael Neale <michael.neale@gmail.com> Signed-off-by: Tyler Longwell <tlongwell@block.xyz> Co-authored-by: Michael Neale <michael.neale@gmail.com> Co-authored-by: npub1qyvc0c5kl4gqv2fd97fsk46tu378sqgy35vc83rvgfwne90sel7s0ed67d <011987e296fd5006292d2f930b574be47c7801048d1983c46c425d3c95f0cffd@buzz.block.builderlab.xyz> Co-authored-by: Tyler Longwell <tlongwell@block.xyz>
69 lines
2.9 KiB
Rust
69 lines
2.9 KiB
Rust
use serde::{Deserialize, Serialize};
|
|
|
|
/// Host-side "who is using the compute I'm sharing" snapshot.
|
|
///
|
|
/// Read-only projection of the serving node's own runtime metrics (the same
|
|
/// `routing_metrics` / `inflight_requests` the SDK already exposes on the local
|
|
/// console). No new trust surface: it reads the node's own status payload.
|
|
///
|
|
/// The local/remote/endpoint attempt split is what distinguishes *my own*
|
|
/// agent (local) from *another member consuming my compute* (remote/endpoint).
|
|
#[derive(Debug, Clone, Serialize, Deserialize, PartialEq, Default)]
|
|
#[serde(rename_all = "camelCase")]
|
|
pub struct MeshServingUsage {
|
|
/// Requests being served right now.
|
|
pub inflight: u64,
|
|
/// Highest concurrent in-flight seen this session.
|
|
pub peak_inflight: u64,
|
|
/// Total requests routed through this node.
|
|
pub requests_served: u64,
|
|
/// Completion tokens produced.
|
|
pub tokens_served: u64,
|
|
/// Recent decode throughput.
|
|
pub tokens_per_second: f64,
|
|
/// Requests served for this machine's own agents.
|
|
pub local_attempts: u64,
|
|
/// Requests served for a remote peer (someone else consuming my compute).
|
|
pub remote_attempts: u64,
|
|
/// Requests served via an advertised endpoint (also a remote consumer).
|
|
pub endpoint_attempts: u64,
|
|
/// Other nodes currently visible as peers.
|
|
pub peers: u64,
|
|
}
|
|
|
|
/// Pure extractor: project a raw SDK status payload into [`MeshServingUsage`].
|
|
///
|
|
/// Every field is read defensively (missing → 0) so an SDK shape change
|
|
/// degrades to "no usage shown" rather than an error. Kept pure so it can be
|
|
/// unit-tested against a captured payload without a live runtime.
|
|
pub fn serving_usage_from_payload(payload: &serde_json::Value) -> MeshServingUsage {
|
|
let u64_at = |v: &serde_json::Value| v.as_u64().unwrap_or(0);
|
|
let rm = payload.get("routing_metrics");
|
|
let local = rm.and_then(|m| m.get("local_node"));
|
|
let get_u64 = |obj: Option<&serde_json::Value>, key: &str| {
|
|
obj.and_then(|o| o.get(key)).map(u64_at).unwrap_or(0)
|
|
};
|
|
MeshServingUsage {
|
|
inflight: local
|
|
.and_then(|l| l.get("current_inflight_requests"))
|
|
.map(u64_at)
|
|
.or_else(|| payload.get("inflight_requests").map(u64_at))
|
|
.unwrap_or(0),
|
|
peak_inflight: get_u64(local, "peak_inflight_requests"),
|
|
requests_served: get_u64(rm, "request_count"),
|
|
tokens_served: get_u64(rm, "completion_tokens_observed"),
|
|
tokens_per_second: rm
|
|
.and_then(|m| m.get("avg_tokens_per_second"))
|
|
.and_then(serde_json::Value::as_f64)
|
|
.unwrap_or(0.0),
|
|
local_attempts: get_u64(local, "local_attempt_count"),
|
|
remote_attempts: get_u64(local, "remote_attempt_count"),
|
|
endpoint_attempts: get_u64(local, "endpoint_attempt_count"),
|
|
peers: payload
|
|
.get("peers")
|
|
.and_then(serde_json::Value::as_array)
|
|
.map(|a| a.len() as u64)
|
|
.unwrap_or(0),
|
|
}
|
|
}
|