mirror of
https://github.com/block/buzz.git
synced 2026-08-18 06:50:31 +02:00
## Problem Observer telemetry is the noisiest client of the relay: the old pacer (167ms spacing + 90/min rolling cap) let a busy session bill up to 6 events/second against the owner's message quota, and the rolling cap silently *dropped* frames once exceeded. Ruling from the rate-limiting investigation thread (channel `826fc99b-1472-40e7-a529-6b9db8943b8c`): pace at 1/s, always emit, minimal PR. **Review round 1 (Max, Sami)** found the first cut wrong in three ways — tick burst (all pending frames per tick), startup burst (`interval` fires at t=0), and per-channel quota arithmetic. All fixed and mutation-verified in round 1. **Review round 2 (Sami, Max)** found two more against the round-1 head: 1. **Drain-rate collapse (Sami, blocker):** front-run-only packing meant a frame held ONE event whenever channels interleaved — measured 275 B/s vs 63.5 KB/s, so an ordinary 2-channel session fell minutes behind with zero drops and no warning. Silent unbounded latency. 2. **Coalescer byte-cap bypass (Max):** chunks pending in `ObserverChunkCoalescer` were unbounded and outside the 4 MiB cap — 500 distinct-`messageId` 50KB chunks retained ~48MB with `pending_bytes == 0` and zero drops. **Review round 3 (Max)** found the drop accounting undercounted merged chunks: a coalescer entry that merged N same-`messageId` chunks counted as **1** in `dropped_events` when evicted (50 merged 1KB chunks evicted → counter read 1, 49 generated events unaccounted). Fixed: accounting is now denominated in **source (generated) observer events** end to end — each pending entry tracks how many chunks it absorbed, eviction charges that count, and the count survives flush into the publish FIFO. **Review round 4 (Sami, Max)** found three more against the round-3 head: 1. **Coalescer byte undercount (both, independently):** a pending merged entry retains its first chunk's text **twice** until flush — once inside the serialized event skeleton and once in the extracted text accumulator — but was charged only `serialized_len`, so true retention overshot the 4 MiB cap ~2× (measured 8.3 MB). Fixed: `push_pending` charges `serialized_len(&event) + text.len()`. 2. **Cap regressions asserted the accumulator against itself,** which is how the undercount hid. All three cap tests now assert on independently **walked** retained bytes (`serialized_len` per FIFO entry + `serialized_len + text.len()` per coalescer entry), with a secondary `accumulator >= walked` sanity check. Reverting the fix makes them fail at exactly 8,328,386 / 8,328,272 bytes. 3. **FIFO-arm source accounting was implemented but untested (Sami M13; Max reproduced at `cc9333b7c` with 102/151):** the round-3 regression only evicted a merged entry while still in the coalescer. New test forces a merged entry (50×1KB, `source_events=50`) through flush into the publish FIFO, then evicts it from there — mutating the FIFO eviction to `dropped += 1` fails with the reviewers' exact numbers (102 vs 151). **Review round 5 (Sami 9/9/9, Max 9/9/9)** — production judged merge-safe by both; remaining items are tests only, all landed at `63d821620`: 1. **The walker instrument was itself unverified (Sami M17–M20; Max independently confirmed the `return 0` mutant survives):** every cap test asks `walked_retained_bytes()` only for `<= CAP`, so a blinded walker passes everything — and paired with a reverted `push_pending` fix the two mutations cancel, hiding exactly the 8.3 MB overshoot it exists to detect. New two-sided pin: the walker must SEE the first chunk's text twice, and must agree with the accumulator EXACTLY while both stores are non-empty. Kills M17, M18, M19, M20. 2. **Two pre-existing snapshot-clone siblings (Sami D5b/D5c; byte-identical at merge-base `7334ad1e1` — not this PR's regression, but the PR made the class visible):** aliasing the inner turns map leaks a post-save turn into the snapshot; aliasing the inner tombstones map leaks a post-save terminal that blocks a legitimate post-restore resurrection. Two isolation tests with in-test controls — all three inner-map clones in `saveActiveAgentTurnsForCommunity` are now pinned. ## Change **Harness (`crates/buzz-acp`)** - **Global pacer: AT MOST ONE relay frame per second**, regardless of channel count or backlog size. `interval_at(now + 1s)` restores the no-startup-burst property; `MissedTickBehavior::Skip` is now pinned by a paused-time test (a stalled tick arm fires one catch-up frame, not one per missed deadline). At 1 frame/s telemetry spends ≤60/min of the shared 120/min quota; `OBSERVER_PUBLISH_TICK` documents the tradeoff as the knob. - **`ObserverPublishQueue` with gather-packing:** events wait as byte-accounted events (FIFO). `next_frame()` packs the front event's channel **gathered queue-wide in FIFO order** — frames never mix channels, and each channel's events keep their FIFO order, but cross-channel frame order MAY differ from arrival order. That is what keeps the drain rate in **bytes per slot** (one ~64KB frame/s) instead of front-run-length events per slot. **Null-channel events (`agent_panic`-class) are packing barriers** nothing gathers across, so causally-global events keep exact order against every channel. - **One byte cap over BOTH stores:** the event FIFO and the coalescer's pending chunk buffer count against the 4 MiB budget together; eviction is oldest-first across both (queue front, then coalescer front — structural age order) with accounting (warn + counter). A high-cardinality chunk flood is bounded exactly like a plain event flood. Coalescer entries are charged their **true** retention (`serialized_len + text.len()` — the first chunk's text lives in both the serialized skeleton and the extracted accumulator until flush). - **Shutdown is not a burst bypass:** paced one-frame-per-tick drain until empty. **Desktop** - `unwrapObserverBatch` expands envelopes on the live relay path and archive-ingest seam (round 1, unchanged). - **`activeAgentTurnsStore` watermark re-keyed per (agent, channel)** with a dedicated null-channel bucket: the per-agent `(timestamp, seq)` gate would silently skip a delayed channel's frames as stale under gather-packing's intentional cross-channel reorder. Safe because every turn-mutating path is channel-scoped by the event's own `channelId` (endTurn's null-turnId fallback matches `turn.channelId`; resurrectTurn keys on `event.channelId`), so per-channel serialization preserves each guard the per-agent gate provided. The tombstone-cap justification is rewritten for the new keying (worst case for an evicted tombstone is a ghost badge the prune reaps — bounded cosmetic staleness, not corruption). Community-switch save/restore deep-clones the nested map. Other per-agent maps stay agent-keyed: the clock offset is a running minimum (order-insensitive); turns/tombstones mutate only through channel-scoped paths. ## Version skew — old desktop + new harness Gather-packing *intentionally* emits cross-channel-reordered frames. An **old desktop** (per-agent watermark) against a **new harness** will silently skip a delayed channel's turn-state events as stale — working badges on that channel can go stale/missing until its next fresh event. Transcript and archive are unaffected (the transcript store sorts + rebuilds on out-of-order arrival; the archive is per-channel by construction). Ship desktop and harness together; skew degrades badges only, not data at rest. ## Throughput ceiling — "lossless" is qualified Sustained lossless rate is what fits in one ~64KB frame per second, now genuinely in bytes under interleaving: | event payload | events per frame | sustained ceiling | |---|---|---| | 100 B | 250 | 250 ev/s | | 500 B | 99 | 99 ev/s | | 2 KB | 30 | 30 ev/s | | 10 KB | 6 | 6 ev/s | With C channels producing concurrently, publish slots round-robin between them: per-channel drain is ~64KB/C per second and the 4 MiB burst budget (~64s single-channel) shortens accordingly. Beyond budget, oldest-first drops **with accounting** — visible, designed loss. **Accounting semantics:** `dropped_events` counts SOURCE (generated) observer events, not retained entries — evicting a coalesced entry that merged N chunks charges N. On the published side, a merged entry ships all N sources' text in ONE event, so the reconciliation invariant is `ingested == dropped_events + Σ source_events over published events` (for unmerged events, source_events = 1). ## Verification At `63d821620d3513505e8766ac691a8002f9d4a96f` (this head; `git rev-parse HEAD` matched in the same shell as every run), rustc 1.95.0: - `cargo test -p buzz-acp`: **689 lib + 9 integration, 0 failed** — regressions: interleaved 2-channel drain packs into ≤4 frames not 200 slots; null-channel barrier; queue-wide gather with within-channel FIFO; distinct-key 50KB chunk flood bounded by the cap with event-level accounting (published + dropped == ingested, survivors newest); paused-time `MissedTickBehavior::Skip` pin (verified to fail under `Burst`: 3 frames vs 1); merged-key eviction accounts every absorbed source chunk in BOTH arms — coalescer-side (Max's round-3 probe) and post-flush FIFO-side (Sami M13 / Max's round-4 probe: fails 102 vs 151 under `+= 1`). All three cap tests assert on independently walked retained bytes, not the accumulator (verified to fail without the `+text.len()` fix: 8,328,386 / 8,328,272 vs 4 MiB); the walker itself is pinned two-sided against the accumulator (all four blinding mutants M17–M20 verified to fail it, including the walker+fix cancellation pair). - `cargo clippy -p buzz-acp --all-targets -- -D warnings` clean, `cargo fmt --check` clean - Desktop: `tsc --noEmit` clean; node tests **4366 passed, 0 failed** — snapshot-clone family fully pinned: watermark aliasing (round 4), turns aliasing and tombstone aliasing (round 5, pre-existing gaps; each mutant verified to fail exactly its target test with an in-test control). Prior rounds: cross-channel reorder processed, cross-channel-delayed null-turnId `turn_error` evicts only its own channel's turn, null-bucket replay idempotency, same-channel stale/duplicate still skipped, watermark survives community-switch save/restore - All pre-push hooks green at the pushed commit (branch-skew, desktop-check, desktop-test, rust-tests, desktop-tauri-checks) Part of the rate-limiting fix stack; independent of `eva/rate-limit-fixes` by design (separate minimal PR per Tyler's ruling). --------- Signed-off-by: Eva <011987e296fd5006292d2f930b574be47c7801048d1983c46c425d3c95f0cffd@buzz.block.builderlab.xyz> Signed-off-by: Sami <f4a42a97e594b77bdbd8ee35191c8b28a94a4cb871d96f32921558275421fb68@buzz.block.builderlab.xyz> Co-authored-by: Eva <011987e296fd5006292d2f930b574be47c7801048d1983c46c425d3c95f0cffd@buzz.block.builderlab.xyz> Co-authored-by: Sami <f4a42a97e594b77bdbd8ee35191c8b28a94a4cb871d96f32921558275421fb68@buzz.block.builderlab.xyz>