mirror of
https://github.com/block/buzz.git
synced 2026-08-18 06:50:31 +02:00
The relay awaited reconcile_nip43_membership_snapshots for EVERY
provisioned community before binding the listener. The sweep is
sequential, two queries per community plus a publication per repair
(~9.5ms/community measured locally). Cost is linear in community count,
so on a deployment with many communities the pre-bind sweep takes
minutes — far past a typical startup probe window. The probe SIGKILLs
the pod mid-sweep, the restart begins again at community number one,
and the pod crashloops forever: low RSS, port never bound, zero log
lines at RUST_LOG=error.
Fix, minimal by design:
- Pre-bind: reconcile ONLY the deployment community (one snapshot check,
bounded), with explicit phase logs carrying elapsed_ms.
- The fleet-wide sweep moves entirely into the periodic post-bind task,
now jittered (random fraction of the interval, same rationale as the
usage-metrics poller) and leader-gated via a session advisory lock so
N replicas do not run N concurrent fleet-wide sweeps. Lock is held
only for the duration of a sweep; a replica dying mid-sweep releases
it with its session.
- Sweep telemetry: start/progress(1000)/complete/fail logs with
elapsed_ms, plus a buzz_nip43_membership_sweep_seconds histogram.
Pagination/resume and per-query deadlines are a deliberate follow-up PR.
Verified: cargo test -p buzz-relay — 837 passed, 1 pre-existing failure
(api::mesh_demo::tests::demo_join_forwarded_arm_round_trips_echo) which
fails identically at origin/main 5e0efb0bb with this diff stashed.
Co-authored-by: npub1qyvc0c5kl4gqv2fd97fsk46tu378sqgy35vc83rvgfwne90sel7s0ed67d <011987e296fd5006292d2f930b574be47c7801048d1983c46c425d3c95f0cffd@buzz.block.builderlab.xyz>
Signed-off-by: npub1qyvc0c5kl4gqv2fd97fsk46tu378sqgy35vc83rvgfwne90sel7s0ed67d <011987e296fd5006292d2f930b574be47c7801048d1983c46c425d3c95f0cffd@buzz.block.builderlab.xyz>