mirror of
https://github.com/block/buzz.git
synced 2026-08-18 06:50:31 +02:00
## Problem On SIGTERM the relay sends every live WebSocket a **1012 Service Restart** close frame via `ConnectionManager::drain_all()` — all in the same instant (`main.rs` shutdown task → `state.rs::drain_all`). On a pod holding thousands of sessions, that makes every client reconnect simultaneously: the thundering-herd reconnect behind the DB pool-timeout bursts observed on each rolling deploy. Client-side jitter can't fix this — the desktop client *resets* its backoff to base on a 1012 and reconnects with only ±25% jitter (`relayClientSession.ts`), so the spread has to come from the server. ## Change Add `BUZZ_DRAIN_JITTER_MS` (default `0` = unchanged behavior). The two paths are kept **deliberately separate** so the default is byte-for-byte the previously shipped shutdown: - **Jitter off (`0`/unset, the default):** the original synchronous, all-at-once `drain_all()` runs unchanged — queue the 1012 on each connection's control channel, cancel, return. No new machinery on the default path. - **Jitter on (`> 0`):** a separate async `drain_all_jittered(jitter_ms)` spreads each connection's restart close over an independent uniform delay in **`[1, jitter_ms]`**. Each delayed close travels a dedicated `RestartClose` channel; the writer flushes the 1012 frame and **acknowledges the flush over a oneshot**, so drain waits for confirmed delivery (up to `RESTART_CLOSE_ACK_TIMEOUT` = 5s) rather than assuming it, falling back to cancellation if the channel is full/closed or the ack times out. The drain future is **owned and awaited** by the shutdown task, and the 30s hard-drain backstop is aborted only after a clean drain — so a clean roll exits `0`. The two methods can be unified and the old one dropped later once the jittered path is proven for all cases. - **`config.rs`** — `drain_jitter_ms`: non-negative parse, clamped to `MAX_DRAIN_JITTER_MS` = **20s** (leaving 10s of the 30s budget for flush). Junk fails loudly at startup; **empty/whitespace-only is treated as unset (jitter off)** so a `BUZZ_DRAIN_JITTER_MS=""` kill switch does not crashloop the relay (matches the sibling env vars in this file). - **`state.rs`** — `drain_all()` (unchanged synchronous default) + `drain_all_jittered()` (jittered + flush-ack). Both set the sticky `draining` flag before the first await. A registration that lands mid-shutdown always self-signals via the **immediate** control-frame + cancel path — jitter smears already-established sockets, not late arrivals. - **`main.rs`** — shutdown task dispatches: `drain_jitter_ms == 0` → `drain_all()`, else `drain_all_jittered(...).await`. ## Safety - **Default off is the currently-committed path.** With jitter unset/0 the shutdown runs the original synchronous `drain_all()` — no restart channel, no ack wait. Safe to deploy dark and dial up. - **Shutdown-boundary race preserved.** Sticky flag set before any await; a late registration self-signals its close with no jitter. - **Owned + backstopped.** The jittered drain future is awaited; the 30s hard-drain `process::exit(1)` remains the ceiling. `MAX_DRAIN_JITTER_MS` (20s) + `RESTART_CLOSE_ACK_TIMEOUT` (5s) = 25s, inside the 30s budget; 5s pre-sleep + 25s = 30s against `terminationGracePeriodSeconds: 60`. ## Known behavior to note (not a blocker, flagged from review) On a **successful** flush the jittered path deliberately does not cancel the connection token — teardown then depends on the client echoing our Close, or on process exit. Compliant clients echo; a silent client rides to the 30s hard exit. The default (jitter-off) path cancels deterministically as before. ## Tests - `config::tests::drain_jitter_defaults_off_and_rejects_junk` — default off, `20000`, clamp `60000`→`20000`, explicit `0`, junk `"soon"` fails, **empty `""` and whitespace-only treated as off**. - `state::tests::drain_all_is_immediate` — default path queues frame + cancels synchronously. - `state::tests::drain_all_sends_restart_close_and_cancels_every_conn`, `drain_all_full_control_buffer_still_cancels`, `register_after_drain_self_signals_restart_close_and_cancel`. - `state::tests::drain_all_jittered_defers_close_until_within_jitter_window` (paused time). - `state::tests::drain_all_jittered_waits_for_writer_acknowledgement_without_cancelling`. - `state::tests::drain_all_jittered_cancels_when_restart_channel_is_full_or_closed`. - `state::tests::drain_all_jittered_cancels_when_flush_ack_times_out` (paused time — the 5s ack-timeout fallback). Validation at `46c690940`: `cargo fmt -p buzz-relay --check`, `cargo clippy -p buzz-relay --all-targets -- -D warnings`, and the drain/config unit suite all clean. Local live SIGTERM test with a real relay process + 200 NIP-42-authenticated sockets — see the PR comment for the before/after distribution and exit codes. ## Rollout Ship with default `0`, then set `BUZZ_DRAIN_JITTER_MS` (e.g. 10000–20000) on bb-block first, watch the roll-window pool-timeout metric, then bb-public. `""` is a safe kill switch. Complements the preStop `sleep` (stops routing before close). --------- Signed-off-by: npub1srl70fhzyu3fsnahl06vw2czvqc2w3ds37hyzvjnk8ve8f03ngcqg9le2w <80ffe7a6e22722984fb7fbf4c72b026030a745b08fae413253b1d993a5f19a30@buzz.block.builderlab.xyz> Signed-off-by: npub128x7j3pwgm4vs8yra3c42fcgcwcvh94g3luwzkqa376du2q6l0esqcrwch <51cde9442e46eac81c83ec71552708c3b0cb96a88ff8e1581d8fb4de281afbf3@buzz.block.builderlab.xyz> Signed-off-by: Brad Seiler <seiler@squareup.com> Co-authored-by: npub1srl70fhzyu3fsnahl06vw2czvqc2w3ds37hyzvjnk8ve8f03ngcqg9le2w <80ffe7a6e22722984fb7fbf4c72b026030a745b08fae413253b1d993a5f19a30@buzz.block.builderlab.xyz> Co-authored-by: npub128x7j3pwgm4vs8yra3c42fcgcwcvh94g3luwzkqa376du2q6l0esqcrwch <51cde9442e46eac81c83ec71552708c3b0cb96a88ff8e1581d8fb4de281afbf3@buzz.block.builderlab.xyz>
400 lines
18 KiB
YAML
400 lines
18 KiB
YAML
# Default values for buzz.
|
|
#
|
|
# Two supported tiers:
|
|
#
|
|
# PRODUCTION (default) — external Postgres/Redis/S3, existingSecret
|
|
# refs everywhere, no chart-side autogeneration, GitOps-safe (ArgoCD/Flux).
|
|
# HA-ready: replicaCount >= 2 (requires Redis; git state is object-store-
|
|
# backed, so no ReadWriteMany volume is needed — RWO per replica is fine).
|
|
#
|
|
# QUICKSTART — bundles in-cluster Postgres + Redis + MinIO and
|
|
# auto-generates relay secrets via the `lookup` pattern (NOT GitOps-safe —
|
|
# see README), single replica, evaluation only. Opt in by enabling each
|
|
# bundled service: postgresql.enabled, redis.enabled, minio.enabled.
|
|
# See ci/quickstart-values.yaml and the README.
|
|
#
|
|
# See examples/argocd-app.yaml and examples/flux-helmrelease.yaml for the
|
|
# canonical GitOps configurations.
|
|
|
|
# Intent marker for the evaluation profile, surfaced in NOTES.txt. It does NOT
|
|
# by itself enable any bundled service — set the per-service .enabled flags
|
|
# (postgresql / redis / minio) to bring them up in-cluster.
|
|
quickstart: false
|
|
|
|
# ── Image ────────────────────────────────────────────────────────────────────
|
|
image:
|
|
repository: ghcr.io/block/buzz
|
|
tag: "" # empty → .Chart.AppVersion
|
|
pullPolicy: IfNotPresent
|
|
pullSecrets: []
|
|
|
|
# ── Topology ────────────────────────────────────────────────────────────────
|
|
# replicaCount > 1 hard-requires Redis for buzz-pubsub (in-cluster or external).
|
|
# It does NOT require ReadWriteMany git storage: git ref/object state is
|
|
# object-store-backed (each request hydrates an ephemeral repo from S3; writer
|
|
# serialization is the object-store pointer CAS), and repo-name uniqueness lives
|
|
# in Postgres. Each replica can use its own ReadWriteOnce volume (or none).
|
|
replicaCount: 1
|
|
|
|
# ── Autoscaling ──────────────────────────────────────────────────────────────
|
|
# Requires Metrics Server for CPU. Optional WebSocket scaling additionally
|
|
# requires a custom-metrics adapter exposing its pod-level Prometheus gauge.
|
|
# Kubernetes HPA uses the larger replica recommendation from enabled metrics.
|
|
autoscaling:
|
|
enabled: false
|
|
minReplicas: 5
|
|
maxReplicas: 15
|
|
targetCPUUtilizationPercentage: 65
|
|
websocketMetricEnabled: true
|
|
websocketMetricName: buzz_ws_connections_active
|
|
targetWebsocketConnections: 5000
|
|
behavior:
|
|
scaleUp:
|
|
stabilizationWindowSeconds: 0
|
|
policies:
|
|
- type: Percent
|
|
value: 100
|
|
periodSeconds: 60
|
|
- type: Pods
|
|
value: 4
|
|
periodSeconds: 60
|
|
selectPolicy: Max
|
|
scaleDown:
|
|
stabilizationWindowSeconds: 600
|
|
policies:
|
|
- type: Pods
|
|
value: 1
|
|
periodSeconds: 120
|
|
selectPolicy: Min
|
|
|
|
# ── Public URL ───────────────────────────────────────────────────────────────
|
|
# Required. The wss:// URL clients use to connect. Drives:
|
|
# - RELAY_URL env (relay-side)
|
|
# - Default mediaBaseUrl (https://<host>/media)
|
|
# - Default ingress host
|
|
relayUrl: ""
|
|
mediaBaseUrl: ""
|
|
|
|
# ── Owner ────────────────────────────────────────────────────────────────────
|
|
# 64-char lowercase hex Nostr pubkey of the relay operator. Required when
|
|
# relay.requireRelayMembership=true (the production default).
|
|
ownerPubkey: ""
|
|
|
|
# ── Chart-managed secrets ────────────────────────────────────────────────────
|
|
# Production / GitOps path: create a Secret out-of-band with these keys and
|
|
# point `secrets.existingSecret` at it. Any key omitted from the existing
|
|
# Secret falls back to chart-side autogen (only effective at first install).
|
|
#
|
|
# Expected keys (all optional unless required by relay config):
|
|
# BUZZ_RELAY_PRIVATE_KEY — 64-char hex; relay identity (rotation = identity change)
|
|
# BUZZ_GIT_HOOK_HMAC_SECRET — 32+ chars; required when replicaCount > 1
|
|
# DATABASE_URL — full Postgres URL (preferred over externalPostgresql.url)
|
|
# READ_DATABASE_URL — optional Postgres read-replica URL; omit to keep all reads on the writer
|
|
# REDIS_URL — full Redis URL with auth
|
|
# BUZZ_S3_ACCESS_KEY — S3 access key
|
|
# BUZZ_S3_SECRET_KEY — S3 secret key
|
|
secrets:
|
|
existingSecret: ""
|
|
# Inline overrides (NOT recommended for production; they land in values).
|
|
relayPrivateKey: ""
|
|
gitHookHmacSecret: ""
|
|
|
|
# ── Relay behavior ───────────────────────────────────────────────────────────
|
|
relay:
|
|
bindAddr: "0.0.0.0:3000"
|
|
maxConnections: 10000
|
|
maxConcurrentHandlers: 1024
|
|
sendBuffer: 1000
|
|
# Graceful-shutdown reconnect jitter. On SIGTERM the relay closes every live
|
|
# WebSocket with a 1012 Service Restart frame; with a rolling deploy this can
|
|
# release a whole pod's sockets at once and stampede reconnects into the DB
|
|
# pool. A positive value (milliseconds) spreads each close over a per-socket
|
|
# random delay in [1, drainJitterMs], smoothing the reconnect herd. 0 (the
|
|
# default) closes all sockets at once, preserving the previous behavior.
|
|
# Values above 20000 are capped to 20000, leaving close-frame delivery
|
|
# headroom under the relay's 30s hard-drain timeout (itself inside the 60s
|
|
# terminationGracePeriodSeconds below).
|
|
drainJitterMs: 0
|
|
requireAuthToken: true
|
|
requireRelayMembership: true
|
|
# Authenticated media reads: relay GET/HEAD /media/* requires Blossom
|
|
# kind 24242 t=get plus relay membership. Enabled by default so private
|
|
# attachments are never publicly readable by URL/hash. Only set false for
|
|
# local development or fully public communities — desktop, mobile, and CLI
|
|
# clients all attach read auth.
|
|
requireMediaGetAuth: true
|
|
allowNipOaAuth: true
|
|
pubkeyAllowlist: false
|
|
corsOrigins: []
|
|
# Huddle audio is safe only for single-pod relay deployments until an SFU
|
|
# exists. null lets the chart render false automatically when
|
|
# replicaCount > 1. Explicit true with replicaCount > 1 means the operator
|
|
# accepts/owns the external multi-pod audio/SFU behavior.
|
|
huddleAudioAvailable: null
|
|
ephemeralTtlOverride: 0
|
|
# Per-upload-event records (`_uploads/` moderation side channel). Off by
|
|
# default. Operators hosting communities for other people may have legal
|
|
# obligations (e.g. NCMEC reporting for US-serving providers) that require
|
|
# recording the network address of an upload; Buzz never collects IPs unless
|
|
# uploadIpHeader is set. When set (e.g. "cf-connecting-ip"), the connecting
|
|
# address reported by YOUR trusted edge is stored in the per-event record
|
|
# only — never served to clients, never in event data. uploadRecords must be
|
|
# true for uploadIpHeader to be valid (startup-checked).
|
|
uploadRecords: false
|
|
uploadIpHeader: ""
|
|
uploadPortHeader: ""
|
|
|
|
livenessProbe:
|
|
httpGet:
|
|
path: /_liveness
|
|
port: health
|
|
initialDelaySeconds: 5
|
|
periodSeconds: 10
|
|
timeoutSeconds: 3
|
|
failureThreshold: 3
|
|
readinessProbe:
|
|
httpGet:
|
|
path: /_readiness
|
|
port: health
|
|
initialDelaySeconds: 5
|
|
periodSeconds: 5
|
|
timeoutSeconds: 3
|
|
failureThreshold: 3
|
|
startupProbe:
|
|
httpGet:
|
|
path: /_liveness
|
|
port: health
|
|
failureThreshold: 60
|
|
periodSeconds: 2
|
|
|
|
resources:
|
|
requests:
|
|
cpu: "500m"
|
|
memory: "512Mi"
|
|
limits:
|
|
cpu: "2"
|
|
memory: "2Gi"
|
|
|
|
podAnnotations: {}
|
|
podLabels: {}
|
|
nodeSelector: {}
|
|
tolerations: []
|
|
affinity: {}
|
|
topologySpreadConstraints: []
|
|
securityContext:
|
|
runAsNonRoot: true
|
|
runAsUser: 65532
|
|
runAsGroup: 65532
|
|
fsGroup: 65532
|
|
seccompProfile:
|
|
type: RuntimeDefault
|
|
containerSecurityContext:
|
|
allowPrivilegeEscalation: false
|
|
capabilities:
|
|
drop: [ALL]
|
|
readOnlyRootFilesystem: false # git writes need a writable repo path
|
|
terminationGracePeriodSeconds: 60
|
|
|
|
# Optional image entrypoint/arguments overrides. Empty arrays preserve the
|
|
# relay image's defaults. Consumers own compatibility with the selected image.
|
|
command: []
|
|
args: []
|
|
# Appended to the chart-owned relay mounts. Names must match extraVolumes (or
|
|
# another volume supplied by the platform) and must not collide with built-ins.
|
|
extraVolumeMounts: []
|
|
|
|
extraEnv: []
|
|
extraEnvFrom: []
|
|
|
|
# ── Pod extensions ──────────────────────────────────────────────────────────
|
|
# Raw Kubernetes fragments appended to the relay Pod. They are rendered with
|
|
# toYaml, not tpl. Init containers must define their own securityContext and
|
|
# resources; names must not collide with chart-owned containers or volumes.
|
|
extraInitContainers: []
|
|
extraVolumes: []
|
|
|
|
# ── Device pairing relay ─────────────────────────────────────────────────────
|
|
# Optional, stateless NIP-AB relay. When enabled, the main relay advertises
|
|
# pairingRelay.url in NIP-11 and Buzz clients use it instead of the legacy
|
|
# same-host /pair convention.
|
|
pairingRelay:
|
|
enabled: false
|
|
url: ""
|
|
replicaCount: 1
|
|
service:
|
|
type: ClusterIP
|
|
port: 5000
|
|
annotations: {}
|
|
podAnnotations: {}
|
|
podLabels: {}
|
|
resources:
|
|
requests:
|
|
cpu: "50m"
|
|
memory: "32Mi"
|
|
limits:
|
|
cpu: "250m"
|
|
memory: "128Mi"
|
|
|
|
# ── Service ──────────────────────────────────────────────────────────────────
|
|
service:
|
|
type: ClusterIP
|
|
port: 3000
|
|
healthPort: 8080
|
|
metricsPort: 9102
|
|
annotations: {}
|
|
|
|
serviceAccount:
|
|
create: true
|
|
name: ""
|
|
annotations: {}
|
|
|
|
podDisruptionBudget:
|
|
enabled: true
|
|
minAvailable: 1
|
|
maxUnavailable: ""
|
|
|
|
# ── Ingress (classic) ────────────────────────────────────────────────────────
|
|
# Mutually exclusive with httproute.enabled.
|
|
ingress:
|
|
enabled: false
|
|
className: ""
|
|
annotations: {}
|
|
hosts: [] # empty → derived from relayUrl
|
|
tls: [] # [{hosts: [...], secretName: "..."}]
|
|
|
|
# ── Gateway API (HTTPRoute) ──────────────────────────────────────────────────
|
|
httproute:
|
|
enabled: false
|
|
parentRefs: []
|
|
hostnames: []
|
|
rules: [] # empty → default match-all → service
|
|
|
|
# ── Git scratch volume ───────────────────────────────────────────────────────
|
|
# Ephemeral working space only. No persistent git state lives here — reads/writes
|
|
# hydrate ephemeral repos from object storage per request, and repo-name
|
|
# uniqueness lives in Postgres.
|
|
#
|
|
# enabled: true → mount a PVC at mountPath (durable across pod restarts, but a
|
|
# single ReadWriteOnce PVC binds to one node, so it does NOT support multi-pod
|
|
# scheduling across nodes on a Deployment).
|
|
# enabled: false → mount a per-pod emptyDir at mountPath (pure scratch), bounded
|
|
# by size. This is the correct choice for multi-replica HA: each pod gets its
|
|
# own local working space, nothing is shared, and there is no volume to
|
|
# multi-attach. Safe because the object store + Postgres are the sources of
|
|
# truth, not this disk.
|
|
persistence:
|
|
git:
|
|
enabled: true
|
|
mountPath: /var/lib/buzz/git
|
|
storageClass: ""
|
|
accessMode: ReadWriteOnce
|
|
size: 10Gi # PVC capacity or emptyDir sizeLimit
|
|
annotations: {}
|
|
existingClaim: ""
|
|
|
|
# ── Postgres ─────────────────────────────────────────────────────────────────
|
|
# Eval-only CloudPirates subchart. The relay's DATABASE_URL is composed in the
|
|
# chart-managed Secret with a chart-generated password; auth.existingSecret
|
|
# points this subchart at that same Secret/key so server and client agree.
|
|
postgresql:
|
|
enabled: false
|
|
auth:
|
|
database: buzz
|
|
username: buzz
|
|
existingSecret: '{{ if contains "buzz" .Release.Name }}{{ .Release.Name }}-relay{{ else }}{{ .Release.Name }}-buzz-relay{{ end }}'
|
|
secretKeys:
|
|
adminPasswordKey: postgres-password
|
|
persistence:
|
|
enabled: true
|
|
size: 10Gi
|
|
externalPostgresql:
|
|
url: "" # postgres://user:pass@host:5432/db — placeholder example, sadscan:disable np.postgres.1
|
|
|
|
# ── Redis ────────────────────────────────────────────────────────────────────
|
|
# Eval-only CloudPirates subchart (standalone). REDIS_URL is composed in the
|
|
# chart-managed Secret; auth.existingSecret points the subchart at that Secret
|
|
# so the server password matches the URL the relay dials.
|
|
redis:
|
|
enabled: false
|
|
auth:
|
|
existingSecret: '{{ if contains "buzz" .Release.Name }}{{ .Release.Name }}-relay{{ else }}{{ .Release.Name }}-buzz-relay{{ end }}'
|
|
existingSecretPasswordKey: redis-password
|
|
persistence:
|
|
enabled: true
|
|
size: 4Gi
|
|
externalRedis:
|
|
url: "" # redis://:pass@host:6379
|
|
|
|
# ── S3 / object storage (media) ──────────────────────────────────────────────
|
|
# Production: point endpoint/bucket at an external S3-compatible service and
|
|
# supply credentials (inline below or via secrets.existingSecret).
|
|
# Quickstart (`minio.enabled: true`): the chart runs an in-cluster, eval-only
|
|
# MinIO Deployment, creates the bucket via a post-install Job, and composes
|
|
# the endpoint + autogenerated credentials automatically.
|
|
#
|
|
# Storage metrics (hourly bucket sweep, BUZZ_STORAGE_METRICS — see env docs):
|
|
# the credentials above must additionally grant `s3:ListBucket` on the bucket
|
|
# ARN itself (bucket-level; distinct from the object-level GetObject/
|
|
# PutObject/DeleteObject perms already required for media). Without it the
|
|
# first sweep fails AccessDenied and buzz_storage_sweep_ok stays 0 — no other
|
|
# media functionality is affected. Set BUZZ_STORAGE_METRICS=off to disable
|
|
# the sweep entirely on a deployment that can't grant it.
|
|
# Note: buzz_storage_sweep_failures is a process-local gauge — on leader
|
|
# failover it resets to the new leader's local count, not a global total.
|
|
# Note: on a failed sweep attempt, the next retry fires on the next usage tick
|
|
# (default 300 s BUZZ_USAGE_METRICS_INTERVAL_SECS), not at sweep-interval
|
|
# cadence — so a permanently missing s3:ListBucket yields one cheap LIST call
|
|
# per tick until the permission is added.
|
|
s3:
|
|
endpoint: ""
|
|
bucket: "buzz-media"
|
|
# Optional SigV4 signing region. Leave empty to preserve the relay's
|
|
# AWS_REGION fallback; set the provider's credential value when needed.
|
|
region: ""
|
|
# path: https://endpoint/bucket/key (bundled MinIO-compatible default)
|
|
# virtual: https://bucket.endpoint/key (standard S3; required by new Railway buckets)
|
|
addressingStyle: path
|
|
accessKey: ""
|
|
secretKey: ""
|
|
|
|
# In-cluster MinIO for the quickstart profile only. Production deploys leave
|
|
# this disabled and use s3.* (or secrets.existingSecret) against managed S3.
|
|
minio:
|
|
enabled: false # quickstart: set true for bundled in-cluster MinIO
|
|
image: minio/minio:RELEASE.2025-09-07T16-13-09Z
|
|
mcImage: minio/mc:RELEASE.2025-08-13T08-35-41Z
|
|
persistence:
|
|
enabled: true
|
|
size: 10Gi
|
|
|
|
# ── Git server config ────────────────────────────────────────────────────────
|
|
git:
|
|
maxPackBytes: 524288000 # 500 MiB
|
|
packCachePath: /var/cache/buzz/git-packs
|
|
packCacheMaxBytes: 5368709120 # 5 GiB
|
|
packCacheMaxConcurrentPopulations: 2
|
|
packCacheVolumeSize: 7Gi # per-pod emptyDir; includes cold-population staging
|
|
maxReposPerPubkey: 100
|
|
maxConcurrentOps: 20
|
|
|
|
# ── Migrations ───────────────────────────────────────────────────────────────
|
|
# Relay runs sqlx migrations at startup via BUZZ_AUTO_MIGRATE=true.
|
|
migrate:
|
|
autoMigrate: true
|
|
preUpgradeJob:
|
|
enabled: false
|
|
resources: {}
|
|
backoffLimit: 3
|
|
activeDeadlineSeconds: 600
|
|
|
|
# ── Monitoring ───────────────────────────────────────────────────────────────
|
|
serviceMonitor:
|
|
enabled: false
|
|
namespace: ""
|
|
interval: 30s
|
|
scrapeTimeout: 10s
|
|
labels: {}
|
|
|
|
# ── Free-form extra manifests ────────────────────────────────────────────────
|
|
extraManifests: []
|