Files
tutabridge/crates
Anthony MandGitHub 1863144627 Server robustness: SMTP size limits, resilient accept loop, backup offload (#9)
* smtp: enforce message size and line length limits

The server advertised SIZE 26214400 in EHLO but never enforced it, and
the DATA loop appended every line into an in-memory buffer with no cap,
so a single local client could grow the process memory without bound (a
line with no terminator was read unboundedly too).

Enforce both: reject a MAIL FROM that declares an over-limit SIZE, stop
buffering and reply 552 once a message exceeds the cap, and bound each
protocol line. handle_connection is now generic over the stream so the
whole conversation can be exercised over an in-memory pipe; tests cover
the size param, the DATA cap, the per-line cap, and a normal send.

* backup: run mail decryption and writes off the async runtime

export_eml decrypted each cached .eml.enc and wrote the output file
inline on the async task. A GUI backup reuses the running bridge's
runtime, so over a large already-synced mailbox that tight, non-yielding
loop pinned a worker and froze the live IMAP/SMTP servers for the whole
export (the same failure class as the cached-folder load).

Wrap the per-mail decrypt and file write in block_in_place so the worker
hands its other tasks off and the servers stay responsive. The backup
integration tests run on a multi-thread runtime now (block_in_place
requires it) and still assert the same cache/server/resume behaviour.

* net: tolerant accept loop with a connection cap and handshake timeout

Both servers ran `loop { listener.accept().await? }`. A single transient
accept error (EMFILE, ECONNABORTED, ...) propagated out and stopped the
server for good, nothing bounded concurrent connections, and a stalled
TLS handshake was never timed out (a client that connects but never
negotiates parked a task and a file descriptor forever).

Extract a shared net::accept_loop that logs and retries a failed accept,
caps concurrency with a semaphore (64 connections), and wrap each
handshake in a 15s timeout. The loop is transport agnostic so it is unit
tested without TLS: one test proves it keeps accepting across
connections, another that it bounds concurrency at the cap.

* event-bus: recover poisoned last_batch_ids lock instead of panicking

last_batch_ids is a std Mutex shared between the bridge, the event
handler, and the SDK's reconnect path. Every accessor used
.lock().unwrap(), so one panic while holding it would poison the mutex
and make every later lock (the SDK reconnect included) panic, killing
realtime sync for the rest of the process's life.

Add util::lock_recover (locks, recovering the guard from poisoning) and
use it at the bridge-side accessors. Tested against a poisoned mutex.
2026-06-14 21:12:56 +02:00
..