mirror of
https://github.com/CloakHQ/CloakBrowser.git
synced 2026-06-23 11:41:46 +02:00
* feat(lambda): cold-start hardening + handler-side retry orchestration
Two related improvements based on benchmarking the integration at scale
(3454-site sample, multiple iterations).
Cold-start hardening (lambda-entrypoint.sh + lambda_handler.py):
- Clean stale Xvfb lock file before starting the X server. We observed
that under cold-start storms, a previous Xvfb sometimes died and left
/tmp/.X99-lock + /tmp/.X11-unix/X99 behind, so the next start failed
with "Server is already active for display 99". Removing both files
makes Xvfb start cleanly every time.
- Replace `sleep 0.5` with a poll-for-X11-socket loop (up to 10s) plus
a 200ms post-socket buffer for listen()/accept() to settle. The
fixed sleep lost the race during concurrent cold inits, surfacing as
"Looks like you launched a headed browser without having a XServer
running" failures (~10% rate at 100-concurrent cold-start storm).
- Add _launch_with_retry helper in the handler: 3 attempts with linear
backoff (0.3s, 0.6s) on launch_context_async failures. Belt-and-
suspenders for whatever the entrypoint fix doesn't catch — a retry on
a now-warm container almost always succeeds.
Handler-side retry orchestration (lambda_handler.py):
- Add _classify_error() — maps Playwright errors to retry-strategy
overrides:
ERR_CERT_* -> --ignore-certificate-errors + 60s goto
Timeout exceeded -> 90s goto + 25s smart_wait cap
ERR_CONNECTION_TIMED_OUT -> same as Timeout
Returns None for unrecoverable site issues (DNS, SSL, refused, HTTP
4xx/5xx) — those bail immediately without burning a retry slot.
- Add _attempt_scrape() — extracted scrape body so the retry loop can
call it with overridden event dicts. Each attempt relaunches the
browser; uniform behavior across strategies.
- Rewrite _run() as a retry loop: first attempt uses event verbatim;
on a classifiable failure, merge the strategy's overrides into the
event and retry. Bounded by the new `retries` event field (default 1;
set to 0 to disable retry).
- Add _raise_with_history() — surfaces a final failure with a
retry_history block embedded in the error message so callers see
exactly what was tried before bailing. Successful invocations return
the standard response shape unchanged — no surprise fields.
INSTRUCTIONS.md updates:
- Bump function timeout recommendation from 60-120s to 120-180s. Under
retry, a Timeout-class first failure (30s) plus a longer-budget retry
(90s) plus cleanup can total ~120-130s; 180s leaves headroom.
- Document the new `retries` event field in the schema.
- Add a "Retry orchestration" subsection covering both layers (launch
retries and strategy retries) with the full strategy table.
Bench results on the 3454-site sample (seed=1):
v1 baseline (no fixes, c=100): 13.5% failure rate, $1.07
v2 (entrypoint Xvfb poll only, c=100): 9.9% failure rate, $1.11
v3 (cold-start fix + bench-side retry): 3.3% failure rate, $1.32
This change (handler retry, c=250): 2.1% failure rate, $1.13
The remaining 2.1% are all genuinely unrecoverable: DNS doesn't exist,
broken SSL, connection refused, 4xx/5xx responses, payload >6MB Lambda
limit. No retry logic can fix those.
* fix(lambda): merge extra_args on strategy retry instead of clobbering
A flat dict spread replaced caller-supplied extra_args (e.g.
--proxy-server=...) with the strategy's extra_args on a cert retry.
Append both lists so caller flags survive the merge.