fix(tiers): recognize more block/error page variants, add Tier 4 captcha parity, surface proxy/timing info

Found while running trawl against a large batch of real-world URLs: several
cases where the API returned 200 with content that was actually a blocked
page, an empty challenge stub, or Firefox's own error page. Each was a
detection gap where a tier didn't recognize the failure and reported it as a
successful scrape.

- Recognize Firefox's about:neterror/about:certerror page (browser never
  reached a server), Cloudflare's static "you have been blocked" WAF-deny
  page, and a lean CF challenge stub (blank title/body, just the bootstrap
  script) — the stub check is gated on page size since the same script
  snippet also appears on ordinary, fully-loaded CF pages as bot-management
  telemetry.
- Wire the existing isBlocked() status-code check (403/429/202) into Tiers 2
  and 3 — previously only Tier 1 checked status code, so a generic non-CF WAF
  deny that escalated to a browser tier was reported as a success.
- Bring Tier 4 up to parity with Tier 3: captcha solving and the same block
  detection. Sites that need Tier 4 for IP reputation can just as easily have
  an in-page captcha widget.
- Add proxyUsed: boolean to the response, set from the actual proxy used by
  the winning tier — previously the only signal was inferring from tier === 4,
  which doesn't distinguish "no proxy" from Tier 3's datacenter proxy.
- Attach the per-tier timings array to thrown errors via a new ScrapeError,
  and return it in /scrape's error response. The array was already being
  built in memory; it just never survived the throw, so failed requests gave
  a flat error string with no way to see which tier failed or why.
- Add process-level uncaughtException/unhandledRejection handlers. One target
  site's page threw a JS error that Camoufox/Firefox reports in a shape
  playwright-core's dispatcher doesn't expect, which crashed the entire
  process and dropped every in-flight request across all clients.
- Update the native API docs for the new response fields and error shape.

All additive — no existing fields changed shape. Full existing test suite
passes (58/58), and this is rebuilt/smoke-tested against latest dev.
This commit is contained in:
Erik Dasque
2026-07-07 15:58:54 +00:00
parent d7476274a2
commit 7d3204351c
9 changed files with 144 additions and 11 deletions
+18 -1
View File
@@ -46,6 +46,8 @@ interface ScrapeResult {
sessionCached: boolean // true if a cached session was used
timings: TierResult[] // per-tier attempt history
totalMs: number
captchasSolved?: string[] // captcha types solved on the page itself (e.g. ['turnstile'])
proxyUsed?: boolean // true if the winning tier routed through a proxy (Tier 3 datacenter pool or Tier 4 residential pool/override)
}
interface TierResult {
@@ -151,4 +153,19 @@ For 429 pool-exhaustion errors, the body is a **FlareSolverr v2 envelope** (same
}
```
For 400 / 503 / 500 the body is the native shape `{ "error": "Human-readable message" }`.
For 400 / 503 the body is the native shape `{ "error": "Human-readable message" }`.
For 500 errors raised after at least one tier was attempted, the body also includes the
per-tier attempt history, so a failed request is still fully diagnosable from the response
alone — no need to check server logs:
```json
{
"error": "All tiers exhausted. Last failure: http-403",
"timings": [
{ "tier": 1, "status": "needs-js", "durationMs": 50, "reason": "cloudflare-challenge" },
{ "tier": 3, "status": "blocked", "durationMs": 2942, "reason": "http-403" },
{ "tier": 4, "status": "blocked", "durationMs": 7890, "reason": "http-403" }
]
}
```