mirror of
https://github.com/germondai/trawl.git
synced 2026-08-17 12:11:23 +02:00
fix(tiers): recognize more block/error page variants, add Tier 4 captcha parity, surface proxy/timing info
Found while running trawl against a large batch of real-world URLs: several cases where the API returned 200 with content that was actually a blocked page, an empty challenge stub, or Firefox's own error page. Each was a detection gap where a tier didn't recognize the failure and reported it as a successful scrape. - Recognize Firefox's about:neterror/about:certerror page (browser never reached a server), Cloudflare's static "you have been blocked" WAF-deny page, and a lean CF challenge stub (blank title/body, just the bootstrap script) — the stub check is gated on page size since the same script snippet also appears on ordinary, fully-loaded CF pages as bot-management telemetry. - Wire the existing isBlocked() status-code check (403/429/202) into Tiers 2 and 3 — previously only Tier 1 checked status code, so a generic non-CF WAF deny that escalated to a browser tier was reported as a success. - Bring Tier 4 up to parity with Tier 3: captcha solving and the same block detection. Sites that need Tier 4 for IP reputation can just as easily have an in-page captcha widget. - Add proxyUsed: boolean to the response, set from the actual proxy used by the winning tier — previously the only signal was inferring from tier === 4, which doesn't distinguish "no proxy" from Tier 3's datacenter proxy. - Attach the per-tier timings array to thrown errors via a new ScrapeError, and return it in /scrape's error response. The array was already being built in memory; it just never survived the throw, so failed requests gave a flat error string with no way to see which tier failed or why. - Add process-level uncaughtException/unhandledRejection handlers. One target site's page threw a JS error that Camoufox/Firefox reports in a shape playwright-core's dispatcher doesn't expect, which crashed the entire process and dropped every in-flight request across all clients. - Update the native API docs for the new response fields and error shape. All additive — no existing fields changed shape. Full existing test suite passes (58/58), and this is rebuilt/smoke-tested against latest dev.
This commit is contained in:
@@ -46,6 +46,8 @@ interface ScrapeResult {
|
||||
sessionCached: boolean // true if a cached session was used
|
||||
timings: TierResult[] // per-tier attempt history
|
||||
totalMs: number
|
||||
captchasSolved?: string[] // captcha types solved on the page itself (e.g. ['turnstile'])
|
||||
proxyUsed?: boolean // true if the winning tier routed through a proxy (Tier 3 datacenter pool or Tier 4 residential pool/override)
|
||||
}
|
||||
|
||||
interface TierResult {
|
||||
@@ -151,4 +153,19 @@ For 429 pool-exhaustion errors, the body is a **FlareSolverr v2 envelope** (same
|
||||
}
|
||||
```
|
||||
|
||||
For 400 / 503 / 500 the body is the native shape `{ "error": "Human-readable message" }`.
|
||||
For 400 / 503 the body is the native shape `{ "error": "Human-readable message" }`.
|
||||
|
||||
For 500 errors raised after at least one tier was attempted, the body also includes the
|
||||
per-tier attempt history, so a failed request is still fully diagnosable from the response
|
||||
alone — no need to check server logs:
|
||||
|
||||
```json
|
||||
{
|
||||
"error": "All tiers exhausted. Last failure: http-403",
|
||||
"timings": [
|
||||
{ "tier": 1, "status": "needs-js", "durationMs": 50, "reason": "cloudflare-challenge" },
|
||||
{ "tier": 3, "status": "blocked", "durationMs": 2942, "reason": "http-403" },
|
||||
{ "tier": 4, "status": "blocked", "durationMs": 7890, "reason": "http-403" }
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
Reference in New Issue
Block a user