mirror of
https://github.com/gbessoni/seobuild-onpage.git
synced 2026-06-23 11:58:37 +02:00
Adds Massive (render.joinmassive.com) as the primary competitor content
parser when MASSIVE_API_TOKEN is configured. Returns clean rendered
markdown including JS-loaded content, which DataForSEO's
content_parsing/live endpoint has always missed.
Architecture:
- scripts/lib/massive.py wraps the /browser endpoint and outputs the
same shape as DataForSEOClient._extract_content (title, word_count,
headings, plain_text_size) so it's a drop-in replacement.
- scripts/research.py branches on Massive availability per URL. If
Massive errors or returns empty for any single URL, that URL falls
back to DataForSEO -- a partial Massive outage cannot break a run.
- env.py and .env.example pick up MASSIVE_API_TOKEN. New
creds["has_massive"] flag.
Scope intentionally narrow:
- SERP organic, PAA, and keyword data continue to come from DataForSEO.
- Massive's /search endpoint as of v1.9.0 only returns "also-searched"
query suggestions, not organic results. Tested directly to confirm
before scoping the integration.
Observability:
- research output now carries a content_parsers field summarizing
which parser handled each URL (e.g. {"massive": 4,
"dataforseo-fallback": 1}).
- Per-URL stderr log lines tag the parser: "Parsing content
(1/5, via massive)" or "via dataforseo".
Tests:
- 8 new unit tests in tests/test_massive.py covering the markdown
parser, the shape contract with DataForSEOClient, and client
construction.
- All 5 test files pass.
Live smoke-tested against airport parking JFK:
- With token: content_parsers = {"massive": 5}, real word counts on
all 5 competitors (Massive sees full body content; e.g. SpotHero
1132 words vs DataForSEO's previous 245).
- Without token: content_parsers = {"dataforseo": 4}, graceful
fallback to existing behavior.
No real token committed -- .env.example uses an empty placeholder.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>