mirror of
https://github.com/gbessoni/seobuild-onpage.git
synced 2026-06-23 11:58:37 +02:00
The on_page/content_parsing/live endpoint returns headings inside main_topic[] / secondary_topic[] arrays (each with h_title + level), not as flat h1/h2/h3 arrays. Old extractor looked for keys that don't exist, so every research run reported Avg H2s: 0 / Avg H3s: 0 regardless of competitor depth. Word count and title were broken for the same reason. Fixes: - _extract_headings now walks main_topic + secondary_topic by level - _count_words computes from page_as_markdown with topic-text fallback - _extract_title pulls H1 from markdown, falls back to main_topic[0] - Replaced regression test that was asserting the broken shape - Added 7 new tests covering the real DataForSEO response Verified live: SpotHero JFK page now extracts 19 headings (11 H2 + 8 H3) where it previously returned 0. Full pipeline: Avg H2s 9.2 / Avg H3s 3.5. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>