scripts/research.py:
- New --differentiators CLI flag accepts a comma-separated list of brand
USPs (e.g. "women-owned, 24/7 service, no hidden fees"). Flows through
to research output as research.differentiators for the writing agent
to enforce verbatim in body content + AI Summary Nugget.
- New extract_missing_spokes() walks the top 3 ranking competitors'
internal-link anchors, filters out generic navigation (Home, Contact,
Privacy, FAQ, Login, social media, etc.) and broken-markdown image-
link leakage ([](href) nesting), and outputs a ranked
missing_spokes list. The list is the client's build-order priority
for filling the topical-silo gap.
- Compact + brief outputs both surface differentiators and missing_spokes.
scripts/lib/massive.py:
- _parse_markdown now returns a links list ([{text, url}]) from standard
[text](url) syntax. Skips image links, hash anchors, mailto/tel/javascript.
scripts/lib/dataforseo.py:
- New _extract_links static method walks page_content.main_topic and
secondary_topic, pulling anchor_text + url from primary_content[].urls.
- content_parse output now includes links field matching MassiveClient.
SKILL.md (v1.9.0 -> v1.9.1):
- Execution Protocol step 2 brief template gains Brand Differentiators /
USPs field. New paragraph instructs the agent to STOP and ASK the user
for differentiators if not provided up front.
- Section 12 Hub & Spoke Internal Linking gets a new Missing Spoke
Detection subsection requiring every generated page to append a
"## Recommended Spoke Pages" block built from missing_spokes data.
- Section 14 checklist expands 48 -> 51 points:
#49 Decision Fit (heading structure maps to buyer stage)
#50 Brand Identity (differentiators verbatim in chunks + nugget)
#51 Topical Silo (Recommended Spoke Pages block appended)
Passing threshold raised to 42/51.
references/quality-checklist.md: new v1.9.1 section detailing the three
new checks. Top-section reference updated to 51-point.
README.md, CHANGELOG.md, CLAUDE.md: version bumped, release-notes block
added, capability list updated. Historical version blocks restored to
their version-of-the-time checklist sizes (28, 34, 38, 41, 45, 48)
after over-greedy replace_all in prior commits.
Tests:
- 16 new tests in tests/test_research_v191.py covering --differentiators
parsing, domain normalization, generic-anchor filtering (including
nested-image-link leakage regression test), missing-spokes extraction
(same-domain filter, top-N respect, empty-input safety), markdown
link parsing in MassiveClient, and topic-tree link extraction in
DataForSEOClient.
- All 6 test files green.
Live smoke-tested against airport parking JFK with both flags:
- differentiators populated in compact output
- missing_spokes returned 12 semantic anchors after filtering
(SpotHero for Business, Reserve your spot, Parking details by lot,
EV charging stations, Learn about the JFK AirTrain, etc.)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds Massive (render.joinmassive.com) as the primary competitor content
parser when MASSIVE_API_TOKEN is configured. Returns clean rendered
markdown including JS-loaded content, which DataForSEO's
content_parsing/live endpoint has always missed.
Architecture:
- scripts/lib/massive.py wraps the /browser endpoint and outputs the
same shape as DataForSEOClient._extract_content (title, word_count,
headings, plain_text_size) so it's a drop-in replacement.
- scripts/research.py branches on Massive availability per URL. If
Massive errors or returns empty for any single URL, that URL falls
back to DataForSEO -- a partial Massive outage cannot break a run.
- env.py and .env.example pick up MASSIVE_API_TOKEN. New
creds["has_massive"] flag.
Scope intentionally narrow:
- SERP organic, PAA, and keyword data continue to come from DataForSEO.
- Massive's /search endpoint as of v1.9.0 only returns "also-searched"
query suggestions, not organic results. Tested directly to confirm
before scoping the integration.
Observability:
- research output now carries a content_parsers field summarizing
which parser handled each URL (e.g. {"massive": 4,
"dataforseo-fallback": 1}).
- Per-URL stderr log lines tag the parser: "Parsing content
(1/5, via massive)" or "via dataforseo".
Tests:
- 8 new unit tests in tests/test_massive.py covering the markdown
parser, the shape contract with DataForSEOClient, and client
construction.
- All 5 test files pass.
Live smoke-tested against airport parking JFK:
- With token: content_parsers = {"massive": 5}, real word counts on
all 5 competitors (Massive sees full body content; e.g. SpotHero
1132 words vs DataForSEO's previous 245).
- Without token: content_parsers = {"dataforseo": 4}, graceful
fallback to existing behavior.
No real token committed -- .env.example uses an empty placeholder.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Research pipeline (scripts/research.py):
- extract_meta_entities() mines bolded query-matched phrases from the
`highlighted` field of competitor SERP results, with inline
<b>/<strong>/** fallback parsing. These are the entities Google's
snippet generator already validated as relevant.
- extract_target_ngrams() tokenizes top 3 competitors' headings + titles,
filters via inlined English stopwords, returns top 5 bigrams/trigrams.
- detect_secondary_intent() maps the funnel-next intent (Orcas 1 dual
intent) with overrides for transactional title signals and brand-
domain dominance.
- All four signals (primary_intent, secondary_intent, meta_entities,
target_ngrams) now surface at top level of research output for direct
brief consumption. Compact + brief output formats updated.
- DataForSEOClient._extract_serp now passes through the `highlighted`
field per organic result (was being dropped silently).
Framework (SKILL.md):
- 41-point checklist -> 45-point checklist with Meta Entity Isolation,
N-Gram AI Alignment, Dual-Intent, and Status Code Governance checks.
Passing threshold raised to 36/45.
- New "Technical Codebase Execution Rules" section: when run inside a
project repo, detect framework (Next.js, Astro, Hugo, Jekyll, etc.),
inject semantic HTML into source files, emit Apache/.htaccess +
Nginx + next.config.js + Vercel snippets for 301/410 redirects.
- HARD RULES rewritten as positive naming guidance.
Docs:
- README.md "What It Actually Does" block expanded 13 -> 14 steps
reflecting dual-intent mapping, n-gram seeding, 301/410 governance.
- references/quality-checklist.md adds the 4 new pass/fail checks with
field references back to research.py output.
- CLAUDE.md framework features list updated.
Tests:
- New tests/test_research_v171.py with 13 tests covering meta-entity
extraction (highlighted + inline-tag + dedup), n-gram extraction
(stopword filter, top-N limit, empty input), tokenizer, and
secondary-intent funnel + overrides.
- Live smoke-tested against airport parking JFK: meta_entities returns
8 real bolded SERP phrases; target_ngrams returns "jfk airport",
"airport parking", "uncovered valet" etc. as expected.
All existing test files still pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds the off-page architecture missing from prior versions. Modern AEO
weighs Knowledge Graph inclusion and AI Overview impression share as
primary success signals, both gated by third-party corroboration. The
"inspector" layer cross-checks external mentions before trusting your
domain -- on-page perfection without off-page validation underperforms.
Changes:
- Core Belief #7: AEO Entity Validation via Owned Tier 1 Assets
- New SKILL.md Section 11A: Tributary Trust Protocol
- Tier 1 / Tier 2 asset definitions
- Companion content rule (substantive, not snippets)
- Network topology with upstream + lateral linking
- Derivation matrix mapping money-page chunks to tributary types
- Sequencing: tributaries before or same-day as money page
- Execution Protocol step 11: tributary deployment now mandatory for
commercial-intent and local pages, with all quality gates enforced
equally on off-page content
- scripts/tributary_gen.py: CLI that derives companion briefs from a
money page's 500-token chunks, inheriting {{VERIFY}} tags and entity
list, output to ~/Documents/SEO-AGI/tributaries/<slug>/
- CLAUDE.md updated with new script and framework features
Smoke-tested: generates 5 Tier 1 briefs (Google Sites, Medium,
Subreddit, Sheets, LinkedIn) with manifest, no warnings.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The on_page/content_parsing/live endpoint returns headings inside
main_topic[] / secondary_topic[] arrays (each with h_title + level),
not as flat h1/h2/h3 arrays. Old extractor looked for keys that don't
exist, so every research run reported Avg H2s: 0 / Avg H3s: 0
regardless of competitor depth. Word count and title were broken for
the same reason.
Fixes:
- _extract_headings now walks main_topic + secondary_topic by level
- _count_words computes from page_as_markdown with topic-text fallback
- _extract_title pulls H1 from markdown, falls back to main_topic[0]
- Replaced regression test that was asserting the broken shape
- Added 7 new tests covering the real DataForSEO response
Verified live: SpotHero JFK page now extracts 19 headings (11 H2 + 8 H3)
where it previously returned 0. Full pipeline: Avg H2s 9.2 / Avg H3s 3.5.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Writes pages that rank on Google AND get cited by LLMs. 500-token chunk
architecture, RAG targeting, Reddit Test quality gates, verification tags,
competitive data from DataForSEO/Ahrefs/SEMRush/GSC.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
GEO framework that writes pages ranking on Google AND getting cited by
LLMs. 500-token chunk architecture, Reddit Test quality gates,
verification tags, Not For You blocks, information gain enforcement.
Data layer: DataForSEO, GSC, Ahrefs MCP, SEMRush MCP.
21 files, all tests passing.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>