mirror of
https://github.com/gbessoni/seobuild-onpage.git
synced 2026-06-23 11:58:37 +02:00
scripts/research.py:
- New --differentiators CLI flag accepts a comma-separated list of brand
USPs (e.g. "women-owned, 24/7 service, no hidden fees"). Flows through
to research output as research.differentiators for the writing agent
to enforce verbatim in body content + AI Summary Nugget.
- New extract_missing_spokes() walks the top 3 ranking competitors'
internal-link anchors, filters out generic navigation (Home, Contact,
Privacy, FAQ, Login, social media, etc.) and broken-markdown image-
link leakage ([](href) nesting), and outputs a ranked
missing_spokes list. The list is the client's build-order priority
for filling the topical-silo gap.
- Compact + brief outputs both surface differentiators and missing_spokes.
scripts/lib/massive.py:
- _parse_markdown now returns a links list ([{text, url}]) from standard
[text](url) syntax. Skips image links, hash anchors, mailto/tel/javascript.
scripts/lib/dataforseo.py:
- New _extract_links static method walks page_content.main_topic and
secondary_topic, pulling anchor_text + url from primary_content[].urls.
- content_parse output now includes links field matching MassiveClient.
SKILL.md (v1.9.0 -> v1.9.1):
- Execution Protocol step 2 brief template gains Brand Differentiators /
USPs field. New paragraph instructs the agent to STOP and ASK the user
for differentiators if not provided up front.
- Section 12 Hub & Spoke Internal Linking gets a new Missing Spoke
Detection subsection requiring every generated page to append a
"## Recommended Spoke Pages" block built from missing_spokes data.
- Section 14 checklist expands 48 -> 51 points:
#49 Decision Fit (heading structure maps to buyer stage)
#50 Brand Identity (differentiators verbatim in chunks + nugget)
#51 Topical Silo (Recommended Spoke Pages block appended)
Passing threshold raised to 42/51.
references/quality-checklist.md: new v1.9.1 section detailing the three
new checks. Top-section reference updated to 51-point.
README.md, CHANGELOG.md, CLAUDE.md: version bumped, release-notes block
added, capability list updated. Historical version blocks restored to
their version-of-the-time checklist sizes (28, 34, 38, 41, 45, 48)
after over-greedy replace_all in prior commits.
Tests:
- 16 new tests in tests/test_research_v191.py covering --differentiators
parsing, domain normalization, generic-anchor filtering (including
nested-image-link leakage regression test), missing-spokes extraction
(same-domain filter, top-N respect, empty-input safety), markdown
link parsing in MassiveClient, and topic-tree link extraction in
DataForSEOClient.
- All 6 test files green.
Live smoke-tested against airport parking JFK with both flags:
- differentiators populated in compact output
- missing_spokes returned 12 semantic anchors after filtering
(SpotHero for Business, Reserve your spot, Parking details by lot,
EV charging stations, Learn about the JFK AirTrain, etc.)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
14 KiB
14 KiB
Changelog
All notable changes to seo-agi are documented here.
[1.9.1] - 2026-05-21
Added
--differentiatorsCLI flag onscripts/research.py. Comma-separated brand USPs (e.g.--differentiators="women-owned, 24/7 service, no hidden fees") flow through to the research JSON + brief output asresearch.differentiatorsfor the writing agent to enforce verbatim.- Missing Spoke Detection in
scripts/research.py(extract_missing_spokes()). Walks the top 3 ranking competitors' internal-link anchors, filters out generic navigation (Home, Contact, Privacy, FAQ, etc.) and nested-image-link leakage, returns a rankedmissing_spokeslist in the output payload. - Link extraction in both content parsers:
MassiveClient._parse_markdownnow returns alinkslist ([{text, url}]) from standard markdown link syntax;DataForSEOClient._extract_linkswalksprimary_content[].urls[]acrossmain_topicandsecondary_topic. - SKILL.md Section 12 (Hub & Spoke) now requires every generated page to append a
## Recommended Spoke Pagesblock built frommissing_spokes. - SKILL.md Execution Protocol now stops and asks the user for Brand Differentiators / USPs if not provided up front. Pages built without explicit differentiators are flagged as Reddit-Test failure risk.
- 51-point quality checklist with #49 Decision Fit, #50 Brand Identity, #51 Topical Silo. Passing threshold raised to 42/51.
- 15 new tests in
tests/test_research_v191.pycovering differentiator parsing, domain normalization, generic-anchor filtering (including nested-image-link leakage), missing-spokes extraction, and link extraction in both content parsers.
Changed
references/quality-checklist.mdappended with a v1.9.1 section detailing the three new checks and threshold.
[1.9.0] - 2026-05-20
Added
- Massive Web Render integration (
scripts/lib/massive.py). WhenMASSIVE_API_TOKENis configured, the/browserendpoint becomes the primary competitor content parser. Returns clean rendered markdown including JS-loaded content -- the gap DataForSEO'scontent_parsing/livehas always had. - Per-URL graceful fallback: if Massive errors or returns empty for any single URL, that URL falls back to DataForSEO. Partial Massive outage cannot break a run.
content_parsersfield in research output records which parser handled each URL (e.g.{"massive": 4, "dataforseo-fallback": 1}). Visible in saved JSON and printed in compact output's per-URL log lines (via massive/via dataforseo).MASSIVE_API_TOKENadded to.env.exampleandenv.pycredential loader. Newcreds["has_massive"]flag.- 8 new unit tests in
tests/test_massive.pycovering the markdown parser, output-shape contract with DataForSEOClient, and client construction.
Changed
scripts/research.pycontent-parsing step now branches on Massive availability per URL.- SKILL.md Data Cascade table reorganized to surface Massive as Priority 1 for content parsing, with DataForSEO as the SERP + keyword data backbone and per-URL fallback.
Not Changed
- SERP organic results, PAA, and keyword data still come from DataForSEO. Massive's
/searchendpoint as of v1.9.0 only returns "also-searched" suggestions, not organic results. SERP path is unchanged. - All existing tests pass. Heading extraction, n-gram analysis, meta-entity isolation, intent detection, and the 48-point checklist behave identically whichever content parser is used.
[1.8.0] - 2026-05-08
Added
- DOM Vectoring & Shard Extraction Compliance (SKILL.md Section 6): documents the constraint that Google's AI Overviews are now built by Gemini 3.5 Flash via a RAG pipeline that extracts structural "shards" from the raw HTML DOM. Critical data must live in front-facing
<table>markup or inline RDFa spans -- JSON-LD in<head>is no longer sufficient on its own. - Trust Pilot as Tier 1 Tributary (Section 11A): added to the Tributary Trust Protocol table as a first-class Tier 1 asset. Cited as a highly weighted trust/relevance signal for LLMs that shifts brand description vectoring in Gemini and ChatGPT within 48 hours.
- Off-Page Schema Injection (Section 11A): new Tier 1 tactic. Embedding
OrganizationandPersonJSON-LD schema in Cloud Pages and PRs with explicit GBP CID backlinks blocks Google NavBoost from rank-shuffling the money page during A/B exposure tests. - 48-point quality checklist: adds #46 (Trust Pilot entity profiling with target bigrams), #47 (off-page Organization/Person schema mapping), #48 (DOM-visible critical data points). Passing threshold raised to 39/48.
Changed
- Tributary network spread requirement updated from "4 of the 5 Tier 1 assets" to "5 of the 7 Tier 1 assets" to account for the two new asset types.
- SKILL.md frontmatter description now mentions Gemini 3.5 Flash RAG optimization.
[1.7.1] - 2026-05-07
Added
- Meta-Specific Entity Isolation:
extract_meta_entities()mines bolded query-matched phrases from thehighlightedfield of competitor SERP results (with inline<b>/<strong>/**fallback parsing). These are the entities Google's snippet generator already validated as relevant -- a stronger signal than body-content extraction. - Bigram/Trigram AI Alignment:
extract_target_ngrams()tokenizes the top 3 ranking competitors' headings + titles, filters via an inlined English stopword list, and returns the top 5 bigrams and top 5 trigrams. Output seeds the AI Summary Nugget for LLM-retrieval token overlap. - Primary + Secondary Intent (Orcas 1):
detect_secondary_intent()maps the funnel-next intent (informational → commercial → transactional → navigational), with overrides for transactional title signals (book/reserve/buy) and brand-domain dominance in the top 5. - 45-point quality checklist: adds Meta Entity Isolation, N-Gram Alignment, Dual-Intent, and Status Code Governance checks. Passing threshold raised to 36/45.
- Technical Codebase Execution Rules (SKILL.md): when the skill runs inside a project repo, it detects the framework (Next.js, Astro, Gatsby, Hugo, etc.), injects semantic HTML into source files, and emits
.htaccess/ Nginx /next.config.jsredirect snippets for 301/410 recommendations. - The 410 Prune Protocol in the rewrite workflow: every legacy URL gets an explicit 301 (preserve equity) or 410 (prune) recommendation. Silent leave-as-is is no longer acceptable.
Changed
dataforseo.py:_extract_serpnow passes through thehighlightedfield on every organic result (was being dropped). Required for Meta Entity Isolation.research.py: research output now surfacesprimary_intent,secondary_intent,meta_entities, andtarget_ngramsat top level (not buried inanalysis) for direct brief consumption.- README.md "What It Actually Does" block expanded from 13 to 14 steps reflecting dual-intent mapping (step 5), n-gram seeding (step 8), and 301/410 governance (step 12).
- HARD RULES rewritten in SKILL.md: replaced the negative "never use codename X" rule with positive naming guidance ("framework is seo-agi / seobuild-onpage; no prior internal codenames in any output").
Tests
- Added
tests/test_research_v171.pywith 13 new tests covering meta-entity extraction (highlighted field + inline-tag fallback + dedup), n-gram extraction (stopword filtering, top-N limiting, empty-input safety), tokenizer behavior, and secondary-intent funnel + override logic.
[1.7.0] - 2026-04-30
Added
- Tributary Trust Protocol (Section 11A): off-page architecture for AEO entity validation. Defines Tier 1 owned assets (Google Sites, Medium, Subreddits, Google Sheets, LinkedIn) as tributaries that feed entity signal to the money page. Includes companion-content rules, network topology, derivation matrix, and sequencing requirements.
- Core Belief #7: AEO Entity Validation via Owned Tier 1 Assets. Knowledge Graph inclusion and AI Overview impression share are now primary success signals, gated by off-page corroboration.
scripts/tributary_gen.py: CLI tool that reads a money page, extracts its 500-token chunks + entities +{{VERIFY}}tags, and outputs derived companion briefs to~/Documents/SEO-AGI/tributaries/<slug>/with a manifest mapping each draft to its host platform.- Execution Protocol step 11: tributary deployment is now mandatory for commercial-intent and local pages. All quality gates (Reddit Test, Information Gain, Prove-It,
{{VERIFY}}resolution, Section 9 banned patterns, Entity Consensus) apply equally to off-page content -- thin tributaries net-harm the money page's entity signal.
Changed
- CLAUDE.md directory structure now lists
scripts/tributary_gen.py - CLAUDE.md framework features updated with Tributary Trust capability
[1.6.1] - 2026-04-28
Fixed
- Heading extraction returned empty for every competitor page. DataForSEO's
on_page/content_parsing/liveendpoint returns headings insidemain_topic[]andsecondary_topic[]arrays (each withh_titleandlevel), not as flath1/h2/h3arrays. The old_extract_headingslooked for keys that don't exist, soAvg H2s: 0, Avg H3s: 0for every research run regardless of competitor depth. Now correctly walks the topic tree. - Word count returned 0 for every competitor.
plain_text_word_countis not a field DataForSEO returns. Now computes frompage_as_markdown(preferred) with a fallback walkingprimary_content[].text. - Title extraction returned empty.
header.titledoesn't exist; replaced with markdown H1 detection plus amain_topic[0].h_titlefallback. - Added regression tests covering the real API response shape so this can't silently break again.
[1.6.0] - 2026-04-15
Added
- Ideal Customer Persona (ICP): Page brief template now requires a defined ICP with demographics, psychographics, and specific pain points. Content maps to the actual reader, not a generic audience.
- Deep Entity History & Identity Tags: Founding dates, generational ownership, and identity attributes (women-owned, veteran-owned) are now explicit entity signals in Section 4 SEAT Signals.
- The Self-Placement Rule: Listicle guidance now explicitly allows ranking the client #1, provided the entry is objective with a specific use-case and honest tradeoffs.
- Keyword Cannibalization Governance: Section 9 "Never Do" list now prohibits creating pages that compete with existing URLs for the same intent. Sales-focused duplicates get tagged with
noindexrecommendation.
Changed
- Quality checklist expanded from 38 to 41 items (ICP alignment, entity history, cannibalization)
- Minimum passing score raised to 33/41
- Both brief templates (Page Brief Template + Execution Protocol) updated with ICP field
[1.3.0] - 2026-03-25
Added
- AI Summary Nugget: Mandatory 200-character fact-dense block at top of every page, designed for LLM scrapers (Perplexity, Gemini, ChatGPT) to cite as a consensus source
- Original Research Block: Every page must include a data experiment or first-hand observation section to satisfy Google's Experience (E-E-A-T) signal
- Map Traffic Shifting: Local SEO instruction to link from high-traffic informational pages to map embeds, shifting user interaction signals toward local intent
- Spam Resilience Logic: Quality checklist now prioritizes technical relevance density over "human tone" -- factually perfect content is not downgraded for sounding clinical
- Recursive Fact-Checking: New execution step validates every claim against 2+ high-ranking sources for Entity Consensus before delivery
Changed
- Quality checklist expanded from 24 to 28 items
- Minimum passing score raised to 22/28
- Execution protocol now has 11 steps (was 10)
[1.2.0] - 2026-03-25
Added
- Anti-spam ranking signals: single H1 rule, no EMQ in meta descriptions, no keyword-stuffed alt text, no duplicate content, internal linking requirements
- EMQ allowed in title and URL, banned in H2/H3/H4 subheadings
- Interactive elements section (cost calculators, widgets) to defend against AI Overview traffic loss
- Broken backlink monitoring guidance
- "Boat anchor" page culling (410 status for unindexable cruft)
Changed
- Quality checklist expanded from 20 to 24 items (added H1, meta desc, heading, alt text checks)
[1.1.0] - 2026-03-24
Added
- Hard rule: framework is named seo-agi only; prior internal codenames must not appear in output
- Mandatory printed scorecard at end of every page output (no exceptions)
- FAQ/PAA section required (3+ questions, FAQPage schema)
- JSON-LD schema block required per page type
- Hub/spoke internal links required in every output
- RAG Targeting section (write for AI retrieval, not keyword volume)
- Topical Circle Audit (stay inside core service topic, noindex strays)
- Off-Page Sequencing (establish external brand footprint before on-page)
- Reddit Subdomain Indexing (subdomains over standard posts for AI retrieval)
- Ask Maps / Conversational GBP Optimization
Changed
- Quality checklist expanded from base to 20 items
- Scorecard enforcement language strengthened ("INCOMPLETE without this table")
[1.0.0] - 2026-03-23
Added
- Initial release: GEO framework skill for Claude Code, OpenClaw, and Codex
- 500-token chunk architecture for Google AI retrieval
- SEAT signals (Semantic + E-E-A-T + Entity/Knowledge Graph)
- Google AI Search 7 ranking signals (Gecko, Jetstream, BM25, PCTR, Freshness, Boost/Bury)
- Reddit Test, Prove-It Details, Not For You, Information Gain quality gates
- Verification tagging system (VERIFY, RESEARCH NEEDED, SOURCE NEEDED)
- DataForSEO integration with graceful no-creds fallback
- MCP tool integration (Ahrefs, SEMRush)
- Google Search Console pull scripts
- Skill root discovery loop (works across Claude Code, OpenClaw, Codex, Gemini)
- Reference files: page templates, schema patterns, quality checklist
- Mock mode for testing without API keys