14 Commits
Author SHA1 Message Date
Greg BessoniandClaude Opus 4.7 36ff1607bc v1.9.1: Decision Fit + Brand Voice + Missing Spoke Detection
scripts/research.py:
- New --differentiators CLI flag accepts a comma-separated list of brand
  USPs (e.g. "women-owned, 24/7 service, no hidden fees"). Flows through
  to research output as research.differentiators for the writing agent
  to enforce verbatim in body content + AI Summary Nugget.
- New extract_missing_spokes() walks the top 3 ranking competitors'
  internal-link anchors, filters out generic navigation (Home, Contact,
  Privacy, FAQ, Login, social media, etc.) and broken-markdown image-
  link leakage ([![alt](img)](href) nesting), and outputs a ranked
  missing_spokes list. The list is the client's build-order priority
  for filling the topical-silo gap.
- Compact + brief outputs both surface differentiators and missing_spokes.

scripts/lib/massive.py:
- _parse_markdown now returns a links list ([{text, url}]) from standard
  [text](url) syntax. Skips image links, hash anchors, mailto/tel/javascript.

scripts/lib/dataforseo.py:
- New _extract_links static method walks page_content.main_topic and
  secondary_topic, pulling anchor_text + url from primary_content[].urls.
- content_parse output now includes links field matching MassiveClient.

SKILL.md (v1.9.0 -> v1.9.1):
- Execution Protocol step 2 brief template gains Brand Differentiators /
  USPs field. New paragraph instructs the agent to STOP and ASK the user
  for differentiators if not provided up front.
- Section 12 Hub & Spoke Internal Linking gets a new Missing Spoke
  Detection subsection requiring every generated page to append a
  "## Recommended Spoke Pages" block built from missing_spokes data.
- Section 14 checklist expands 48 -> 51 points:
    #49 Decision Fit (heading structure maps to buyer stage)
    #50 Brand Identity (differentiators verbatim in chunks + nugget)
    #51 Topical Silo (Recommended Spoke Pages block appended)
  Passing threshold raised to 42/51.

references/quality-checklist.md: new v1.9.1 section detailing the three
new checks. Top-section reference updated to 51-point.

README.md, CHANGELOG.md, CLAUDE.md: version bumped, release-notes block
added, capability list updated. Historical version blocks restored to
their version-of-the-time checklist sizes (28, 34, 38, 41, 45, 48)
after over-greedy replace_all in prior commits.

Tests:
- 16 new tests in tests/test_research_v191.py covering --differentiators
  parsing, domain normalization, generic-anchor filtering (including
  nested-image-link leakage regression test), missing-spokes extraction
  (same-domain filter, top-N respect, empty-input safety), markdown
  link parsing in MassiveClient, and topic-tree link extraction in
  DataForSEOClient.
- All 6 test files green.

Live smoke-tested against airport parking JFK with both flags:
- differentiators populated in compact output
- missing_spokes returned 12 semantic anchors after filtering
  (SpotHero for Business, Reserve your spot, Parking details by lot,
   EV charging stations, Learn about the JFK AirTrain, etc.)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-03 10:15:19 -04:00
Greg BessoniandClaude Opus 4.7 3ce6455b5c v1.9.0: Massive Web Render as primary content parser
Adds Massive (render.joinmassive.com) as the primary competitor content
parser when MASSIVE_API_TOKEN is configured. Returns clean rendered
markdown including JS-loaded content, which DataForSEO's
content_parsing/live endpoint has always missed.

Architecture:
- scripts/lib/massive.py wraps the /browser endpoint and outputs the
  same shape as DataForSEOClient._extract_content (title, word_count,
  headings, plain_text_size) so it's a drop-in replacement.
- scripts/research.py branches on Massive availability per URL. If
  Massive errors or returns empty for any single URL, that URL falls
  back to DataForSEO -- a partial Massive outage cannot break a run.
- env.py and .env.example pick up MASSIVE_API_TOKEN. New
  creds["has_massive"] flag.

Scope intentionally narrow:
- SERP organic, PAA, and keyword data continue to come from DataForSEO.
- Massive's /search endpoint as of v1.9.0 only returns "also-searched"
  query suggestions, not organic results. Tested directly to confirm
  before scoping the integration.

Observability:
- research output now carries a content_parsers field summarizing
  which parser handled each URL (e.g. {"massive": 4,
  "dataforseo-fallback": 1}).
- Per-URL stderr log lines tag the parser: "Parsing content
  (1/5, via massive)" or "via dataforseo".

Tests:
- 8 new unit tests in tests/test_massive.py covering the markdown
  parser, the shape contract with DataForSEOClient, and client
  construction.
- All 5 test files pass.

Live smoke-tested against airport parking JFK:
- With token: content_parsers = {"massive": 5}, real word counts on
  all 5 competitors (Massive sees full body content; e.g. SpotHero
  1132 words vs DataForSEO's previous 245).
- Without token: content_parsers = {"dataforseo": 4}, graceful
  fallback to existing behavior.

No real token committed -- .env.example uses an empty placeholder.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-20 16:46:14 -04:00
Greg BessoniandClaude Opus 4.7 2e8e03e110 v1.8.0: Gemini 3.5 Flash RAG optimization + off-page trust expansion
SKILL.md:
- New Section 6 subsection "DOM Vectoring & Shard Extraction
  Compliance" -- Google's AI Overviews are now built by Gemini 3.5
  Flash via a RAG pipeline that extracts shards from the raw HTML
  DOM. JSON-LD in <head> is no longer sufficient on its own; critical
  data points must live in front-facing <table> markup or inline
  RDFa spans visible to a clean-session crawler.
- Section 11A Tributary Trust Protocol gains two new Tier 1 assets:
    * Trust Pilot -- shifts brand description vectoring in Gemini /
      ChatGPT inside 48 hours of publication
    * Off-Page Schema Injection -- Organization/Person JSON-LD on
      Cloud Pages and PRs with GBP CID backlinks blocks NavBoost
      rank-shuffling during A/B exposure tests
- Network spread requirement updated from 4/5 to 5/7 Tier 1 assets.
- Frontmatter description now mentions Gemini 3.5 Flash RAG.

Section 14 checklist (45 -> 48 points, threshold 36/45 -> 39/48):
  #46 Trust Pilot entity profiling with target bigrams
  #47 Off-page cross-cutting Organization/Person schema to GBP
  #48 Critical data points visible in raw HTML DOM (not just JSON-LD)

README.md, CHANGELOG.md, CLAUDE.md, references/quality-checklist.md
updated to match.

All existing tests pass (test_dataforseo, test_env, test_serp_analyze,
test_research_v171). No code changes -- v1.8.0 is a framework/protocol
release that extends existing structural rules.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-20 16:23:35 -04:00
Greg BessoniandClaude Opus 4.7 a80df5cd34 v1.7.1: LLM retrieval signals + codebase execution rules
Research pipeline (scripts/research.py):
- extract_meta_entities() mines bolded query-matched phrases from the
  `highlighted` field of competitor SERP results, with inline
  <b>/<strong>/** fallback parsing. These are the entities Google's
  snippet generator already validated as relevant.
- extract_target_ngrams() tokenizes top 3 competitors' headings + titles,
  filters via inlined English stopwords, returns top 5 bigrams/trigrams.
- detect_secondary_intent() maps the funnel-next intent (Orcas 1 dual
  intent) with overrides for transactional title signals and brand-
  domain dominance.
- All four signals (primary_intent, secondary_intent, meta_entities,
  target_ngrams) now surface at top level of research output for direct
  brief consumption. Compact + brief output formats updated.
- DataForSEOClient._extract_serp now passes through the `highlighted`
  field per organic result (was being dropped silently).

Framework (SKILL.md):
- 41-point checklist -> 45-point checklist with Meta Entity Isolation,
  N-Gram AI Alignment, Dual-Intent, and Status Code Governance checks.
  Passing threshold raised to 36/45.
- New "Technical Codebase Execution Rules" section: when run inside a
  project repo, detect framework (Next.js, Astro, Hugo, Jekyll, etc.),
  inject semantic HTML into source files, emit Apache/.htaccess +
  Nginx + next.config.js + Vercel snippets for 301/410 redirects.
- HARD RULES rewritten as positive naming guidance.

Docs:
- README.md "What It Actually Does" block expanded 13 -> 14 steps
  reflecting dual-intent mapping, n-gram seeding, 301/410 governance.
- references/quality-checklist.md adds the 4 new pass/fail checks with
  field references back to research.py output.
- CLAUDE.md framework features list updated.

Tests:
- New tests/test_research_v171.py with 13 tests covering meta-entity
  extraction (highlighted + inline-tag + dedup), n-gram extraction
  (stopword filter, top-N limit, empty input), tokenizer, and
  secondary-intent funnel + overrides.
- Live smoke-tested against airport parking JFK: meta_entities returns
  8 real bolded SERP phrases; target_ngrams returns "jfk airport",
  "airport parking", "uncovered valet" etc. as expected.

All existing test files still pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-07 08:05:24 -04:00
Greg BessoniandClaude Opus 4.7 9a9e424c33 v1.6.0: ICP-driven content + local trust signals
- Page brief template now requires Ideal Customer Persona (ICP)
- Deep Entity History & Identity Tags in Section 4 SEAT signals
- Self-Placement Rule for listicles (objective #1 with tradeoffs)
- Keyword Cannibalization governance in Section 9
- Quality checklist expanded 38 -> 41, passing threshold 33/41

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-28 10:06:21 -04:00
Greg Bessoni 713a428bfd Update README for v1.5.0: semantic HTML containers, proof-term proximity, structural signals section 2026-04-03 09:36:57 -04:00
Greg Bessoni 1ae0a07911 v1.5.0: rename to seobuild-onpage, forensic SEO protocols, 38-point checklist 2026-04-02 10:05:05 -04:00
Greg BessoniandClaude Opus 4.6 ce774cb6a2 Update install URLs to new repo name (on-page-agent)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-01 08:40:09 -04:00
Greg BessoniandClaude Opus 4.6 6f59755a0e v1.4.0: March 2026 update protocols - NavBoost click signals, AI Overview optimization, FHASS, geographic relevance, banned patterns
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-01 08:28:55 -04:00
Greg BessoniandClaude Opus 4.6 3575d057f0 Add version header and changelog callouts to README (v1.3.0)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 09:03:51 -04:00
Greg BessoniandClaude Opus 4.6 dfae798583 Update README with v1.3.0 features: AI nuggets, experiment blocks, 28-point checklist
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 09:00:49 -04:00
Greg BessoniandClaude Opus 4.6 f773d53a96 Add hard rules: ban "beefy" in all output, enforce mandatory scorecard
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-20 05:33:21 -04:00
Greg BessoniandClaude Opus 4.6 b5c1417af2 SEO-AGI: GEO framework skill for Claude Code, OpenClaw, and Codex
Writes pages that rank on Google AND get cited by LLMs. 500-token chunk
architecture, RAG targeting, Reddit Test quality gates, verification tags,
competitive data from DataForSEO/Ahrefs/SEMRush/GSC.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-19 08:11:06 -04:00
Greg BessoniandClaude Opus 4.6 ae47fcccba seo-agi: The GEO framework skill for Claude Code + OpenClaw
GEO framework that writes pages ranking on Google AND getting cited by
LLMs. 500-token chunk architecture, Reddit Test quality gates,
verification tags, Not For You blocks, information gain enforcement.

Data layer: DataForSEO, GSC, Ahrefs MCP, SEMRush MCP.
21 files, all tests passing.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-18 10:56:28 -04:00