seo-agi: The GEO framework skill for Claude Code + OpenClaw

GEO framework that writes pages ranking on Google AND getting cited by
LLMs. 500-token chunk architecture, Reddit Test quality gates,
verification tags, Not For You blocks, information gain enforcement.

Data layer: DataForSEO, GSC, Ahrefs MCP, SEMRush MCP.
21 files, all tests passing.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
Greg Bessoni
2026-03-18 10:56:28 -04:00
co-authored by Claude Opus 4.6
commit ae47fcccba
21 changed files with 3139 additions and 0 deletions
+7
View File
@@ -0,0 +1,7 @@
__pycache__/
*.pyc
.env
*.egg-info/
dist/
build/
.DS_Store
+59
View File
@@ -0,0 +1,59 @@
# CLAUDE.md
This is **seo-agi** (The Beefy SEO Skill) -- a Claude Code skill for Generative Engine Optimization. It writes pages that rank on Google AND get cited by LLMs.
This is not a generic SEO prompt. It enforces 500-token chunk architecture, Reddit Test quality gates, verification tags, "Not For You" blocks, and real competitive data from DataForSEO/Ahrefs/SEMRush/GSC.
## Structure
```
SKILL.md -- The framework (beefy SEO brain + data layer integration)
SPEC.md -- Technical architecture for the data layer
scripts/
research.py -- CLI: SERP research via DataForSEO
gsc_pull.py -- CLI: Google Search Console data
setup.py -- Interactive first-run config
lib/
env.py -- Config loader
dataforseo.py -- DataForSEO REST API client
serp_analyze.py -- Content gap analysis engine
gsc_client.py -- Google Search Console client
references/
page-templates.md -- Structural templates by page type
schema-patterns.md -- JSON-LD schema patterns
fixtures/ -- Mock data for testing without API calls
tests/ -- Unit tests
```
## The Framework (SKILL.md)
The SKILL.md is the living document. It contains:
- Core belief system (anti-generic, LLM retrieval, entity consensus)
- Google AI Search 7 ranking signals
- 500-token chunk architecture
- SEAT signals (Semantic + E-E-A-T + Entity/Knowledge Graph)
- Quality gates: Reddit Test, Prove-It Details, Not For You, Information Gain Test
- Verification tagging system ({{VERIFY}}, {{RESEARCH NEEDED}}, {{SOURCE NEEDED}})
- Vertical-specific instructions (airport/parking, local service, listicle, comparison)
- LLM/AEO citation strategy
- Hub & spoke internal linking
- Execution protocol with data layer integration
## Running Tests
```bash
python3 tests/test_env.py
python3 tests/test_serp_analyze.py
python3 tests/test_dataforseo.py
# Mock mode research:
python3 scripts/research.py "test keyword" --mock
```
## Style
- No em dashes in any output
- No "nestled", no "in today's fast-paced world", no "whether you're a...or a..."
- Numbers and specifics over adjectives
- Every claim tagged with {{VERIFY}} or {{SOURCE NEEDED}}
- Tables are mandatory for comparisons -- never simulate with bullet points
+190
View File
@@ -0,0 +1,190 @@
# SEO-AGI
### One command. Competitive data in. Ranking pages out.
Most SEO tools tell you what's wrong with your site. This one writes the pages.
`/seoagi "airport parking JFK"` pulls the current SERP, analyzes what's ranking, finds the gaps in their content, and writes you a complete page -- with the heading structure, depth, FAQ section, and schema markup that actually competes. Not thin content. Not keyword-stuffed filler. Pages backed by live data from the tools the pros use.
**I built this because I got tired of the gap between "SEO audit" and "published page."** I've been doing SEO for 20+ years in ground transportation (1M+ bookings, 2M+ rides across my companies). The workflow was always the same: pull SERP data, analyze competitors, find gaps, write brief, write page, add schema, publish. Over and over. So I turned that entire workflow into a single skill that any AI agent can execute.
The result? I used this to research a competitor's best-performing pages, built equivalent content with `/seoagi`, bought the exact-match domains, and every single page is ranking on page 1. That's not theory. That's the workflow.
---
## What It Actually Does
```
You: /seoagi "best project management tools 2026"
SEO-AGI:
1. Pulls SERP top 10 via DataForSEO
2. Parses competitor content (word count, headings, topics covered)
3. Extracts People Also Ask questions
4. Pulls related keywords with search volumes
5. Detects search intent (informational vs commercial vs transactional)
6. Generates a data-driven content brief
7. Writes the complete page (Markdown + YAML frontmatter)
8. Adds FAQ section from real PAA data
9. Generates JSON-LD schema markup
10. Validates against a quality checklist
```
For rewrites, point it at any URL. It compares your page against the current top 3 ranking competitors, identifies exactly what you're missing, and rewrites with a change summary explaining every edit.
---
## The SEO Knowledge Inside
This isn't a wrapper around "write me an SEO article." The skill encodes strategies from the best in the game:
**Traditional SEO**
- Intent-first content architecture (match what searchers actually want, not what you think the keyword means)
- Competitive word count targeting (page length based on what's ranking, not arbitrary "write 2000 words")
- Heading hierarchy derived from SERP analysis (not templates, not guesswork)
- People Also Ask coverage as FAQ sections (answer the questions Google already knows people are asking)
- Schema markup patterns by page type (FAQPage, LocalBusiness, HowTo, Product, BreadcrumbList)
- Internal linking suggestions based on actual site data from GSC
**GEO / LLM SEO (Generative Engine Optimization)**
- Content structured for AI citation (Perplexity, ChatGPT, Google AI Overviews)
- Entity-rich writing that LLMs can extract and reference
- Depth-over-length philosophy (comprehensive coverage that becomes the authoritative source)
- FAQ patterns that match how AI systems parse and surface answers
- Data-backed claims that AI systems prefer to cite over vague assertions
**The quality checklist every page runs through:**
- Title tag <60 chars with target keyword? Check.
- Meta description <155 chars with CTA? Check.
- Single H1 mirroring the title? Check.
- Logical H2 > H3 hierarchy? Check.
- At least 3 PAA questions answered? Check.
- Specific data/statistics (not vague claims)? Check.
- JSON-LD schema appropriate to page type? Check.
- Word count within competitive range? Check.
- No keyword stuffing? Check.
Pages scoring below 80% get flagged with specific items to fix. Below 60% = rewrite.
---
## Data Integrations (BYOK)
Bring your own API keys. Use one, use all. The skill adapts:
| Integration | What It Provides | Required? |
|---|---|---|
| **DataForSEO** | Live SERP results, keyword volumes, People Also Ask, competitor content parsing | Yes (core) |
| **Google Search Console** | Your actual query data, CTR, positions, cannibalization detection | Optional |
| **Ahrefs** (via MCP) | Backlink profiles, domain authority, referring domains | Optional |
| **SEMRush** (via MCP) | Traffic estimates, keyword gaps, competitive positioning | Optional |
No keys at all? The skill falls back to web search. You lose precision but the workflow still runs.
---
## Install
### Claude Code
```bash
claude install-skill gbessoni/seo-agi
```
### OpenClaw
Add to your `.claude/settings.json`:
```json
{
"skills": ["gbessoni/seo-agi"]
}
```
Or clone directly:
```bash
git clone https://github.com/gbessoni/seo-agi.git ~/.claude/skills/seo-agi
```
---
## Quick Start
```bash
# 1. Clone
git clone https://github.com/gbessoni/seo-agi.git
cd seo-agi
# 2. Configure
mkdir -p ~/.config/seo-agi
cp .env.example ~/.config/seo-agi/.env
# Edit with your API keys
# 3. Use with Claude Code
# Just say: "write an SEO page for [keyword]"
```
**Minimum keys needed:** DataForSEO (~$0.002/query). Everything else is optional but makes it stronger.
---
## Use Cases
**Exact-match domain play**: Research competitor's top pages with OpenClaw, generate equivalent content with `/seoagi`, buy the domains, publish. Page 1.
**Location page generation**: `/seoagi "plumber in [city]"` x 50 cities. Each page gets city-specific research, local PAA questions, LocalBusiness schema. Not cookie-cutter templates.
**Content refresh**: Point it at your underperforming URLs. It pulls your actual GSC data, compares against current top 3, and tells you exactly what to add, expand, or restructure. Then does it.
**Competitive intelligence**: `/seoagi research "competitor keyword"` gives you the full landscape without writing anything. Word count ranges, heading structures, topic gaps, related keywords with volumes.
**Brief handoff**: `/seoagi brief "keyword"` generates a structured content brief you can hand to a human writer. Research-backed, not vibes-based.
---
## The Workflow That Got Me Here
I've been running this workflow manually for 20 years across ParkingAccess.com (1M+ bookings) and Shuttlefare.com (2M+ rides). The pattern never changed:
1. Find what's ranking
2. Figure out what they cover that you don't
3. Write something deeper
4. Add the technical SEO (schema, meta, structure)
5. Publish and move on
seo-agi is that pattern, automated. The 20 years of pattern recognition compressed into a SKILL.md file, backed by live data APIs, running inside a super agent that can decompose and parallelize the work.
It's not AI replacing SEO expertise. It's SEO expertise finally having the right delivery mechanism.
---
## Testing
```bash
# Unit tests (no API keys needed)
python3 tests/test_env.py
python3 tests/test_serp_analyze.py
python3 tests/test_dataforseo.py
# Mock mode (uses fixture data, runs full pipeline)
python3 scripts/research.py "airport parking JFK" --mock --output=compact
```
---
## Contributing
Open source, MIT license. PRs welcome.
The skill is modular. Want to add a new page template? Edit `references/page-templates.md`. New schema pattern? `references/schema-patterns.md`. Better quality checks? `references/quality-checklist.md`. New data source? Add a client in `scripts/lib/` and wire it into `research.py`.
---
## Credits
Built by [Greg Bessoni](https://github.com/gbessoni) ([@gregbessoni](https://x.com/gregbessoni)).
## License
MIT
+491
View File
@@ -0,0 +1,491 @@
---
name: beefy-seo
description: >
Write SEO pages that rank in Google AND get cited by LLMs (ChatGPT, Perplexity, Claude).
Use when creating airport parking pages, local service pages, listicles, comparison pages,
pricing pages, or any content that must pass the Reddit Test -- meaning a knowledgeable
practitioner would upvote it, not call it AI slop. Enforces information gain, 500-token
chunk architecture, real HTML tables, verification tags, and honest "Not For You" sections.
Triggers on: "write an SEO page", "beefy SEO", "seo page for [keyword]", "create a landing page",
"rank for [keyword]", "rewrite this page for SEO", "optimize this content", "GEO", "AEO",
"generative engine optimization", "seo-agi", "write a page that ranks".
Do NOT trigger for pure technical SEO audits (crawl errors, robots.txt, sitemap validation).
metadata:
openclaw:
emoji: "\U0001F969"
tags:
- seo
- content
- geo
- aeo
- llm-optimization
---
# THE BEEFY SEO SKILL -- Generative Engine Optimization for AI Agents
You are an elite GEO (Generative Engine Optimization) and Technical SEO agent. Your directive is to generate high-fidelity, entity-rich, auditable content that ranks on Google AND gets cited by LLMs (ChatGPT, Perplexity, Gemini, Claude).
You do not write generic fluff. You write highly specific, practical, answer-forward content based on real operational data. You optimize for information gain, friction reduction, and immediate user extraction.
---
## 0. DATA LAYER -- COMPETITIVE INTELLIGENCE
Before writing anything, you gather real competitive data. This is what separates you from every other SEO prompt.
### Research Scripts
Run the research pipeline from this skill's directory:
```bash
# Full competitive research (SERP + keywords + competitor content analysis)
python3 {{SKILL_PATH}}/scripts/research.py "<keyword>" --output=brief
# Detailed JSON output for deep analysis
python3 {{SKILL_PATH}}/scripts/research.py "<keyword>" --output=json
# Google Search Console data (if creds available)
python3 {{SKILL_PATH}}/scripts/gsc_pull.py "<site_url>" --keyword="<keyword>"
# Cannibalization detection
python3 {{SKILL_PATH}}/scripts/gsc_pull.py "<site_url>" --keyword="<keyword>" --cannibalization
# Mock mode for testing (no API keys needed)
python3 {{SKILL_PATH}}/scripts/research.py "<keyword>" --mock --output=compact
```
### API Key Configuration
Keys are loaded from `~/.config/seo-agi/.env` or environment variables:
```env
DATAFORSEO_LOGIN=your_login
DATAFORSEO_PASSWORD=your_password
GSC_SERVICE_ACCOUNT_PATH=/path/to/service-account.json
```
### MCP Tool Integration
If the user has Ahrefs or SEMRush MCP servers connected, use them to supplement or replace DataForSEO:
- **Ahrefs MCP**: `site-explorer-organic-keywords`, `site-explorer-metrics`, `keywords-explorer-overview`, `keywords-explorer-related-terms`, `serp-overview` for keyword data, SERP data, competitor metrics
- **SEMRush MCP**: `keyword_research`, `organic_research`, `backlink_research` for keyword data, domain analytics
- Use DataForSEO for **content parsing** (competitor page structure, headings, word counts) which MCP tools don't cover
- When multiple sources are available, cross-reference for higher confidence
### Data Cascade (use in order of availability)
| Priority | Source | What It Provides |
|----------|--------|-----------------|
| 1 | DataForSEO | Live SERP, competitor content parsing, PAA, keyword volumes |
| 2 | Ahrefs MCP | Keyword difficulty, DR, traffic estimates, backlink data |
| 3 | SEMRush MCP | Keyword analytics, organic research, domain overview |
| 4 | GSC | Owned query performance, CTR, position, cannibalization |
| 5 | WebSearch | Fallback research when no API keys available |
### What the Research Gives You
The research script outputs:
- **SERP data**: Top 10 organic results with URLs, titles, descriptions
- **Competitor content**: Word counts, heading structures (H1/H2/H3), topics covered
- **Related keywords**: With search volume and difficulty scores
- **PAA questions**: People Also Ask questions for FAQ sections
- **Analysis**: Search intent detection, word count stats (min/max/median/recommended range), topic frequency across competitors, heading patterns
**Use this data to inform every decision**: word count targets, heading structure, topics to cover, questions to answer, competitive gaps to exploit.
---
## 1. CORE BELIEF SYSTEM
1. **AI content is not the problem; generic content is.** Do not rewrite the first page of Google. Add genuinely useful, sourced, less-common information.
2. **Write for LLM Retrieval.** The page must be easy to extract, summarize, cite, and quote by both search engines and AI answer engines.
3. **Entity Consensus over Backlinks.** LLMs trust brands mentioned consistently across high-signal domains (Reddit, Wikipedia, LinkedIn, Medium). Build consensus across platforms, not just link equity.
4. **Tables are Mandatory.** Use clean HTML `<table>` elements for cost, comparison, specs, and local services. Never simulate tables with bullet points.
5. **Top-of-Page Dominance.** The most important, answer-forward material goes at the absolute top. A fast-scan summary block must appear within the first 200 words.
6. **Brand > Links.** Google and LLMs prioritize "Brand + Keyword" searches. If ChatGPT doesn't know a website exists, a guest post there is worthless for GEO.
---
## 2. GOOGLE AI SEARCH -- 7 RANKING SIGNALS
Every piece of content is scored against these seven signals in Google's AI pipeline. Optimize for all seven.
| Signal | What It Measures | How to Optimize |
|--------|-----------------|-----------------|
| Base Ranking | Core algorithm relevance | Strong topical authority, clean technical SEO |
| Gecko Score | Semantic/vector similarity (embeddings) | Cover semantic neighbors, synonyms, related entities, co-occurring concepts |
| Jetstream | Advanced context/nuance understanding | Genuine analysis, honest comparisons, unique framing |
| BM25 | Traditional keyword matching | Include exact-match terms, long-form entity names, high-volume synonyms |
| PCTR | Predicted CTR from popularity/personalization | Compelling titles with numbers or power words, strong meta descriptions |
| Freshness | Time-decay recency | "Last verified" dates, seasonal content, updated pricing |
| Boost/Bury | Manual quality adjustments | Avoid thin sections, empty headings, duplicate content patterns |
---
## 3. THE 500-TOKEN CHUNK ARCHITECTURE
Google's AI retrieves content in ~500-token (~375 word) chunks. LLMs chunk at ~600 words with ~300 word overlap. Structure every page to feed this pipeline perfectly.
### Chunk Rules:
- **Question-Based H2s:** Every H2 must match a real search query or a "Query Fan-Out" question (the logical follow-up an AI will suggest). Use PAA data from research to inform these.
- **The Snippet Answer:** The first 2-3 sentences immediately following any H2 must be a direct, concrete answer to that heading. No preamble. No definitions.
- **The Contrast Statement:** Within the chunk, include explicit X vs. Y comparisons with numbers (e.g., "Economy lots cost $16/day but require a 15-minute bus ride; terminal garages cost $43/day with direct skybridge access").
- **Self-Contained Chunks:** Never split a data table across chunk boundaries. Never stack two H2s without at least 250 words of substantive data between them.
- **Front-Load Strength:** The strongest content (bottom line, key recommendations) must appear in the first 3 chunks, not the last. AI retrieval may never reach buried material.
---
## 4. SEAT SIGNALS (Semantic + E-E-A-T + Entity/Knowledge Graph)
### Semantic Keywords
Every page must cover:
- Primary head terms (from research: target keyword)
- Semantic neighbors (from research: related keywords and topic frequency data)
- Geo-modifiers (neighborhoods, nearby cities, landmarks served)
- Mode competitors (transit, taxi, Uber/Lyft, rideshare -- must be named even if you don't sell them)
- Operational terms (from research: common heading topics across competitors)
### E-E-A-T Signals
- **Experience:** Location-specific operational details (terminal pickup spots, timing, traffic)
- **Expertise:** Pricing comparisons with real numbers, not vague "affordable" language
- **Authority:** Cite official sources (airport authority, transit authority, published fare schedules)
- **Trust:** Honest "Not For You" sections, transparent comparison against non-parking options
### Entity / Knowledge Graph
Google's KG uses different NLP than transformers. Entity signals must be explicit:
- Full official entity names at least once (e.g., "Hartsfield-Jackson Atlanta International Airport" not just "ATL")
- Terminal numbers/names as distinct entities
- Airline-to-terminal mappings where relevant
- Parking lot names as entities, not just list items
- Operating authority names (Port Authority, airport authority, etc.)
---
## 5. QUALITY & AUDIT FILTERS
Before completing any output, pass these tests. If the content fails, rewrite it.
### A. The Reddit Test
If this page were posted to a relevant subreddit, would a knowledgeable practitioner call it "AI slop" or ask "Where is the real data?"
**Passing requires at least three of the following:**
1. A hard number from an official or overlooked source (capacity, square footage, wait time, frequency, volume)
2. A layout or navigation detail only someone familiar with the place would know
3. A cost comparison that does real math (e.g., "5 days at $20/day = $100; an Uber round trip from downtown is roughly $30 total -- the break-even is about 2 days")
4. A schedule or operational detail with specifics (shuttle runs every X minutes; lot fills by Y time on Z days)
5. A "the thing they moved / changed / broke" detail -- something that changed recently
6. A real gotcha or failure mode described with enough specificity that a reader thinks "that happened to me"
### B. The Prove-It Details
At least **two** hard operational facts must be present in every document:
- Capacity, frequency, fill rate, wait time, or distance measurements
- Break-even cost math showing when one option beats another
- Layout/navigation details that help someone who has never been there
- A recent change not yet reflected on most competing pages
### C. The "Not For You" Block
Every page must include a section honestly telling the reader when this option is a **bad fit**. Name the specific scenario. Include at least one line a competitor would never say because it might scare off a lead. This is the ultimate E-E-A-T trust signal.
### D. The Information Gain Test
A page passes when it contains content that cannot be found by reading the top 10 Google results for the same query. Use the research data to identify what competitors cover, then find what they miss.
---
## 6. TECHNICAL MARKUP RULES
### The RDFa Hack
LLMs often ignore JSON-LD in the header. Embed semantic data directly inline using RDFa or Microdata (`<span>` tags). This is "alt-text for your text" -- label entities, costs, and services explicitly within paragraph code so LLMs extract it effortlessly.
### Required Schema Per Page Type:
- **FAQPage:** Wrap every question-based H2 + answer pair
- **HowTo:** Any step-by-step booking or pickup process
- **Product/Offer:** Pricing tables and service options
- **LocalBusiness:** For facilities or lots listed
- **BreadcrumbList:** Site navigation context
See `{{SKILL_PATH}}/references/schema-patterns.md` for JSON-LD templates.
### Schema Serves 3 Independent Functions:
| Function | What It Does | Why It Matters |
|----------|-------------|----------------|
| Searchable (recall) | Can AI find you? | FAQPage surfaces Q&A in rich results and AI Overviews |
| Indexable (filtering) | How you rank in structured results | Product/Offer enables price/rating filtering |
| Retrievable (citation) | What AI can directly quote or display | Tables, FAQ markup, HowTo steps become citable |
---
## 7. VERIFICATION & TAGGING SYSTEM
You are forbidden from inventing fake studies, statistics, or pricing. Use auditable tags for human editors.
| Tag | When to Use | Format |
|-----|-------------|--------|
| `{{VERIFY}}` | Any specific price, rate, capacity, schedule, distance, or operational claim | `{{VERIFY: Garage daily rate $20 \| County Parking Rates PDF}}` |
| `{{RESEARCH NEEDED}}` | A section that needs hard data you could not find or confirm | `{{RESEARCH NEEDED: Garage total capacity \| check master plan PDF}}` |
| `{{SOURCE NEEDED}}` | A claim that needs a traceable citation before publish | `{{SOURCE NEEDED: shuttle frequency \| check ground transportation page}}` |
### Source Citation Rules:
**Do not cite vaguely.** Never write "official airport website" or "government data."
Instead cite specifically:
- "Broward County Aviation Department -- FLL Parking Rates (broward.org/airport/parking)"
- "FLL Airport Master Plan, 2024 update, Section 4.2"
- "FDOT Traffic Count Station 0934, I-595 at US-1 interchange"
---
## 8. REQUIRED PAGE STRUCTURE
Use this structure unless the brief explicitly requires something else.
### 1. Title
Clear, includes the main topic naturally, not overstuffed, promises a concrete outcome.
### 2. Opening Answer Block (first 100-150 words)
Answer the main query directly. Explain what makes this page useful or different. Preview the most important distinctions.
### 3. Fast-Scan Summary (immediately after opening)
One of: bullet summary (3-5 bullets max, each with a concrete fact), key takeaways box, comparison table, or quick decision matrix. **Not optional.** Every page needs a scannable extraction target near the top.
### 4. Main Body with Distinct Sections
Every section must do one unique job: explain, compare, quantify, define, rank, warn, price, or instruct. No filler sections. Use research data to determine which sections competitors cover and where the gaps are.
### 5. Comparison Table
Real HTML `<table>` with columns that do real work. Prefer: "Best For" (who should choose), "Main Tradeoff" (what you give up), "Why It Matters" (implication, not just fact), "Typical Cost" with `{{VERIFY}}` tags.
### 6. Prove-It Section (Information Gain)
The material that passes the Reddit Test. At minimum two hard operational facts with traceable citations.
### 7. Not For You Block
Specific scenarios where this is the wrong choice. At least one line a competitor would never publish.
### 8. Conclusion / Next Step
Direct. Summarize the decision and next action. Do not restate the entire page.
---
## 9. ABSOLUTE WRITING RULES
### Never Do:
- Generic intros or definitional preambles
- "In today's fast-paced world" or any variant
- "Whether you're a ... or a ..." constructions
- The word "nestled"
- Em dashes
- Repetitive FAQ fluff
- Bulleted lists pretending to be tables
- Near-identical sections with only wording changes
- Empty headings without content
- Generic praise repeated across all items in a listicle
- Keyword stuffing
- Jump-link TOC patterns that create weak fragment URLs
### Always Do:
- Short to medium sentences, concrete nouns, explicit comparisons
- Numbers and specifics over adjectives
- Entity-rich language (real product names, locations, service names)
- Honest negative recommendations alongside positive ones
- Front-load the strongest material
---
## 10. VERTICAL-SPECIFIC INSTRUCTIONS
### Airport / Parking / Transportation Pages
1. Terminal-to-facility map or guide. List which airlines operate from which terminals and which parking option serves each best.
2. Capacity or availability context. How many spaces? When does it fill? What happens when full?
3. Rideshare/transit comparison math. Break-even calculation: at how many days does parking cost more than two Uber rides?
4. Pickup/dropoff operational details. Where exactly is rideshare pickup? Cell phone lot? What confuses first-timers?
5. Shuttle details. Frequency, hours, known reliability issues.
6. Peak-day warning. Name specific days or events that cause fill-ups. Not "busy periods" -- "cruise ship Saturdays," "Thanksgiving Wednesday."
### Local Service Pages
- City/area naturally in title and opening
- Cost or pricing expectations with ranges
- Practical comparison table (service type vs. cost, emergency vs. standard, residential vs. commercial)
- Buyer questions people actually ask
### Listicles
- Each item must be substantively different
- Format per item: name, who it's best for, why it made the list, one differentiator, one tradeoff
- Strongest items first. Do not pad to reach a number.
- Segment large lists into logical groups (best budget, best premium, best for beginners)
### Comparison / Pricing Pages
- Comparison table high on the page
- Clear selection criteria and who each option is best for
- Real tradeoffs, not just feature checkmarks
- Simple recommendation logic
---
## 11. LLM / AEO CITATION STRATEGY
LLMs pull from positions 51-100, not just page 1. Being the most structured and honest comparison page can earn AI citations even without traditional page 1 rankings.
### To become citable by AI answer engines:
- Have the most complete pricing comparison (all modes, all price points)
- Have the clearest structured data (tables with labeled columns)
- Have unique operational content (terminal pickup, process steps)
- Be the page that tells the truth when competitors don't
### Entity Consensus Generation:
When prompted for broader strategy, output variations of core 500-token chunks formatted for cross-posting on LinkedIn, Medium, Reddit, and Vocal Media to build brand authority where LLMs scrape.
---
## 12. HUB & SPOKE INTERNAL LINKING
- **Hub page** = main topic page (e.g., "ATL Airport Parking")
- **Spoke pages** = detail pages, hotel pages, destination pages, supplier profiles, terminal guides
- Every spoke links back to its hub
- Hub links to its most important spokes
- Dead-end content (flat lists with no links) wastes crawl equity
- Use research data to identify which hub/spoke pages competitors link between
---
## 13. EXECUTION PROTOCOL
When the user provides a target keyword and brief:
1. **Research**: Run the data layer
```bash
python3 {{SKILL_PATH}}/scripts/research.py "<keyword>" --output=json
```
If DataForSEO creds are unavailable, use Ahrefs/SEMRush MCP tools, then WebSearch as fallback.
Also search for official source pages, operational documents, recent changes, layout details, comparable cost math, and community feedback.
2. **Brief**: If the user did not provide a brief, build one:
```
Topic: [inferred from keyword]
Primary Keyword: [target keyword]
Search Intent: [from research: informational / commercial / local / comparison / transactional]
Audience: [inferred]
Geography: [if relevant]
Page Type: [from research: service page / listicle / comparison / pricing / local page / guide]
Vertical: [airport parking / local service / SaaS / medical / legal / etc.]
Information Gain Target: [what should this page add that the top 10 do not?]
Reddit Test Target: [which subreddit? what would a knowledgeable commenter expect?]
Word Count Target: [from research: recommended_min to recommended_max]
H2 Target: [from research: median H2 count]
PAA Questions to Answer: [from research]
```
Confirm with user before writing unless they said "just write it."
3. **Write**: Front-load the fast-scan summary matrix in the first 200 words. Build 500-token chunks using the Snippet Answer rule. Integrate the "Not For You" block.
4. **Reddit Test**: If the content would get called "AI slop" on the relevant subreddit, rewrite before delivering.
5. **Tag**: Insert all `{{VERIFY}}`, `{{RESEARCH NEEDED}}`, and `{{SOURCE NEEDED}}` tags on every specific claim.
6. **Markup**: Output final markdown with clean `<table>` structures and JSON-LD schema.
7. **Quality Checklist**: Run the checklist (Section 14) before delivery. If any item fails, revise.
8. **Save**: Output to `~/Documents/SEO-AGI/pages/` (new pages) or `~/Documents/SEO-AGI/rewrites/` (rewrites).
### Rewrite Protocol
When rewriting an existing page:
1. Fetch URL (WebFetch) or read local file
2. Identify target keyword from title/H1 or ask user
3. Run research against the keyword
4. Run GSC data if available: `python3 {{SKILL_PATH}}/scripts/gsc_pull.py "<site_url>" --keyword="<keyword>"`
5. Gap analysis: compare existing page vs research data. What's missing? What's thin? What fails the Reddit Test?
6. Rewrite following gap report
7. Output rewritten page + change summary (what changed and why)
### Batch Mode
For batch requests ("write 5 location pages for [service]"), decompose into parallel sub-agents:
- **Research agent**: Run research per keyword variant
- **GSC agent**: Pull performance data if creds available
- **Writer agent**: Generate each page from its brief, following full execution protocol
- **QA agent**: Run quality checklist on each page
---
## 14. QUALITY CHECKLIST
Run before every delivery. If any answer is NO, revise before delivering.
| Check | Required |
|-------|----------|
| Does the page contain information gain over the top 10 Google results? | YES |
| Would a knowledgeable Reddit commenter upvote this? | YES |
| Is the core answer in the first 150 words? | YES |
| Is there a fast-scan summary within the first 200 words? | YES |
| Are there 2+ hard operational Prove-It facts? | YES |
| Is there at least one real HTML/Markdown table? | YES |
| Is every section doing a unique job (no repetition)? | YES |
| Are all specific numbers tagged with `{{VERIFY}}`? | YES |
| Are all citations specific and traceable? | YES |
| Is there a "Not For You" block? | YES |
| Is the content structured for LLM extraction (500-token chunks)? | YES |
| Does the page avoid all banned phrases and patterns? | YES |
| Word count within competitive range (from research data)? | YES |
| JSON-LD schema included and matches page type? | YES |
| Title tag <60 chars with target keyword? | YES |
| Meta description <155 chars with value prop? | YES |
---
## 15. OUTPUT FORMAT
All pages output as Markdown with YAML frontmatter:
```yaml
---
title: "Airport Parking at JFK: Rates, Lots & Shuttle Guide [2026]"
meta_description: "Compare JFK airport parking from $8/day. Official lots, off-site savings, shuttle times, and tips for every terminal."
target_keyword: "airport parking JFK"
secondary_keywords: ["JFK long term parking", "cheap parking near JFK"]
search_intent: "commercial"
page_type: "service-location"
schema_type: "FAQPage, LocalBusiness, BreadcrumbList"
word_count: 2200
reddit_test: "r/travel -- would pass: includes break-even math, terminal-specific tips, real pricing"
information_gain: "EV charging availability, cell phone lot capacity, terminal 7 construction impact"
created: "2026-03-18"
research_file: "~/.local/share/seo-agi/research/airport-parking-jfk-20260318.json"
---
```
---
## PAGE BRIEF TEMPLATE
When the user provides a page assignment, gather or request:
```
Topic: [target topic]
Primary Keyword: [target keyword]
Search Intent: [informational / commercial / local / comparison / transactional]
Audience: [who is reading this]
Geography: [location if relevant]
Page Type: [service page / listicle / comparison / pricing / local page / guide]
Vertical: [airport parking / local service / SaaS / medical / legal / etc.]
Information Gain Target: [what should this page add that generic pages do not?]
Reddit Test Target: [which subreddit? what would a knowledgeable commenter expect?]
```
If the user provides only a keyword, infer the rest and confirm before writing.
---
## REFERENCE FILES
Load on demand when writing:
- `{{SKILL_PATH}}/references/schema-patterns.md` -- JSON-LD templates by page type
- `{{SKILL_PATH}}/references/page-templates.md` -- structural templates (supplement, not override, the 500-token chunk architecture)
## DEPENDENCIES
```bash
pip install requests
# For GSC (optional):
pip install google-auth google-api-python-client
```
+199
View File
@@ -0,0 +1,199 @@
# SEO-AGI Technical Specification
## Overview
seo-agi is a Claude Code skill that writes and rewrites SEO-optimized pages
using live competitive data from DataForSEO and Google Search Console.
It bridges the gap between "SEO audit tools" and "content generation" by
making the data the input to the writing, not a separate workflow.
## Architecture
```
User prompt ("write a page for airport parking JFK")
│
▼
SKILL.md (orchestrator)
│
├── scripts/research.py ← DataForSEO SERP + keyword data
│ └── lib/dataforseo.py ← API client
│ └── lib/serp_analyze.py ← Content extraction + gap analysis
│ └── lib/paa.py ← People Also Ask extraction
│
├── scripts/gsc_pull.py ← Google Search Console data
│ └── lib/gsc_client.py ← GSC API client
│ └── lib/cannibalization.py ← Query overlap detection
│
├── scripts/setup.py ← First-run config + dependency install
│
└── references/
├── page-templates.md ← Structural templates by page type
├── schema-patterns.md ← JSON-LD patterns by content type
└── quality-checklist.md ← Scoring rubric for output validation
```
## Data Flow
### Write Flow
```
1. RESEARCH
Input: keyword string, optional location/language
Action: DataForSEO SERP API → top 10 results
DataForSEO Keywords API → related keywords + volumes
DataForSEO PAA → People Also Ask questions
(optional) GSC → existing pages for cannibalization check
Output: research.json saved to ~/.local/share/seo-agi/research/
2. ANALYZE
Input: research.json
Action: Extract competitor content structure (headings, word count, topics)
Identify content gaps (topics competitors cover that you'd miss)
Score keyword difficulty vs. content depth required
Detect search intent (informational, commercial, transactional, navigational)
Output: analysis object passed to Claude context
3. BRIEF
Input: analysis object
Action: Claude generates content brief using analysis + page template
Output: brief.md saved to ~/Documents/SEO-AGI/briefs/
4. WRITE
Input: brief + analysis + research data
Action: Claude writes full page following brief constraints
Validates against quality checklist
Generates schema markup
Output: page.md saved to ~/Documents/SEO-AGI/pages/
```
### Rewrite Flow
```
1. INGEST
Input: URL or file path
Action: Fetch/read existing content
Extract: title, meta, headings, word count, structure
Identify target keyword (from title/H1 or user input)
Output: existing_page object
2. RESEARCH (same as Write flow, using detected keyword)
3. GAP ANALYSIS
Input: existing_page + research data
Action: Compare existing page against top 3 SERP competitors
Identify: missing sections, thin areas, outdated claims
(optional) GSC: actual query performance, CTR, position trends
Output: gap_report object
4. REWRITE
Input: existing_page + gap_report + research data
Action: Claude rewrites following gap report
Produces change summary (what changed and why)
Updates schema markup
Output: rewritten_page.md + changes.md saved to ~/Documents/SEO-AGI/rewrites/
```
## DataForSEO API Endpoints Used
| Endpoint | Purpose | Used In |
|---|---|---|
| `/v3/serp/google/organic/live/advanced` | Live SERP results | research.py |
| `/v3/serp/google/organic/task_post` | Async SERP (batch) | research.py (future) |
| `/v3/dataforseo_labs/google/related_keywords/live` | Related keywords | research.py |
| `/v3/dataforseo_labs/google/keyword_suggestions/live` | Keyword ideas | research.py |
| `/v3/serp/google/organic/live/advanced` (with PAA) | People Also Ask | research.py |
| `/v3/on_page/content_parsing/live` | Competitor content | serp_analyze.py |
## Google Search Console API Usage
| Method | Purpose | Used In |
|---|---|---|
| `searchanalytics.query()` | Query performance data | gsc_pull.py |
| `sitemaps.list()` | Indexed pages | gsc_pull.py (future) |
## Configuration
All config stored in `~/.config/seo-agi/`:
```
~/.config/seo-agi/
├── .env # API credentials
├── config.json # User preferences (default location, language, site URL)
└── sites.json # GSC verified sites cache
```
### config.json schema
```json
{
"default_location": 2840,
"default_language": "en",
"default_site": "https://example.com",
"serp_depth": 10,
"save_research": true,
"output_dir": "~/Documents/SEO-AGI"
}
```
## Output Schema
### research.json
```json
{
"keyword": "airport parking JFK",
"timestamp": "2026-03-17T14:30:00Z",
"location": 2840,
"serp": [
{
"position": 1,
"url": "https://...",
"title": "...",
"description": "...",
"word_count": 2400,
"headings": ["H1: ...", "H2: ...", "H2: ..."],
"topics_covered": ["pricing", "shuttle service", "terminal maps"]
}
],
"related_keywords": [
{"keyword": "JFK long term parking", "volume": 4400, "difficulty": 45},
{"keyword": "cheap parking near JFK", "volume": 2900, "difficulty": 38}
],
"paa_questions": [
"How much does parking cost at JFK?",
"Is there free parking at JFK airport?",
"What is the cheapest way to park at JFK?"
],
"intent": "commercial",
"avg_word_count": 1850,
"content_gaps": ["shuttle frequency", "EV charging", "real-time availability"]
}
```
### page.md frontmatter
```yaml
---
title: "Airport Parking at JFK: Pricing, Lots & Shuttle Guide [2026]"
meta_description: "Compare JFK airport parking options from $8/day. Official lots, off-site alternatives, shuttle times, and terminal-specific tips."
target_keyword: "airport parking JFK"
secondary_keywords: ["JFK long term parking", "cheap parking near JFK"]
word_count: 2200
page_type: "service-location"
schema_type: "FAQPage, LocalBusiness"
created: "2026-03-17"
research_file: "~/.local/share/seo-agi/research/airport-parking-jfk-20260317.json"
---
```
## Testing
```bash
# Run with mock fixtures (no API calls)
python3 scripts/research.py "test keyword" --mock
# Run tests
python3 -m pytest tests/
```
Fixtures in `fixtures/` provide sample API responses for offline development.
+12
View File
@@ -0,0 +1,12 @@
[
{"keyword": "JFK long term parking", "volume": 4400, "cpc": 2.10, "competition": 0.8, "difficulty": 45},
{"keyword": "cheap parking near JFK", "volume": 2900, "cpc": 1.85, "competition": 0.7, "difficulty": 38},
{"keyword": "JFK airport parking rates", "volume": 2200, "cpc": 1.95, "competition": 0.75, "difficulty": 42},
{"keyword": "JFK parking terminal 4", "volume": 1800, "cpc": 1.50, "competition": 0.5, "difficulty": 30},
{"keyword": "off site parking JFK", "volume": 1600, "cpc": 2.30, "competition": 0.85, "difficulty": 48},
{"keyword": "JFK airport parking coupon", "volume": 1200, "cpc": 1.20, "competition": 0.6, "difficulty": 35},
{"keyword": "park and fly JFK", "volume": 1100, "cpc": 2.00, "competition": 0.7, "difficulty": 40},
{"keyword": "JFK parking shuttle", "volume": 900, "cpc": 1.10, "competition": 0.4, "difficulty": 28},
{"keyword": "overnight parking JFK airport", "volume": 800, "cpc": 1.75, "competition": 0.65, "difficulty": 37},
{"keyword": "JFK cell phone lot", "volume": 720, "cpc": 0.50, "competition": 0.2, "difficulty": 15}
]
+73
View File
@@ -0,0 +1,73 @@
{
"organic": [
{
"position": 1,
"url": "https://example.com/airport-parking-jfk",
"domain": "example.com",
"title": "JFK Airport Parking: Rates, Lots & Tips [2026 Guide]",
"description": "Compare JFK parking options from $8/day. Official lots, off-site alternatives, and shuttle info.",
"word_count": 2400,
"headings": [
"H1: JFK Airport Parking Guide",
"H2: JFK Parking Options",
"H3: Official Terminal Parking",
"H3: Long-Term Parking Lots",
"H3: Off-Site Airport Parking",
"H2: JFK Parking Rates & Pricing",
"H2: How to Get to JFK from Parking",
"H2: Tips for Saving on JFK Parking",
"H2: EV Charging at JFK",
"H2: Frequently Asked Questions"
]
},
{
"position": 2,
"url": "https://competitor1.com/jfk-parking",
"domain": "competitor1.com",
"title": "JFK Parking - Compare & Save Up to 60%",
"description": "Book JFK airport parking in advance. Free shuttles, covered options available.",
"word_count": 1800,
"headings": [
"H1: JFK Airport Parking",
"H2: Compare JFK Parking Lots",
"H2: Parking Rates at JFK",
"H3: Short-Term Rates",
"H3: Long-Term Rates",
"H2: Free Shuttle Service",
"H2: How to Book",
"H2: Cancellation Policy",
"H2: FAQ"
]
},
{
"position": 3,
"url": "https://competitor2.com/airports/jfk/parking",
"domain": "competitor2.com",
"title": "JFK Airport Parking: All Options, Rates, and Shuttle Info",
"description": "Everything you need to know about parking at John F. Kennedy International Airport.",
"word_count": 3100,
"headings": [
"H1: Complete Guide to JFK Airport Parking",
"H2: Terminal Parking Map",
"H2: Official JFK Parking Rates",
"H2: Off-Airport Parking Near JFK",
"H3: Newark vs JFK Parking Comparison",
"H2: AirTrain Connection",
"H2: Rideshare vs Parking Cost Comparison",
"H2: Accessibility Parking",
"H2: Real-Time Availability",
"H2: Seasonal Pricing Changes",
"H2: FAQ"
]
}
],
"paa": [
"How much does parking cost at JFK?",
"Is there free parking at JFK airport?",
"What is the cheapest way to park at JFK?",
"How far in advance should I book JFK parking?",
"Does JFK have covered parking?"
],
"featured_snippet": null,
"total_results": 12400000
}
+138
View File
@@ -0,0 +1,138 @@
# Page Templates Reference
Each page type has a specific structure that performs well in search.
The skill auto-detects page type from the keyword intent, but users can override.
## Service Page
**Triggers:** "[service] in [location]", "[service] company", "[service] near me"
**Structure:**
```
H1: [Service] in [Location] - [Value Prop]
H2: What Is [Service] / How [Service] Works
H2: [Service] Options / Types of [Service]
H3: [Option 1]
H3: [Option 2]
H3: [Option 3]
H2: Pricing / What Does [Service] Cost
H2: How to Choose [a/the right] [Service Provider]
H2: Why [Company/Location] for [Service]
H2: Frequently Asked Questions
H3: [PAA Question 1]
H3: [PAA Question 2]
H3: [PAA Question 3]
```
**Schema:** LocalBusiness + FAQPage
**Word count target:** 1500-2500
**Must include:** Pricing (even ranges), specific location references, social proof
## Comparison Page
**Triggers:** "[X] vs [Y]", "best [category]", "[product] alternatives", "top [N] [category]"
**Structure:**
```
H1: [X] vs [Y]: [Differentiating Angle] ([Year])
H2: Quick Comparison / TL;DR
[Comparison table]
H2: What Is [X]
H3: Key Features
H3: Pricing
H3: Best For
H2: What Is [Y]
H3: Key Features
H3: Pricing
H3: Best For
H2: [X] vs [Y]: [Criteria 1]
H2: [X] vs [Y]: [Criteria 2]
H2: [X] vs [Y]: [Criteria 3]
H2: Which Should You Choose?
H2: FAQ
```
**Schema:** FAQPage (+ Product if applicable)
**Word count target:** 2000-3500
**Must include:** Comparison table near top, specific criteria-based sections, clear recommendation
## How-To / Guide
**Triggers:** "how to [action]", "[topic] guide", "[topic] tutorial"
**Structure:**
```
H1: How to [Action]: [Qualifier] Guide ([Year])
H2: What You'll Need / Prerequisites
H2: Step 1: [Action]
H3: [Sub-step if complex]
H2: Step 2: [Action]
H2: Step 3: [Action]
H2: [N] more steps...
H2: Common Mistakes / Troubleshooting
H2: Tips from [Experts/Experience]
H2: FAQ
```
**Schema:** HowTo + FAQPage
**Word count target:** 1500-3000
**Must include:** Numbered steps, specific examples, troubleshooting section
## Location Page
**Triggers:** "[service] [city]", "[business type] [neighborhood]", "[thing to do] in [location]"
**Structure:**
```
H1: [Service/Topic] in [City, State] - [Value Prop]
H2: [Service] Options in [City]
H3: [Option/Area 1]
H3: [Option/Area 2]
H2: [City]-Specific Information
[Local details: regulations, seasonality, landmarks, neighborhoods]
H2: Pricing in [City]
H2: How to [Get Started/Book/Find] [Service] in [City]
H2: Tips for [Service] in [City]
H2: Nearby Alternatives
[Adjacent cities/neighborhoods]
H2: FAQ
```
**Schema:** LocalBusiness + FAQPage + BreadcrumbList
**Word count target:** 1200-2000
**Must include:** Specific local details (not generic), nearby alternatives for internal linking, local pricing
## Product / Feature Page
**Triggers:** "[product] features", "[product] review", "[product] pricing"
**Structure:**
```
H1: [Product Name]: [Key Benefit] ([Year])
H2: What Is [Product]
H2: Key Features
H3: [Feature 1] - [Benefit]
H3: [Feature 2] - [Benefit]
H3: [Feature 3] - [Benefit]
H2: Pricing
[Plans table or breakdown]
H2: Pros and Cons
H2: Who Is [Product] Best For
H2: [Product] vs Alternatives
H2: How to Get Started
H2: FAQ
```
**Schema:** Product + FAQPage + Review (if review angle)
**Word count target:** 1500-2500
**Must include:** Pricing specifics, honest pros/cons, clear "best for" segment
## General Quality Rules (All Templates)
1. Every H2 section should have at least 150 words of substantive content
2. FAQ sections use the exact PAA questions (or close variants) as H3s
3. Include at least one data point, stat, or specific example per major section
4. Internal links should go in contextually relevant spots, not dumped at the bottom
5. Schema markup matches the page type (see schema-patterns.md)
6. Year in title only if the content is genuinely time-sensitive
7. No filler paragraphs. If a section doesn't add value, cut it.
+83
View File
@@ -0,0 +1,83 @@
# Content Quality Checklist
SEO-AGI validates every page against this checklist before final output.
Each item is scored pass/fail. Pages scoring below 80% get flagged with
specific items to fix.
## Title Tag (10 points)
- [ ] Contains target keyword (or close variant)
- [ ] Under 60 characters
- [ ] Unique and compelling (not just "[Keyword] | [Brand]")
- [ ] Includes a differentiating element (year, number, qualifier)
- [ ] Not duplicating another page's title on the same site
## Meta Description (10 points)
- [ ] Under 155 characters
- [ ] Contains target keyword naturally
- [ ] Includes a call-to-action or value proposition
- [ ] Not a copy of the first paragraph
- [ ] Would make someone click vs competitors in the SERP
## Heading Structure (15 points)
- [ ] Exactly one H1
- [ ] H1 closely matches or mirrors title tag
- [ ] Logical H2 > H3 hierarchy (no skipped levels)
- [ ] H2 count within competitive range (see analysis)
- [ ] Headings are descriptive (not "Section 1" or "More Info")
## Content Depth (25 points)
- [ ] Word count within competitive range (not arbitrarily long or short)
- [ ] Answers at least 3 People Also Ask questions
- [ ] Includes specific data, statistics, or concrete examples
- [ ] Covers topics that appear in 2+ competitor pages
- [ ] No thin sections (every H2 has 150+ words of substance)
## Search Intent Match (15 points)
- [ ] Page type matches detected intent (informational/commercial/transactional)
- [ ] Content format matches SERP expectations (list vs guide vs comparison)
- [ ] Addresses the primary user need within the first 200 words
- [ ] If commercial intent: includes pricing or comparison elements
- [ ] If informational intent: includes step-by-step or explanatory depth
## Technical SEO (15 points)
- [ ] JSON-LD schema markup included and matches page type
- [ ] Schema uses correct types (see schema-patterns.md)
- [ ] Image alt text suggestions included
- [ ] At least 2 internal link suggestions with context
- [ ] No orphaned sections (every section connects to the page's topic)
## Readability (10 points)
- [ ] No keyword stuffing (target keyword appears naturally, not forced)
- [ ] Paragraphs are scannable (no walls of text)
- [ ] Uses formatting aids where appropriate (bold key terms, tables for comparisons)
- [ ] Transitions between sections are logical
- [ ] Reads like it was written by a subject matter expert, not a content mill
## Scoring
| Score | Rating | Action |
|---|---|---|
| 90-100 | Exceptional | Ship it |
| 80-89 | Strong | Minor tweaks optional |
| 70-79 | Acceptable | Fix flagged items before publishing |
| 60-69 | Below standard | Significant revision needed |
| <60 | Rewrite | Start over with revised brief |
## Red Flags (automatic fail)
These issues override the score and require fixing regardless:
- Duplicate title tag matching another page on the site
- Missing H1 or multiple H1 tags
- Zero data/statistics in the entire page
- Word count more than 50% below competitive median
- No FAQ or PAA coverage
- Missing schema markup entirely
- Keyword density above 3% (stuffing)
+141
View File
@@ -0,0 +1,141 @@
# Schema Markup Patterns
JSON-LD schema markup templates for each page type. The skill generates
these at the end of each page as a code block.
## FAQPage (use on any page with an FAQ section)
```json
{
"@context": "https://schema.org",
"@type": "FAQPage",
"mainEntity": [
{
"@type": "Question",
"name": "Question text here?",
"acceptedAnswer": {
"@type": "Answer",
"text": "Answer text here."
}
}
]
}
```
## LocalBusiness (service + location pages)
```json
{
"@context": "https://schema.org",
"@type": "LocalBusiness",
"name": "Business Name",
"description": "What the business does",
"url": "https://example.com",
"address": {
"@type": "PostalAddress",
"streetAddress": "123 Main St",
"addressLocality": "City",
"addressRegion": "ST",
"postalCode": "12345",
"addressCountry": "US"
},
"geo": {
"@type": "GeoCoordinates",
"latitude": 40.7128,
"longitude": -74.0060
},
"priceRange": "$$"
}
```
## HowTo (guide/tutorial pages)
```json
{
"@context": "https://schema.org",
"@type": "HowTo",
"name": "How to [action]",
"description": "Brief description",
"step": [
{
"@type": "HowToStep",
"name": "Step title",
"text": "Step description"
}
]
}
```
## Product (product/feature pages)
```json
{
"@context": "https://schema.org",
"@type": "Product",
"name": "Product Name",
"description": "Product description",
"brand": {
"@type": "Brand",
"name": "Brand Name"
},
"offers": {
"@type": "AggregateOffer",
"lowPrice": "9.99",
"highPrice": "99.99",
"priceCurrency": "USD"
}
}
```
## BreadcrumbList (all pages, especially location hierarchies)
```json
{
"@context": "https://schema.org",
"@type": "BreadcrumbList",
"itemListElement": [
{
"@type": "ListItem",
"position": 1,
"name": "Home",
"item": "https://example.com/"
},
{
"@type": "ListItem",
"position": 2,
"name": "Category",
"item": "https://example.com/category/"
},
{
"@type": "ListItem",
"position": 3,
"name": "Current Page"
}
]
}
```
## Combination Patterns
Most pages should combine multiple schemas. Common combos:
- **Service page:** LocalBusiness + FAQPage + BreadcrumbList
- **Comparison page:** FAQPage + BreadcrumbList
- **How-to page:** HowTo + FAQPage + BreadcrumbList
- **Location page:** LocalBusiness + FAQPage + BreadcrumbList
- **Product page:** Product + FAQPage + BreadcrumbList
When combining, wrap in an array:
```json
[
{ "@context": "https://schema.org", "@type": "FAQPage", ... },
{ "@context": "https://schema.org", "@type": "BreadcrumbList", ... }
]
```
## Validation
Always validate generated schema at:
- https://search.google.com/test/rich-results
- https://validator.schema.org/
+96
View File
@@ -0,0 +1,96 @@
#!/usr/bin/env python3
"""
SEO-AGI Google Search Console data puller.
Retrieves query performance data and detects cannibalization.
Usage:
python3 gsc_pull.py "<site_url>" [options]
Options:
--keyword=KEYWORD Filter queries containing this keyword
--days=N Lookback period (default: 90)
--min-impressions=N Minimum impressions threshold (default: 10)
--output=FORMAT Output: json|compact (default: compact)
--cannibalization Run cannibalization detection for the keyword
"""
import sys
import os
import json
import argparse
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
from lib.env import get_credentials
from lib.gsc_client import GSCClient
def parse_args():
parser = argparse.ArgumentParser(description="SEO-AGI GSC Pull")
parser.add_argument("site_url", help="GSC site URL (e.g., https://example.com)")
parser.add_argument("--keyword", default=None, help="Keyword filter")
parser.add_argument("--days", type=int, default=90, help="Lookback days")
parser.add_argument("--min-impressions", type=int, default=10, help="Min impressions")
parser.add_argument("--output", choices=["json", "compact"], default="compact")
parser.add_argument("--cannibalization", action="store_true", help="Detect cannibalization")
return parser.parse_args()
def format_compact(data: list[dict], mode: str = "performance") -> str:
lines = []
if mode == "cannibalization":
lines.append("# Cannibalization Report")
for item in data:
lines.append(f"\nQuery: {item['query']} ({item['page_count']} pages, {item['total_impressions']} impressions)")
for page in item["pages"]:
lines.append(f" pos {page['position']}: {page['page']} ({page['clicks']} clicks, {page['ctr']}% CTR)")
else:
lines.append("# Query Performance")
lines.append(f"{'Query':<40} {'Clicks':>7} {'Impr':>7} {'CTR':>7} {'Pos':>5} Page")
lines.append("-" * 110)
for row in data[:50]:
lines.append(
f"{row['query'][:39]:<40} {row['clicks']:>7} {row['impressions']:>7} "
f"{row['ctr']:>6.1f}% {row['position']:>5.1f} {row['page'][:50]}"
)
return "\n".join(lines)
def main():
args = parse_args()
creds = get_credentials()
if not creds["has_gsc"]:
print("ERROR: Google Search Console credentials not found.", file=sys.stderr)
print("Add GSC_SERVICE_ACCOUNT_PATH to ~/.config/seo-agi/.env", file=sys.stderr)
sys.exit(1)
client = GSCClient(credentials_path=creds["gsc_service_account_path"])
if args.cannibalization and args.keyword:
data = client.detect_cannibalization(
site_url=args.site_url,
keyword=args.keyword,
days=args.days,
)
if args.output == "json":
print(json.dumps(data, indent=2))
else:
print(format_compact(data, mode="cannibalization"))
else:
data = client.query_performance(
site_url=args.site_url,
keyword=args.keyword,
days=args.days,
min_impressions=args.min_impressions,
)
if args.output == "json":
print(json.dumps(data, indent=2))
else:
print(format_compact(data))
if __name__ == "__main__":
main()
+1
View File
@@ -0,0 +1 @@
# seo-agi scripts library
+244
View File
@@ -0,0 +1,244 @@
"""
DataForSEO API client for SEO-AGI.
Handles SERP results, keyword data, People Also Ask, and content parsing.
"""
import json
import base64
import urllib.request
import urllib.error
from typing import Optional
class DataForSEOClient:
"""Client for DataForSEO REST API v3."""
BASE_URL = "https://api.dataforseo.com/v3"
def __init__(self, login: str, password: str):
self.login = login
self.password = password
self._auth_header = self._make_auth_header(login, password)
@staticmethod
def _make_auth_header(login: str, password: str) -> str:
token = base64.b64encode(f"{login}:{password}".encode()).decode()
return f"Basic {token}"
def _request(self, endpoint: str, payload: list[dict]) -> dict:
"""Make a POST request to DataForSEO API."""
url = f"{self.BASE_URL}{endpoint}"
data = json.dumps(payload).encode("utf-8")
req = urllib.request.Request(
url,
data=data,
headers={
"Authorization": self._auth_header,
"Content-Type": "application/json",
},
method="POST",
)
try:
with urllib.request.urlopen(req, timeout=30) as resp:
return json.loads(resp.read().decode())
except urllib.error.HTTPError as e:
body = e.read().decode() if e.fp else ""
raise RuntimeError(
f"DataForSEO API error {e.code}: {body}"
) from e
except urllib.error.URLError as e:
raise RuntimeError(
f"DataForSEO connection error: {e.reason}"
) from e
def serp_live(
self,
keyword: str,
location_code: int = 2840,
language_code: str = "en",
depth: int = 10,
) -> dict:
"""
Get live SERP results for a keyword.
Returns organic results with position, URL, title, description.
"""
payload = [
{
"keyword": keyword,
"location_code": location_code,
"language_code": language_code,
"depth": depth,
"se_type": "organic",
}
]
result = self._request(
"/serp/google/organic/live/advanced", payload
)
return self._extract_serp(result)
def related_keywords(
self,
keyword: str,
location_code: int = 2840,
language_code: str = "en",
limit: int = 30,
) -> list[dict]:
"""Get related keywords with search volume and difficulty."""
payload = [
{
"keyword": keyword,
"location_code": location_code,
"language_code": language_code,
"limit": limit,
}
]
result = self._request(
"/dataforseo_labs/google/related_keywords/live", payload
)
return self._extract_keywords(result)
def keyword_suggestions(
self,
keyword: str,
location_code: int = 2840,
language_code: str = "en",
limit: int = 30,
) -> list[dict]:
"""Get keyword suggestions (broader ideation)."""
payload = [
{
"keyword": keyword,
"location_code": location_code,
"language_code": language_code,
"limit": limit,
}
]
result = self._request(
"/dataforseo_labs/google/keyword_suggestions/live", payload
)
return self._extract_keywords(result)
def content_parse(self, url: str) -> Optional[dict]:
"""Parse content from a URL (headings, word count, structure)."""
payload = [{"url": url}]
try:
result = self._request(
"/on_page/content_parsing/live", payload
)
return self._extract_content(result)
except RuntimeError:
return None
def _extract_serp(self, raw: dict) -> dict:
"""Extract clean SERP data from API response."""
tasks = raw.get("tasks", [])
if not tasks:
return {"organic": [], "paa": [], "featured_snippet": None}
result = tasks[0].get("result", [])
if not result:
return {"organic": [], "paa": [], "featured_snippet": None}
items = result[0].get("items", [])
organic = []
paa_questions = []
featured_snippet = None
for item in items:
item_type = item.get("type", "")
if item_type == "organic":
organic.append(
{
"position": item.get("rank_absolute", 0),
"url": item.get("url", ""),
"domain": item.get("domain", ""),
"title": item.get("title", ""),
"description": item.get("description", ""),
}
)
elif item_type == "people_also_ask":
for paa_item in item.get("items", []):
q = paa_item.get("title", "")
if q:
paa_questions.append(q)
elif item_type == "featured_snippet":
featured_snippet = {
"url": item.get("url", ""),
"title": item.get("title", ""),
"description": item.get("description", ""),
}
return {
"organic": organic,
"paa": paa_questions,
"featured_snippet": featured_snippet,
"total_results": result[0].get("se_results_count", 0),
}
def _extract_keywords(self, raw: dict) -> list[dict]:
"""Extract keyword data from labs API response."""
tasks = raw.get("tasks", [])
if not tasks:
return []
result = tasks[0].get("result", [])
if not result:
return []
items = result[0].get("items", [])
keywords = []
for item in items:
kw_data = item.get("keyword_data", item)
keyword_info = kw_data.get("keyword_info", {})
keywords.append(
{
"keyword": kw_data.get("keyword", ""),
"volume": keyword_info.get("search_volume", 0),
"cpc": keyword_info.get("cpc", 0),
"competition": keyword_info.get("competition", 0),
"difficulty": kw_data.get(
"keyword_properties", {}
).get("keyword_difficulty", 0),
}
)
return sorted(keywords, key=lambda x: x["volume"], reverse=True)
def _extract_content(self, raw: dict) -> Optional[dict]:
"""Extract content structure from on-page parsing."""
tasks = raw.get("tasks", [])
if not tasks:
return None
result = tasks[0].get("result", [])
if not result:
return None
items = result[0].get("items", [])
if not items:
return None
page = items[0].get("page_content", {})
return {
"title": page.get("header", {}).get("title", ""),
"word_count": page.get("plain_text_word_count", 0),
"headings": self._extract_headings(page),
"plain_text_size": page.get("plain_text_size", 0),
}
@staticmethod
def _extract_headings(page_content: dict) -> list[str]:
"""Pull heading tags from parsed content."""
headings = []
for level in ["h1", "h2", "h3"]:
for heading in page_content.get(level, []):
headings.append(f"{level.upper()}: {heading}")
return headings
+147
View File
@@ -0,0 +1,147 @@
"""
seo-agi environment and configuration loader.
Reads API keys from ~/.config/seo-agi/.env or os.environ.
Resolves paths relative to the skill installation directory.
"""
import os
import json
from pathlib import Path
# Skill root is two levels up from this file (scripts/lib/env.py -> .)
SKILL_DIR = Path(__file__).resolve().parent.parent.parent
OUTPUT_DIR = Path.home() / "Documents" / "SEO-AGI"
DATA_DIR = Path.home() / ".local" / "share" / "seo-agi"
CONFIG_DIR = Path.home() / ".config" / "seo-agi"
ENV_FILE = CONFIG_DIR / ".env"
DEFAULT_CONFIG = {
"default_location": 2840,
"default_language": "en",
"default_site": "",
"serp_depth": 10,
"save_research": True,
"output_dir": str(OUTPUT_DIR),
}
def load_env() -> dict:
"""
Load environment variables.
Reads from ~/.config/seo-agi/.env first, then overlays os.environ.
"""
env = {}
# First, try the config file
if ENV_FILE.exists():
with open(ENV_FILE, "r") as f:
for line in f:
line = line.strip()
if not line or line.startswith("#"):
continue
if "=" in line:
key, _, value = line.partition("=")
key = key.strip()
value = value.strip().strip('"').strip("'")
if value:
env[key] = value
# Then overlay with os.environ
for key in [
"DATAFORSEO_LOGIN", "DATAFORSEO_PASSWORD",
"GSC_SERVICE_ACCOUNT_PATH", "GSC_CLIENT_ID",
"GSC_CLIENT_SECRET", "GSC_REFRESH_TOKEN",
"AHREFS_API_KEY", "SEMRUSH_API_KEY",
]:
val = os.environ.get(key)
if val:
env[key] = val
return env
def load_config() -> dict:
"""Load user config, merged with defaults."""
config = DEFAULT_CONFIG.copy()
config_file = CONFIG_DIR / "config.json"
if config_file.exists():
try:
with open(config_file, "r") as f:
user_config = json.load(f)
config.update(user_config)
except (json.JSONDecodeError, IOError):
pass
return config
def get_credentials() -> dict:
"""Get API credentials with availability flags."""
env = load_env()
creds = {
"dataforseo_login": env.get("DATAFORSEO_LOGIN", ""),
"dataforseo_password": env.get("DATAFORSEO_PASSWORD", ""),
"gsc_service_account_path": env.get("GSC_SERVICE_ACCOUNT_PATH", ""),
"gsc_client_id": env.get("GSC_CLIENT_ID", ""),
"gsc_client_secret": env.get("GSC_CLIENT_SECRET", ""),
"gsc_refresh_token": env.get("GSC_REFRESH_TOKEN", ""),
"ahrefs_api_key": env.get("AHREFS_API_KEY", ""),
"semrush_api_key": env.get("SEMRUSH_API_KEY", ""),
}
creds["has_dataforseo"] = bool(
creds["dataforseo_login"] and creds["dataforseo_password"]
)
creds["has_gsc"] = bool(
creds["gsc_service_account_path"]
or (creds["gsc_client_id"] and creds["gsc_client_secret"])
)
creds["has_ahrefs"] = bool(creds["ahrefs_api_key"])
creds["has_semrush"] = bool(creds["semrush_api_key"])
return creds
def ensure_dirs():
"""Create output directories if they don't exist."""
config = load_config()
output_dir = Path(config["output_dir"]).expanduser()
for subdir in ["research", "briefs", "pages", "rewrites"]:
(output_dir / subdir).mkdir(parents=True, exist_ok=True)
(DATA_DIR / "research").mkdir(parents=True, exist_ok=True)
(DATA_DIR / "cache").mkdir(parents=True, exist_ok=True)
def check_setup() -> dict:
"""Check setup status and return a summary."""
creds = get_credentials()
config = load_config()
return {
"runtime": "claude-code",
"skill_dir": str(SKILL_DIR),
"config_dir_exists": CONFIG_DIR.exists(),
"env_file_exists": ENV_FILE.exists(),
"has_dataforseo": creds["has_dataforseo"],
"has_gsc": creds["has_gsc"],
"has_ahrefs": creds["has_ahrefs"],
"has_semrush": creds["has_semrush"],
"default_location": config["default_location"],
"default_language": config["default_language"],
"mode": _determine_mode(creds),
}
def _determine_mode(creds: dict) -> str:
"""Determine operational mode based on available credentials."""
if creds["has_dataforseo"] and creds["has_gsc"]:
return "full"
elif creds["has_dataforseo"]:
return "dataforseo-only"
elif creds["has_gsc"]:
return "gsc-only"
else:
return "fallback"
+179
View File
@@ -0,0 +1,179 @@
"""
Google Search Console API client for SEO-AGI.
Pulls query performance data, cannibalization detection, and indexing status.
"""
import json
from typing import Optional
from pathlib import Path
class GSCClient:
"""Client for Google Search Console API."""
def __init__(self, credentials_path: str = None, oauth_creds: dict = None):
"""
Initialize GSC client.
Args:
credentials_path: Path to service account JSON file
oauth_creds: Dict with client_id, client_secret, refresh_token
"""
self.credentials_path = credentials_path
self.oauth_creds = oauth_creds
self._service = None
def _get_service(self):
"""Lazy-initialize the GSC API service."""
if self._service is not None:
return self._service
try:
from google.oauth2 import service_account
from googleapiclient.discovery import build
except ImportError:
raise RuntimeError(
"GSC requires google-auth and google-api-python-client. "
"Install with: pip install google-auth google-api-python-client"
)
if self.credentials_path:
creds = service_account.Credentials.from_service_account_file(
self.credentials_path,
scopes=["https://www.googleapis.com/auth/webmasters.readonly"],
)
else:
raise RuntimeError(
"OAuth2 flow not yet implemented. Use a service account."
)
self._service = build("searchconsole", "v1", credentials=creds)
return self._service
def query_performance(
self,
site_url: str,
keyword: str = None,
days: int = 90,
min_impressions: int = 10,
row_limit: int = 100,
) -> list[dict]:
"""
Pull query performance data from GSC.
Args:
site_url: The GSC property URL (e.g., "https://example.com")
keyword: Optional keyword filter (partial match)
days: Lookback period in days
min_impressions: Minimum impressions threshold
row_limit: Max rows to return
Returns:
List of dicts with query, page, clicks, impressions, ctr, position
"""
from datetime import datetime, timedelta
service = self._get_service()
end_date = datetime.now().strftime("%Y-%m-%d")
start_date = (datetime.now() - timedelta(days=days)).strftime("%Y-%m-%d")
request_body = {
"startDate": start_date,
"endDate": end_date,
"dimensions": ["query", "page"],
"rowLimit": row_limit,
"dimensionFilterGroups": [],
}
if keyword:
request_body["dimensionFilterGroups"].append(
{
"filters": [
{
"dimension": "query",
"operator": "contains",
"expression": keyword,
}
]
}
)
response = (
service.searchanalytics()
.query(siteUrl=site_url, body=request_body)
.execute()
)
rows = response.get("rows", [])
results = []
for row in rows:
impressions = row.get("impressions", 0)
if impressions < min_impressions:
continue
results.append(
{
"query": row["keys"][0],
"page": row["keys"][1],
"clicks": row.get("clicks", 0),
"impressions": impressions,
"ctr": round(row.get("ctr", 0) * 100, 2),
"position": round(row.get("position", 0), 1),
}
)
return sorted(results, key=lambda x: x["impressions"], reverse=True)
def detect_cannibalization(
self,
site_url: str,
keyword: str,
days: int = 90,
) -> list[dict]:
"""
Detect keyword cannibalization: multiple pages ranking for same query.
Returns:
List of queries where 2+ pages from the site appear,
sorted by total impressions.
"""
results = self.query_performance(
site_url=site_url,
keyword=keyword,
days=days,
min_impressions=5,
row_limit=500,
)
# Group by query
query_pages = {}
for row in results:
q = row["query"]
if q not in query_pages:
query_pages[q] = []
query_pages[q].append(row)
# Find queries with multiple pages
cannibalized = []
for query, pages in query_pages.items():
if len(pages) > 1:
cannibalized.append(
{
"query": query,
"page_count": len(pages),
"pages": sorted(
pages, key=lambda x: x["position"]
),
"total_impressions": sum(
p["impressions"] for p in pages
),
}
)
return sorted(
cannibalized,
key=lambda x: x["total_impressions"],
reverse=True,
)
+261
View File
@@ -0,0 +1,261 @@
"""
SERP content analyzer for SEO-AGI.
Analyzes competitor content structure and identifies content gaps.
"""
import statistics
from typing import Optional
def analyze_serp(
serp_data: dict,
content_data: list[Optional[dict]],
target_keyword: str,
) -> dict:
"""
Analyze SERP results + parsed competitor content.
Args:
serp_data: Output from DataForSEOClient.serp_live()
content_data: List of DataForSEOClient.content_parse() results
(one per organic result, may contain None for failed parses)
target_keyword: The keyword being targeted
Returns:
Analysis dict with intent, word count stats, topic coverage, gaps.
"""
organic = serp_data.get("organic", [])
paa = serp_data.get("paa", [])
# Word count analysis
word_counts = []
all_headings = []
all_topics = set()
for i, content in enumerate(content_data):
if content is None:
continue
wc = content.get("word_count", 0)
if wc > 0:
word_counts.append(wc)
headings = content.get("headings", [])
all_headings.extend(headings)
# Extract topics from H2/H3 headings
for h in headings:
if h.startswith(("H2:", "H3:")):
topic = h.split(":", 1)[1].strip().lower()
all_topics.add(topic)
# Word count statistics
wc_stats = {}
if word_counts:
wc_stats = {
"min": min(word_counts),
"max": max(word_counts),
"median": int(statistics.median(word_counts)),
"mean": int(statistics.mean(word_counts)),
"recommended_min": int(statistics.median(word_counts) * 0.8),
"recommended_max": int(statistics.median(word_counts) * 1.3),
"count_analyzed": len(word_counts),
}
# Intent detection
intent = detect_intent(target_keyword, organic, serp_data)
# Topic frequency (which topics appear in multiple competitors)
topic_freq = _count_topic_frequency(content_data)
# Heading patterns
heading_patterns = _analyze_heading_patterns(content_data)
return {
"keyword": target_keyword,
"intent": intent,
"word_count_stats": wc_stats,
"paa_questions": paa,
"topic_frequency": topic_freq,
"heading_patterns": heading_patterns,
"competitors_analyzed": len(
[c for c in content_data if c is not None]
),
"total_organic_results": len(organic),
"featured_snippet": serp_data.get("featured_snippet"),
}
def detect_intent(
keyword: str, organic: list[dict], serp_data: dict
) -> str:
"""
Detect search intent from keyword and SERP features.
Returns: informational, commercial, transactional, or navigational
"""
kw_lower = keyword.lower()
# Navigational signals
nav_signals = [
"login",
"sign in",
"website",
"official",
".com",
".org",
]
if any(s in kw_lower for s in nav_signals):
return "navigational"
# Transactional signals
transactional_signals = [
"buy",
"purchase",
"order",
"download",
"subscribe",
"deal",
"discount",
"coupon",
"price",
"pricing",
"cost",
"cheap",
"free trial",
]
if any(s in kw_lower for s in transactional_signals):
return "transactional"
# Commercial investigation signals
commercial_signals = [
"best",
"top",
"review",
"comparison",
"vs",
"versus",
"alternative",
"vs.",
"compared to",
"pros and cons",
]
if any(s in kw_lower for s in commercial_signals):
return "commercial"
# Informational signals
informational_signals = [
"how to",
"what is",
"what are",
"why",
"guide",
"tutorial",
"learn",
"example",
"definition",
"meaning",
"explain",
]
if any(s in kw_lower for s in informational_signals):
return "informational"
# Default: check SERP patterns
if serp_data.get("featured_snippet"):
return "informational"
# If titles contain pricing/comparison language
titles = [r.get("title", "").lower() for r in organic[:5]]
title_text = " ".join(titles)
if any(s in title_text for s in ["best", "top", "review", "vs"]):
return "commercial"
if any(
s in title_text for s in ["how to", "guide", "what", "tutorial"]
):
return "informational"
return "commercial" # default for ambiguous
def _count_topic_frequency(
content_data: list[Optional[dict]],
) -> list[dict]:
"""Count how often topics (H2/H3 headings) appear across competitors."""
topic_counts = {}
for content in content_data:
if content is None:
continue
seen_in_page = set()
for heading in content.get("headings", []):
if heading.startswith(("H2:", "H3:")):
topic = heading.split(":", 1)[1].strip().lower()
# Normalize common variations
topic = _normalize_topic(topic)
if topic and topic not in seen_in_page:
seen_in_page.add(topic)
topic_counts[topic] = topic_counts.get(topic, 0) + 1
# Sort by frequency
sorted_topics = sorted(
topic_counts.items(), key=lambda x: x[1], reverse=True
)
return [
{"topic": t, "competitor_count": c} for t, c in sorted_topics[:30]
]
def _normalize_topic(topic: str) -> str:
"""Basic topic normalization."""
# Remove common filler words at start
for prefix in [
"the ",
"a ",
"an ",
"our ",
"your ",
"my ",
"about ",
]:
if topic.startswith(prefix):
topic = topic[len(prefix) :]
# Strip trailing punctuation
topic = topic.rstrip(".:!?")
return topic.strip()
def _analyze_heading_patterns(
content_data: list[Optional[dict]],
) -> dict:
"""Analyze heading patterns across competitors."""
h2_counts = []
h3_counts = []
for content in content_data:
if content is None:
continue
headings = content.get("headings", [])
h2_count = sum(1 for h in headings if h.startswith("H2:"))
h3_count = sum(1 for h in headings if h.startswith("H3:"))
h2_counts.append(h2_count)
h3_counts.append(h3_count)
return {
"avg_h2_count": (
round(statistics.mean(h2_counts), 1) if h2_counts else 0
),
"avg_h3_count": (
round(statistics.mean(h3_counts), 1) if h3_counts else 0
),
"median_h2_count": (
int(statistics.median(h2_counts)) if h2_counts else 0
),
"median_h3_count": (
int(statistics.median(h3_counts)) if h3_counts else 0
),
}
+334
View File
@@ -0,0 +1,334 @@
#!/usr/bin/env python3
"""
seo-agi research orchestrator.
Pulls SERP data, keyword data, PAA questions, and competitor content analysis
into a single research.json file that feeds content generation.
Usage:
python3 research.py "<keyword>" [options]
Options:
--serp-depth=N Number of SERP results to analyze (default: 10)
--include-paa Include People Also Ask extraction (default: true)
--location=CODE DataForSEO location code (default: 2840 = US)
--language=CODE Language code (default: en)
--output=FORMAT Output: json|compact|brief (default: compact)
--save-dir=PATH Save raw data (default: ~/.local/share/seo-agi/research/)
--content-depth=N Number of top results to parse for content (default: 5)
--mock Use fixture data instead of live API calls
"""
import sys
import os
import json
import argparse
from datetime import datetime, timezone
from pathlib import Path
# Add parent dir to path for lib imports
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
from lib.env import load_env, load_config, get_credentials, ensure_dirs
from lib.dataforseo import DataForSEOClient
from lib.serp_analyze import analyze_serp
def parse_args():
parser = argparse.ArgumentParser(description="SEO-AGI Research")
parser.add_argument("keyword", help="Target keyword or topic")
parser.add_argument(
"--serp-depth", type=int, default=10, help="SERP depth"
)
parser.add_argument(
"--content-depth",
type=int,
default=5,
help="Number of competitors to parse content from",
)
parser.add_argument(
"--location", type=int, default=None, help="Location code"
)
parser.add_argument(
"--language", default=None, help="Language code"
)
parser.add_argument(
"--output",
choices=["json", "compact", "brief"],
default="compact",
help="Output format",
)
parser.add_argument(
"--save-dir", default=None, help="Directory to save research data"
)
parser.add_argument(
"--mock",
action="store_true",
help="Use fixture data (no API calls)",
)
return parser.parse_args()
def load_mock_data(keyword: str) -> dict:
"""Load fixture data for testing without API calls."""
fixtures_dir = (
Path(__file__).parent.parent / "fixtures"
)
serp_fixture = fixtures_dir / "serp_sample.json"
keywords_fixture = fixtures_dir / "keywords_sample.json"
serp_data = {"organic": [], "paa": [], "featured_snippet": None}
related_kw = []
if serp_fixture.exists():
with open(serp_fixture) as f:
serp_data = json.load(f)
if keywords_fixture.exists():
with open(keywords_fixture) as f:
related_kw = json.load(f)
return {
"keyword": keyword,
"timestamp": datetime.now(timezone.utc).isoformat(),
"source": "mock",
"serp": serp_data,
"related_keywords": related_kw,
"analysis": {
"intent": "commercial",
"word_count_stats": {
"min": 800,
"max": 3200,
"median": 1800,
"recommended_min": 1440,
"recommended_max": 2340,
},
"paa_questions": serp_data.get("paa", []),
"topic_frequency": [],
"heading_patterns": {
"avg_h2_count": 6,
"avg_h3_count": 8,
"median_h2_count": 5,
"median_h3_count": 7,
},
},
}
def run_research(args) -> dict:
"""Execute the full research pipeline."""
creds = get_credentials()
config = load_config()
location = args.location or config["default_location"]
language = args.language or config["default_language"]
if args.mock:
return load_mock_data(args.keyword)
if not creds["has_dataforseo"]:
print(
"ERROR: DataForSEO credentials not found.",
file=sys.stderr,
)
print(
"Run: python3 scripts/setup.py",
file=sys.stderr,
)
print(
"Or add DATAFORSEO_LOGIN and DATAFORSEO_PASSWORD to "
"~/.config/seo-agi/.env",
file=sys.stderr,
)
sys.exit(1)
client = DataForSEOClient(
creds["dataforseo_login"], creds["dataforseo_password"]
)
# Step 1: SERP results
print(f"Fetching SERP for: {args.keyword}", file=sys.stderr)
serp_data = client.serp_live(
args.keyword, location, language, args.serp_depth
)
# Step 2: Related keywords
print("Fetching related keywords...", file=sys.stderr)
related_kw = client.related_keywords(args.keyword, location, language)
# Step 3: Parse competitor content (top N)
content_data = []
organic = serp_data.get("organic", [])
parse_count = min(args.content_depth, len(organic))
for i in range(parse_count):
url = organic[i].get("url", "")
if url:
print(
f"Parsing content ({i+1}/{parse_count}): {url[:80]}...",
file=sys.stderr,
)
content = client.content_parse(url)
content_data.append(content)
# Merge content data back into organic result
if content:
organic[i]["word_count"] = content.get("word_count", 0)
organic[i]["headings"] = content.get("headings", [])
else:
content_data.append(None)
# Step 4: Analyze
print("Analyzing competitive landscape...", file=sys.stderr)
analysis = analyze_serp(serp_data, content_data, args.keyword)
# Assemble research output
research = {
"keyword": args.keyword,
"timestamp": datetime.now(timezone.utc).isoformat(),
"location": location,
"language": language,
"source": "dataforseo",
"serp": serp_data,
"related_keywords": related_kw[:20],
"analysis": analysis,
}
return research
def save_research(research: dict, save_dir: str = None):
"""Save research data to disk."""
ensure_dirs()
config = load_config()
if save_dir:
out_dir = Path(save_dir).expanduser()
else:
out_dir = Path.home() / ".local" / "share" / "seo-agi" / "research"
out_dir.mkdir(parents=True, exist_ok=True)
# Filename from keyword + date
slug = (
research["keyword"]
.lower()
.replace(" ", "-")
.replace("/", "-")[:50]
)
date_str = datetime.now().strftime("%Y%m%d")
filename = f"{slug}-{date_str}.json"
filepath = out_dir / filename
with open(filepath, "w") as f:
json.dump(research, f, indent=2)
print(f"Research saved: {filepath}", file=sys.stderr)
return str(filepath)
def format_compact(research: dict) -> str:
"""Format research as compact human-readable output."""
lines = []
kw = research["keyword"]
analysis = research.get("analysis", {})
serp = research.get("serp", {})
organic = serp.get("organic", [])
lines.append(f"# Research: {kw}")
lines.append(f"Intent: {analysis.get('intent', 'unknown')}")
# Word count
wc = analysis.get("word_count_stats", {})
if wc:
lines.append(
f"Competitor word count: {wc.get('min', '?')}-{wc.get('max', '?')} "
f"(median: {wc.get('median', '?')})"
)
lines.append(
f"Recommended range: {wc.get('recommended_min', '?')}-"
f"{wc.get('recommended_max', '?')} words"
)
# Top results
lines.append(f"\n## Top {len(organic)} Results")
for r in organic[:10]:
wc_str = (
f" ({r.get('word_count', '?')} words)"
if r.get("word_count")
else ""
)
lines.append(f" {r['position']}. {r['title']}{wc_str}")
lines.append(f" {r['url']}")
# PAA
paa = analysis.get("paa_questions", serp.get("paa", []))
if paa:
lines.append(f"\n## People Also Ask ({len(paa)})")
for q in paa:
lines.append(f" - {q}")
# Related keywords
related = research.get("related_keywords", [])
if related:
lines.append(f"\n## Related Keywords (top 10)")
for kw_data in related[:10]:
lines.append(
f" - {kw_data['keyword']} "
f"(vol: {kw_data['volume']}, "
f"diff: {kw_data.get('difficulty', '?')})"
)
# Topics
topics = analysis.get("topic_frequency", [])
if topics:
lines.append(f"\n## Common Topics Across Competitors")
for t in topics[:15]:
lines.append(
f" - {t['topic']} (in {t['competitor_count']} pages)"
)
# Heading patterns
hp = analysis.get("heading_patterns", {})
if hp:
lines.append(f"\n## Heading Structure")
lines.append(
f" Avg H2s: {hp.get('avg_h2_count', '?')}, "
f"Avg H3s: {hp.get('avg_h3_count', '?')}"
)
return "\n".join(lines)
def main():
args = parse_args()
research = run_research(args)
# Save
filepath = save_research(research, args.save_dir)
# Output
if args.output == "json":
print(json.dumps(research, indent=2))
elif args.output == "brief":
# Minimal output for piping into content generation
analysis = research.get("analysis", {})
brief_data = {
"keyword": research["keyword"],
"intent": analysis.get("intent"),
"word_count_stats": analysis.get("word_count_stats"),
"paa_questions": analysis.get(
"paa_questions",
research.get("serp", {}).get("paa", []),
),
"topic_frequency": analysis.get("topic_frequency", [])[:10],
"heading_patterns": analysis.get("heading_patterns"),
"research_file": filepath,
}
print(json.dumps(brief_data, indent=2))
else:
print(format_compact(research))
if __name__ == "__main__":
main()
+160
View File
@@ -0,0 +1,160 @@
#!/usr/bin/env python3
"""
SEO-AGI setup script.
Creates config directory, prompts for API keys, installs dependencies.
Usage:
python3 setup.py
"""
import os
import sys
import json
import subprocess
from pathlib import Path
CONFIG_DIR = Path.home() / ".config" / "seo-agi"
ENV_FILE = CONFIG_DIR / ".env"
CONFIG_FILE = CONFIG_DIR / "config.json"
OUTPUT_DIR = Path.home() / "Documents" / "SEO-AGI"
def main():
print("=" * 50)
print(" SEO-AGI Setup")
print("=" * 50)
print()
# Create directories
CONFIG_DIR.mkdir(parents=True, exist_ok=True)
for subdir in ["research", "briefs", "pages", "rewrites"]:
(OUTPUT_DIR / subdir).mkdir(parents=True, exist_ok=True)
data_dir = Path.home() / ".local" / "share" / "seo-agi"
for subdir in ["research", "cache"]:
(data_dir / subdir).mkdir(parents=True, exist_ok=True)
print("[OK] Directories created")
# Install Python dependencies
print("\nInstalling dependencies...")
deps = ["requests"]
try:
subprocess.check_call(
[sys.executable, "-m", "pip", "install", "--quiet"]
+ deps
+ ["--break-system-packages"],
stdout=subprocess.DEVNULL,
stderr=subprocess.DEVNULL,
)
print("[OK] Core dependencies installed (requests)")
except subprocess.CalledProcessError:
try:
subprocess.check_call(
[sys.executable, "-m", "pip", "install", "--quiet"] + deps,
stdout=subprocess.DEVNULL,
stderr=subprocess.DEVNULL,
)
print("[OK] Core dependencies installed (requests)")
except subprocess.CalledProcessError:
print("[WARN] Could not install requests. Install manually:")
print(" pip install requests")
# API Keys
print("\n--- API Keys ---")
print("DataForSEO is REQUIRED. Get credentials at:")
print(" https://app.dataforseo.com/api-dashboard\n")
existing_env = {}
if ENV_FILE.exists():
with open(ENV_FILE) as f:
for line in f:
line = line.strip()
if line and not line.startswith("#") and "=" in line:
k, _, v = line.partition("=")
existing_env[k.strip()] = v.strip()
# DataForSEO
dfs_login = existing_env.get("DATAFORSEO_LOGIN", "")
dfs_pass = existing_env.get("DATAFORSEO_PASSWORD", "")
if dfs_login:
print(f"DataForSEO login found: {dfs_login[:4]}***")
change = input("Update DataForSEO credentials? [y/N]: ").strip().lower()
if change == "y":
dfs_login = input("DataForSEO login (email): ").strip()
dfs_pass = input("DataForSEO password: ").strip()
else:
dfs_login = input("DataForSEO login (email): ").strip()
dfs_pass = input("DataForSEO password: ").strip()
# GSC (optional)
print("\nGoogle Search Console (OPTIONAL - press Enter to skip)")
gsc_path = existing_env.get("GSC_SERVICE_ACCOUNT_PATH", "")
if gsc_path:
print(f"GSC service account found: {gsc_path}")
change = input("Update? [y/N]: ").strip().lower()
if change == "y":
gsc_path = input("Path to service account JSON: ").strip()
else:
gsc_path = input("Path to GSC service account JSON (or Enter to skip): ").strip()
# Write .env
env_lines = [
"# SEO-AGI Configuration",
f"# Generated by setup.py on {__import__('datetime').datetime.now().isoformat()[:10]}",
"",
"# DataForSEO",
f"DATAFORSEO_LOGIN={dfs_login}",
f"DATAFORSEO_PASSWORD={dfs_pass}",
"",
"# Google Search Console",
f"GSC_SERVICE_ACCOUNT_PATH={gsc_path}",
"",
"# Future integrations",
f"AHREFS_API_KEY={existing_env.get('AHREFS_API_KEY', '')}",
f"SEMRUSH_API_KEY={existing_env.get('SEMRUSH_API_KEY', '')}",
"",
"# Defaults",
f"DEFAULT_LOCATION={existing_env.get('DEFAULT_LOCATION', '2840')}",
f"DEFAULT_LANGUAGE={existing_env.get('DEFAULT_LANGUAGE', 'en')}",
]
with open(ENV_FILE, "w") as f:
f.write("\n".join(env_lines) + "\n")
print(f"\n[OK] Credentials saved to {ENV_FILE}")
# Write default config.json if not exists
if not CONFIG_FILE.exists():
default_config = {
"default_location": 2840,
"default_language": "en",
"default_site": "",
"serp_depth": 10,
"save_research": True,
"output_dir": str(OUTPUT_DIR),
}
with open(CONFIG_FILE, "w") as f:
json.dump(default_config, f, indent=2)
print(f"[OK] Config saved to {CONFIG_FILE}")
# Verify
print("\n--- Setup Summary ---")
print(f" DataForSEO: {'configured' if dfs_login else 'NOT SET'}")
print(f" GSC: {'configured' if gsc_path else 'not configured (optional)'}")
print(f" Output dir: {OUTPUT_DIR}")
print(f" Config dir: {CONFIG_DIR}")
if dfs_login:
print("\n[READY] Run a test:")
print(f" python3 {Path(__file__).parent}/research.py \"test keyword\"")
else:
print("\n[WARN] DataForSEO credentials missing. The skill will fall")
print(" back to Claude's web search for basic research.")
print()
if __name__ == "__main__":
main()
+149
View File
@@ -0,0 +1,149 @@
"""Tests for DataForSEO client response parsing."""
import sys
from pathlib import Path
sys.path.insert(
0, str(Path(__file__).parent.parent / "scripts")
)
from lib.dataforseo import DataForSEOClient
def test_extract_serp_empty():
client = DataForSEOClient("test", "test")
result = client._extract_serp({"tasks": []})
assert result["organic"] == []
assert result["paa"] == []
assert result["featured_snippet"] is None
def test_extract_serp_with_organic():
raw = {
"tasks": [{
"result": [{
"se_results_count": 1000000,
"items": [
{
"type": "organic",
"rank_absolute": 1,
"url": "https://example.com/page",
"domain": "example.com",
"title": "Test Page",
"description": "A test description",
},
{
"type": "organic",
"rank_absolute": 2,
"url": "https://other.com/page",
"domain": "other.com",
"title": "Other Page",
"description": "Another description",
},
{
"type": "people_also_ask",
"items": [
{"title": "What is a test?"},
{"title": "How do tests work?"},
],
},
]
}]
}]
}
client = DataForSEOClient("test", "test")
result = client._extract_serp(raw)
assert len(result["organic"]) == 2
assert result["organic"][0]["position"] == 1
assert result["organic"][0]["url"] == "https://example.com/page"
assert result["organic"][1]["title"] == "Other Page"
assert len(result["paa"]) == 2
assert result["paa"][0] == "What is a test?"
assert result["total_results"] == 1000000
def test_extract_keywords_empty():
client = DataForSEOClient("test", "test")
result = client._extract_keywords({"tasks": []})
assert result == []
def test_extract_keywords_with_data():
raw = {
"tasks": [{
"result": [{
"items": [
{
"keyword_data": {
"keyword": "test keyword one",
"keyword_info": {
"search_volume": 5000,
"cpc": 1.50,
"competition": 0.7,
},
"keyword_properties": {
"keyword_difficulty": 42,
},
}
},
{
"keyword_data": {
"keyword": "test keyword two",
"keyword_info": {
"search_volume": 8000,
"cpc": 2.10,
"competition": 0.85,
},
"keyword_properties": {
"keyword_difficulty": 55,
},
}
},
]
}]
}]
}
client = DataForSEOClient("test", "test")
result = client._extract_keywords(raw)
# Should be sorted by volume descending
assert len(result) == 2
assert result[0]["keyword"] == "test keyword two"
assert result[0]["volume"] == 8000
assert result[1]["keyword"] == "test keyword one"
assert result[1]["volume"] == 5000
assert result[1]["difficulty"] == 42
def test_extract_headings():
page_content = {
"h1": ["Main Title"],
"h2": ["Section One", "Section Two"],
"h3": ["Subsection A", "Subsection B"],
}
headings = DataForSEOClient._extract_headings(page_content)
assert "H1: Main Title" in headings
assert "H2: Section One" in headings
assert "H3: Subsection A" in headings
assert len(headings) == 5
def test_auth_header():
client = DataForSEOClient("user@test.com", "mypassword")
assert client._auth_header.startswith("Basic ")
assert len(client._auth_header) > 10
if __name__ == "__main__":
test_extract_serp_empty()
test_extract_serp_with_organic()
test_extract_keywords_empty()
test_extract_keywords_with_data()
test_extract_headings()
test_auth_header()
print("All tests passed.")
+50
View File
@@ -0,0 +1,50 @@
"""Tests for env module."""
import sys
import os
import json
import tempfile
from pathlib import Path
sys.path.insert(
0, str(Path(__file__).parent.parent / "scripts")
)
from lib.env import load_config, _determine_mode, DEFAULT_CONFIG
def test_default_config():
config = DEFAULT_CONFIG.copy()
assert config["default_location"] == 2840
assert config["default_language"] == "en"
assert config["serp_depth"] == 10
assert config["save_research"] is True
def test_determine_mode_full():
creds = {"has_dataforseo": True, "has_gsc": True, "has_ahrefs": False, "has_semrush": False}
assert _determine_mode(creds) == "full"
def test_determine_mode_dataforseo_only():
creds = {"has_dataforseo": True, "has_gsc": False, "has_ahrefs": False, "has_semrush": False}
assert _determine_mode(creds) == "dataforseo-only"
def test_determine_mode_gsc_only():
creds = {"has_dataforseo": False, "has_gsc": True, "has_ahrefs": False, "has_semrush": False}
assert _determine_mode(creds) == "gsc-only"
def test_determine_mode_fallback():
creds = {"has_dataforseo": False, "has_gsc": False, "has_ahrefs": False, "has_semrush": False}
assert _determine_mode(creds) == "fallback"
if __name__ == "__main__":
test_default_config()
test_determine_mode_full()
test_determine_mode_dataforseo_only()
test_determine_mode_gsc_only()
test_determine_mode_fallback()
print("All tests passed.")
+125
View File
@@ -0,0 +1,125 @@
"""Tests for serp_analyze module."""
import sys
import os
import json
from pathlib import Path
# Add scripts dir to path
sys.path.insert(
0, str(Path(__file__).parent.parent / "scripts")
)
from lib.serp_analyze import analyze_serp, detect_intent
def load_fixture(name: str):
fixtures_dir = Path(__file__).parent.parent / "fixtures"
with open(fixtures_dir / name) as f:
return json.load(f)
def test_detect_intent_commercial():
assert detect_intent("best project management tools", [], {}) == "commercial"
assert detect_intent("X vs Y comparison", [], {}) == "commercial"
assert detect_intent("slack alternatives", [], {}) == "commercial"
def test_detect_intent_informational():
assert detect_intent("how to park at JFK", [], {}) == "informational"
assert detect_intent("what is SEO", [], {}) == "informational"
assert detect_intent("python tutorial for beginners", [], {}) == "informational"
def test_detect_intent_transactional():
assert detect_intent("buy parking pass JFK", [], {}) == "transactional"
assert detect_intent("JFK parking coupon discount", [], {}) == "transactional"
assert detect_intent("cheap parking near me", [], {}) == "transactional"
def test_detect_intent_navigational():
assert detect_intent("JFK airport official website", [], {}) == "navigational"
assert detect_intent("parkingaccess.com login", [], {}) == "navigational"
def test_analyze_serp_with_fixtures():
serp_data = load_fixture("serp_sample.json")
# Build content_data from the fixture's embedded word_count/headings
content_data = []
for item in serp_data.get("organic", []):
if item.get("word_count"):
content_data.append({
"word_count": item["word_count"],
"headings": item.get("headings", []),
})
else:
content_data.append(None)
analysis = analyze_serp(serp_data, content_data, "airport parking JFK")
assert analysis["keyword"] == "airport parking JFK"
assert analysis["intent"] in (
"informational", "commercial", "transactional", "navigational"
)
# Word count stats should reflect fixture data
wc = analysis["word_count_stats"]
assert wc["min"] > 0
assert wc["max"] >= wc["min"]
assert wc["median"] > 0
assert wc["recommended_min"] > 0
assert wc["recommended_max"] > wc["recommended_min"]
# PAA should come through
assert len(analysis["paa_questions"]) > 0
# Topic frequency should have entries from headings
assert len(analysis["topic_frequency"]) > 0
# Heading patterns should be populated
hp = analysis["heading_patterns"]
assert hp["avg_h2_count"] > 0
def test_analyze_serp_empty():
"""Handle empty SERP data gracefully."""
analysis = analyze_serp(
{"organic": [], "paa": [], "featured_snippet": None},
[],
"nonexistent keyword",
)
assert analysis["keyword"] == "nonexistent keyword"
assert analysis["word_count_stats"] == {}
assert analysis["competitors_analyzed"] == 0
def test_analyze_serp_partial_content():
"""Handle mix of parsed and failed content parses."""
serp_data = {
"organic": [
{"position": 1, "url": "https://a.com", "title": "A", "description": ""},
{"position": 2, "url": "https://b.com", "title": "B", "description": ""},
],
"paa": ["Question 1?"],
"featured_snippet": None,
}
content_data = [
{"word_count": 2000, "headings": ["H1: Title", "H2: Section One", "H2: Section Two"]},
None, # failed parse
]
analysis = analyze_serp(serp_data, content_data, "test keyword")
assert analysis["competitors_analyzed"] == 1
assert analysis["word_count_stats"]["median"] == 2000
if __name__ == "__main__":
test_detect_intent_commercial()
test_detect_intent_informational()
test_detect_intent_transactional()
test_detect_intent_navigational()
test_analyze_serp_with_fixtures()
test_analyze_serp_empty()
test_analyze_serp_partial_content()
print("All tests passed.")