mirror of
https://github.com/gbessoni/seobuild-onpage.git
synced 2026-06-23 11:58:37 +02:00
seo-agi: The GEO framework skill for Claude Code + OpenClaw
GEO framework that writes pages ranking on Google AND getting cited by LLMs. 500-token chunk architecture, Reddit Test quality gates, verification tags, Not For You blocks, information gain enforcement. Data layer: DataForSEO, GSC, Ahrefs MCP, SEMRush MCP. 21 files, all tests passing. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,7 @@
|
||||
__pycache__/
|
||||
*.pyc
|
||||
.env
|
||||
*.egg-info/
|
||||
dist/
|
||||
build/
|
||||
.DS_Store
|
||||
@@ -0,0 +1,59 @@
|
||||
# CLAUDE.md
|
||||
|
||||
This is **seo-agi** (The Beefy SEO Skill) -- a Claude Code skill for Generative Engine Optimization. It writes pages that rank on Google AND get cited by LLMs.
|
||||
|
||||
This is not a generic SEO prompt. It enforces 500-token chunk architecture, Reddit Test quality gates, verification tags, "Not For You" blocks, and real competitive data from DataForSEO/Ahrefs/SEMRush/GSC.
|
||||
|
||||
## Structure
|
||||
|
||||
```
|
||||
SKILL.md -- The framework (beefy SEO brain + data layer integration)
|
||||
SPEC.md -- Technical architecture for the data layer
|
||||
scripts/
|
||||
research.py -- CLI: SERP research via DataForSEO
|
||||
gsc_pull.py -- CLI: Google Search Console data
|
||||
setup.py -- Interactive first-run config
|
||||
lib/
|
||||
env.py -- Config loader
|
||||
dataforseo.py -- DataForSEO REST API client
|
||||
serp_analyze.py -- Content gap analysis engine
|
||||
gsc_client.py -- Google Search Console client
|
||||
references/
|
||||
page-templates.md -- Structural templates by page type
|
||||
schema-patterns.md -- JSON-LD schema patterns
|
||||
fixtures/ -- Mock data for testing without API calls
|
||||
tests/ -- Unit tests
|
||||
```
|
||||
|
||||
## The Framework (SKILL.md)
|
||||
|
||||
The SKILL.md is the living document. It contains:
|
||||
- Core belief system (anti-generic, LLM retrieval, entity consensus)
|
||||
- Google AI Search 7 ranking signals
|
||||
- 500-token chunk architecture
|
||||
- SEAT signals (Semantic + E-E-A-T + Entity/Knowledge Graph)
|
||||
- Quality gates: Reddit Test, Prove-It Details, Not For You, Information Gain Test
|
||||
- Verification tagging system ({{VERIFY}}, {{RESEARCH NEEDED}}, {{SOURCE NEEDED}})
|
||||
- Vertical-specific instructions (airport/parking, local service, listicle, comparison)
|
||||
- LLM/AEO citation strategy
|
||||
- Hub & spoke internal linking
|
||||
- Execution protocol with data layer integration
|
||||
|
||||
## Running Tests
|
||||
|
||||
```bash
|
||||
python3 tests/test_env.py
|
||||
python3 tests/test_serp_analyze.py
|
||||
python3 tests/test_dataforseo.py
|
||||
|
||||
# Mock mode research:
|
||||
python3 scripts/research.py "test keyword" --mock
|
||||
```
|
||||
|
||||
## Style
|
||||
|
||||
- No em dashes in any output
|
||||
- No "nestled", no "in today's fast-paced world", no "whether you're a...or a..."
|
||||
- Numbers and specifics over adjectives
|
||||
- Every claim tagged with {{VERIFY}} or {{SOURCE NEEDED}}
|
||||
- Tables are mandatory for comparisons -- never simulate with bullet points
|
||||
@@ -0,0 +1,190 @@
|
||||
# SEO-AGI
|
||||
|
||||
### One command. Competitive data in. Ranking pages out.
|
||||
|
||||
Most SEO tools tell you what's wrong with your site. This one writes the pages.
|
||||
|
||||
`/seoagi "airport parking JFK"` pulls the current SERP, analyzes what's ranking, finds the gaps in their content, and writes you a complete page -- with the heading structure, depth, FAQ section, and schema markup that actually competes. Not thin content. Not keyword-stuffed filler. Pages backed by live data from the tools the pros use.
|
||||
|
||||
**I built this because I got tired of the gap between "SEO audit" and "published page."** I've been doing SEO for 20+ years in ground transportation (1M+ bookings, 2M+ rides across my companies). The workflow was always the same: pull SERP data, analyze competitors, find gaps, write brief, write page, add schema, publish. Over and over. So I turned that entire workflow into a single skill that any AI agent can execute.
|
||||
|
||||
The result? I used this to research a competitor's best-performing pages, built equivalent content with `/seoagi`, bought the exact-match domains, and every single page is ranking on page 1. That's not theory. That's the workflow.
|
||||
|
||||
---
|
||||
|
||||
## What It Actually Does
|
||||
|
||||
```
|
||||
You: /seoagi "best project management tools 2026"
|
||||
|
||||
SEO-AGI:
|
||||
1. Pulls SERP top 10 via DataForSEO
|
||||
2. Parses competitor content (word count, headings, topics covered)
|
||||
3. Extracts People Also Ask questions
|
||||
4. Pulls related keywords with search volumes
|
||||
5. Detects search intent (informational vs commercial vs transactional)
|
||||
6. Generates a data-driven content brief
|
||||
7. Writes the complete page (Markdown + YAML frontmatter)
|
||||
8. Adds FAQ section from real PAA data
|
||||
9. Generates JSON-LD schema markup
|
||||
10. Validates against a quality checklist
|
||||
```
|
||||
|
||||
For rewrites, point it at any URL. It compares your page against the current top 3 ranking competitors, identifies exactly what you're missing, and rewrites with a change summary explaining every edit.
|
||||
|
||||
---
|
||||
|
||||
## The SEO Knowledge Inside
|
||||
|
||||
This isn't a wrapper around "write me an SEO article." The skill encodes strategies from the best in the game:
|
||||
|
||||
**Traditional SEO**
|
||||
- Intent-first content architecture (match what searchers actually want, not what you think the keyword means)
|
||||
- Competitive word count targeting (page length based on what's ranking, not arbitrary "write 2000 words")
|
||||
- Heading hierarchy derived from SERP analysis (not templates, not guesswork)
|
||||
- People Also Ask coverage as FAQ sections (answer the questions Google already knows people are asking)
|
||||
- Schema markup patterns by page type (FAQPage, LocalBusiness, HowTo, Product, BreadcrumbList)
|
||||
- Internal linking suggestions based on actual site data from GSC
|
||||
|
||||
**GEO / LLM SEO (Generative Engine Optimization)**
|
||||
- Content structured for AI citation (Perplexity, ChatGPT, Google AI Overviews)
|
||||
- Entity-rich writing that LLMs can extract and reference
|
||||
- Depth-over-length philosophy (comprehensive coverage that becomes the authoritative source)
|
||||
- FAQ patterns that match how AI systems parse and surface answers
|
||||
- Data-backed claims that AI systems prefer to cite over vague assertions
|
||||
|
||||
**The quality checklist every page runs through:**
|
||||
- Title tag <60 chars with target keyword? Check.
|
||||
- Meta description <155 chars with CTA? Check.
|
||||
- Single H1 mirroring the title? Check.
|
||||
- Logical H2 > H3 hierarchy? Check.
|
||||
- At least 3 PAA questions answered? Check.
|
||||
- Specific data/statistics (not vague claims)? Check.
|
||||
- JSON-LD schema appropriate to page type? Check.
|
||||
- Word count within competitive range? Check.
|
||||
- No keyword stuffing? Check.
|
||||
|
||||
Pages scoring below 80% get flagged with specific items to fix. Below 60% = rewrite.
|
||||
|
||||
---
|
||||
|
||||
## Data Integrations (BYOK)
|
||||
|
||||
Bring your own API keys. Use one, use all. The skill adapts:
|
||||
|
||||
| Integration | What It Provides | Required? |
|
||||
|---|---|---|
|
||||
| **DataForSEO** | Live SERP results, keyword volumes, People Also Ask, competitor content parsing | Yes (core) |
|
||||
| **Google Search Console** | Your actual query data, CTR, positions, cannibalization detection | Optional |
|
||||
| **Ahrefs** (via MCP) | Backlink profiles, domain authority, referring domains | Optional |
|
||||
| **SEMRush** (via MCP) | Traffic estimates, keyword gaps, competitive positioning | Optional |
|
||||
|
||||
No keys at all? The skill falls back to web search. You lose precision but the workflow still runs.
|
||||
|
||||
---
|
||||
|
||||
## Install
|
||||
|
||||
### Claude Code
|
||||
|
||||
```bash
|
||||
claude install-skill gbessoni/seo-agi
|
||||
```
|
||||
|
||||
### OpenClaw
|
||||
|
||||
Add to your `.claude/settings.json`:
|
||||
|
||||
```json
|
||||
{
|
||||
"skills": ["gbessoni/seo-agi"]
|
||||
}
|
||||
```
|
||||
|
||||
Or clone directly:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/gbessoni/seo-agi.git ~/.claude/skills/seo-agi
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Quick Start
|
||||
|
||||
```bash
|
||||
# 1. Clone
|
||||
git clone https://github.com/gbessoni/seo-agi.git
|
||||
cd seo-agi
|
||||
|
||||
# 2. Configure
|
||||
mkdir -p ~/.config/seo-agi
|
||||
cp .env.example ~/.config/seo-agi/.env
|
||||
# Edit with your API keys
|
||||
|
||||
# 3. Use with Claude Code
|
||||
# Just say: "write an SEO page for [keyword]"
|
||||
```
|
||||
|
||||
**Minimum keys needed:** DataForSEO (~$0.002/query). Everything else is optional but makes it stronger.
|
||||
|
||||
---
|
||||
|
||||
## Use Cases
|
||||
|
||||
**Exact-match domain play**: Research competitor's top pages with OpenClaw, generate equivalent content with `/seoagi`, buy the domains, publish. Page 1.
|
||||
|
||||
**Location page generation**: `/seoagi "plumber in [city]"` x 50 cities. Each page gets city-specific research, local PAA questions, LocalBusiness schema. Not cookie-cutter templates.
|
||||
|
||||
**Content refresh**: Point it at your underperforming URLs. It pulls your actual GSC data, compares against current top 3, and tells you exactly what to add, expand, or restructure. Then does it.
|
||||
|
||||
**Competitive intelligence**: `/seoagi research "competitor keyword"` gives you the full landscape without writing anything. Word count ranges, heading structures, topic gaps, related keywords with volumes.
|
||||
|
||||
**Brief handoff**: `/seoagi brief "keyword"` generates a structured content brief you can hand to a human writer. Research-backed, not vibes-based.
|
||||
|
||||
---
|
||||
|
||||
## The Workflow That Got Me Here
|
||||
|
||||
I've been running this workflow manually for 20 years across ParkingAccess.com (1M+ bookings) and Shuttlefare.com (2M+ rides). The pattern never changed:
|
||||
|
||||
1. Find what's ranking
|
||||
2. Figure out what they cover that you don't
|
||||
3. Write something deeper
|
||||
4. Add the technical SEO (schema, meta, structure)
|
||||
5. Publish and move on
|
||||
|
||||
seo-agi is that pattern, automated. The 20 years of pattern recognition compressed into a SKILL.md file, backed by live data APIs, running inside a super agent that can decompose and parallelize the work.
|
||||
|
||||
It's not AI replacing SEO expertise. It's SEO expertise finally having the right delivery mechanism.
|
||||
|
||||
---
|
||||
|
||||
## Testing
|
||||
|
||||
```bash
|
||||
# Unit tests (no API keys needed)
|
||||
python3 tests/test_env.py
|
||||
python3 tests/test_serp_analyze.py
|
||||
python3 tests/test_dataforseo.py
|
||||
|
||||
# Mock mode (uses fixture data, runs full pipeline)
|
||||
python3 scripts/research.py "airport parking JFK" --mock --output=compact
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Contributing
|
||||
|
||||
Open source, MIT license. PRs welcome.
|
||||
|
||||
The skill is modular. Want to add a new page template? Edit `references/page-templates.md`. New schema pattern? `references/schema-patterns.md`. Better quality checks? `references/quality-checklist.md`. New data source? Add a client in `scripts/lib/` and wire it into `research.py`.
|
||||
|
||||
---
|
||||
|
||||
## Credits
|
||||
|
||||
Built by [Greg Bessoni](https://github.com/gbessoni) ([@gregbessoni](https://x.com/gregbessoni)).
|
||||
|
||||
## License
|
||||
|
||||
MIT
|
||||
@@ -0,0 +1,491 @@
|
||||
---
|
||||
name: beefy-seo
|
||||
description: >
|
||||
Write SEO pages that rank in Google AND get cited by LLMs (ChatGPT, Perplexity, Claude).
|
||||
Use when creating airport parking pages, local service pages, listicles, comparison pages,
|
||||
pricing pages, or any content that must pass the Reddit Test -- meaning a knowledgeable
|
||||
practitioner would upvote it, not call it AI slop. Enforces information gain, 500-token
|
||||
chunk architecture, real HTML tables, verification tags, and honest "Not For You" sections.
|
||||
Triggers on: "write an SEO page", "beefy SEO", "seo page for [keyword]", "create a landing page",
|
||||
"rank for [keyword]", "rewrite this page for SEO", "optimize this content", "GEO", "AEO",
|
||||
"generative engine optimization", "seo-agi", "write a page that ranks".
|
||||
Do NOT trigger for pure technical SEO audits (crawl errors, robots.txt, sitemap validation).
|
||||
metadata:
|
||||
openclaw:
|
||||
emoji: "\U0001F969"
|
||||
tags:
|
||||
- seo
|
||||
- content
|
||||
- geo
|
||||
- aeo
|
||||
- llm-optimization
|
||||
---
|
||||
|
||||
# THE BEEFY SEO SKILL -- Generative Engine Optimization for AI Agents
|
||||
|
||||
You are an elite GEO (Generative Engine Optimization) and Technical SEO agent. Your directive is to generate high-fidelity, entity-rich, auditable content that ranks on Google AND gets cited by LLMs (ChatGPT, Perplexity, Gemini, Claude).
|
||||
|
||||
You do not write generic fluff. You write highly specific, practical, answer-forward content based on real operational data. You optimize for information gain, friction reduction, and immediate user extraction.
|
||||
|
||||
---
|
||||
|
||||
## 0. DATA LAYER -- COMPETITIVE INTELLIGENCE
|
||||
|
||||
Before writing anything, you gather real competitive data. This is what separates you from every other SEO prompt.
|
||||
|
||||
### Research Scripts
|
||||
|
||||
Run the research pipeline from this skill's directory:
|
||||
|
||||
```bash
|
||||
# Full competitive research (SERP + keywords + competitor content analysis)
|
||||
python3 {{SKILL_PATH}}/scripts/research.py "<keyword>" --output=brief
|
||||
|
||||
# Detailed JSON output for deep analysis
|
||||
python3 {{SKILL_PATH}}/scripts/research.py "<keyword>" --output=json
|
||||
|
||||
# Google Search Console data (if creds available)
|
||||
python3 {{SKILL_PATH}}/scripts/gsc_pull.py "<site_url>" --keyword="<keyword>"
|
||||
|
||||
# Cannibalization detection
|
||||
python3 {{SKILL_PATH}}/scripts/gsc_pull.py "<site_url>" --keyword="<keyword>" --cannibalization
|
||||
|
||||
# Mock mode for testing (no API keys needed)
|
||||
python3 {{SKILL_PATH}}/scripts/research.py "<keyword>" --mock --output=compact
|
||||
```
|
||||
|
||||
### API Key Configuration
|
||||
|
||||
Keys are loaded from `~/.config/seo-agi/.env` or environment variables:
|
||||
|
||||
```env
|
||||
DATAFORSEO_LOGIN=your_login
|
||||
DATAFORSEO_PASSWORD=your_password
|
||||
GSC_SERVICE_ACCOUNT_PATH=/path/to/service-account.json
|
||||
```
|
||||
|
||||
### MCP Tool Integration
|
||||
|
||||
If the user has Ahrefs or SEMRush MCP servers connected, use them to supplement or replace DataForSEO:
|
||||
|
||||
- **Ahrefs MCP**: `site-explorer-organic-keywords`, `site-explorer-metrics`, `keywords-explorer-overview`, `keywords-explorer-related-terms`, `serp-overview` for keyword data, SERP data, competitor metrics
|
||||
- **SEMRush MCP**: `keyword_research`, `organic_research`, `backlink_research` for keyword data, domain analytics
|
||||
- Use DataForSEO for **content parsing** (competitor page structure, headings, word counts) which MCP tools don't cover
|
||||
- When multiple sources are available, cross-reference for higher confidence
|
||||
|
||||
### Data Cascade (use in order of availability)
|
||||
|
||||
| Priority | Source | What It Provides |
|
||||
|----------|--------|-----------------|
|
||||
| 1 | DataForSEO | Live SERP, competitor content parsing, PAA, keyword volumes |
|
||||
| 2 | Ahrefs MCP | Keyword difficulty, DR, traffic estimates, backlink data |
|
||||
| 3 | SEMRush MCP | Keyword analytics, organic research, domain overview |
|
||||
| 4 | GSC | Owned query performance, CTR, position, cannibalization |
|
||||
| 5 | WebSearch | Fallback research when no API keys available |
|
||||
|
||||
### What the Research Gives You
|
||||
|
||||
The research script outputs:
|
||||
- **SERP data**: Top 10 organic results with URLs, titles, descriptions
|
||||
- **Competitor content**: Word counts, heading structures (H1/H2/H3), topics covered
|
||||
- **Related keywords**: With search volume and difficulty scores
|
||||
- **PAA questions**: People Also Ask questions for FAQ sections
|
||||
- **Analysis**: Search intent detection, word count stats (min/max/median/recommended range), topic frequency across competitors, heading patterns
|
||||
|
||||
**Use this data to inform every decision**: word count targets, heading structure, topics to cover, questions to answer, competitive gaps to exploit.
|
||||
|
||||
---
|
||||
|
||||
## 1. CORE BELIEF SYSTEM
|
||||
|
||||
1. **AI content is not the problem; generic content is.** Do not rewrite the first page of Google. Add genuinely useful, sourced, less-common information.
|
||||
2. **Write for LLM Retrieval.** The page must be easy to extract, summarize, cite, and quote by both search engines and AI answer engines.
|
||||
3. **Entity Consensus over Backlinks.** LLMs trust brands mentioned consistently across high-signal domains (Reddit, Wikipedia, LinkedIn, Medium). Build consensus across platforms, not just link equity.
|
||||
4. **Tables are Mandatory.** Use clean HTML `<table>` elements for cost, comparison, specs, and local services. Never simulate tables with bullet points.
|
||||
5. **Top-of-Page Dominance.** The most important, answer-forward material goes at the absolute top. A fast-scan summary block must appear within the first 200 words.
|
||||
6. **Brand > Links.** Google and LLMs prioritize "Brand + Keyword" searches. If ChatGPT doesn't know a website exists, a guest post there is worthless for GEO.
|
||||
|
||||
---
|
||||
|
||||
## 2. GOOGLE AI SEARCH -- 7 RANKING SIGNALS
|
||||
|
||||
Every piece of content is scored against these seven signals in Google's AI pipeline. Optimize for all seven.
|
||||
|
||||
| Signal | What It Measures | How to Optimize |
|
||||
|--------|-----------------|-----------------|
|
||||
| Base Ranking | Core algorithm relevance | Strong topical authority, clean technical SEO |
|
||||
| Gecko Score | Semantic/vector similarity (embeddings) | Cover semantic neighbors, synonyms, related entities, co-occurring concepts |
|
||||
| Jetstream | Advanced context/nuance understanding | Genuine analysis, honest comparisons, unique framing |
|
||||
| BM25 | Traditional keyword matching | Include exact-match terms, long-form entity names, high-volume synonyms |
|
||||
| PCTR | Predicted CTR from popularity/personalization | Compelling titles with numbers or power words, strong meta descriptions |
|
||||
| Freshness | Time-decay recency | "Last verified" dates, seasonal content, updated pricing |
|
||||
| Boost/Bury | Manual quality adjustments | Avoid thin sections, empty headings, duplicate content patterns |
|
||||
|
||||
---
|
||||
|
||||
## 3. THE 500-TOKEN CHUNK ARCHITECTURE
|
||||
|
||||
Google's AI retrieves content in ~500-token (~375 word) chunks. LLMs chunk at ~600 words with ~300 word overlap. Structure every page to feed this pipeline perfectly.
|
||||
|
||||
### Chunk Rules:
|
||||
- **Question-Based H2s:** Every H2 must match a real search query or a "Query Fan-Out" question (the logical follow-up an AI will suggest). Use PAA data from research to inform these.
|
||||
- **The Snippet Answer:** The first 2-3 sentences immediately following any H2 must be a direct, concrete answer to that heading. No preamble. No definitions.
|
||||
- **The Contrast Statement:** Within the chunk, include explicit X vs. Y comparisons with numbers (e.g., "Economy lots cost $16/day but require a 15-minute bus ride; terminal garages cost $43/day with direct skybridge access").
|
||||
- **Self-Contained Chunks:** Never split a data table across chunk boundaries. Never stack two H2s without at least 250 words of substantive data between them.
|
||||
- **Front-Load Strength:** The strongest content (bottom line, key recommendations) must appear in the first 3 chunks, not the last. AI retrieval may never reach buried material.
|
||||
|
||||
---
|
||||
|
||||
## 4. SEAT SIGNALS (Semantic + E-E-A-T + Entity/Knowledge Graph)
|
||||
|
||||
### Semantic Keywords
|
||||
Every page must cover:
|
||||
- Primary head terms (from research: target keyword)
|
||||
- Semantic neighbors (from research: related keywords and topic frequency data)
|
||||
- Geo-modifiers (neighborhoods, nearby cities, landmarks served)
|
||||
- Mode competitors (transit, taxi, Uber/Lyft, rideshare -- must be named even if you don't sell them)
|
||||
- Operational terms (from research: common heading topics across competitors)
|
||||
|
||||
### E-E-A-T Signals
|
||||
- **Experience:** Location-specific operational details (terminal pickup spots, timing, traffic)
|
||||
- **Expertise:** Pricing comparisons with real numbers, not vague "affordable" language
|
||||
- **Authority:** Cite official sources (airport authority, transit authority, published fare schedules)
|
||||
- **Trust:** Honest "Not For You" sections, transparent comparison against non-parking options
|
||||
|
||||
### Entity / Knowledge Graph
|
||||
Google's KG uses different NLP than transformers. Entity signals must be explicit:
|
||||
- Full official entity names at least once (e.g., "Hartsfield-Jackson Atlanta International Airport" not just "ATL")
|
||||
- Terminal numbers/names as distinct entities
|
||||
- Airline-to-terminal mappings where relevant
|
||||
- Parking lot names as entities, not just list items
|
||||
- Operating authority names (Port Authority, airport authority, etc.)
|
||||
|
||||
---
|
||||
|
||||
## 5. QUALITY & AUDIT FILTERS
|
||||
|
||||
Before completing any output, pass these tests. If the content fails, rewrite it.
|
||||
|
||||
### A. The Reddit Test
|
||||
If this page were posted to a relevant subreddit, would a knowledgeable practitioner call it "AI slop" or ask "Where is the real data?"
|
||||
|
||||
**Passing requires at least three of the following:**
|
||||
1. A hard number from an official or overlooked source (capacity, square footage, wait time, frequency, volume)
|
||||
2. A layout or navigation detail only someone familiar with the place would know
|
||||
3. A cost comparison that does real math (e.g., "5 days at $20/day = $100; an Uber round trip from downtown is roughly $30 total -- the break-even is about 2 days")
|
||||
4. A schedule or operational detail with specifics (shuttle runs every X minutes; lot fills by Y time on Z days)
|
||||
5. A "the thing they moved / changed / broke" detail -- something that changed recently
|
||||
6. A real gotcha or failure mode described with enough specificity that a reader thinks "that happened to me"
|
||||
|
||||
### B. The Prove-It Details
|
||||
At least **two** hard operational facts must be present in every document:
|
||||
- Capacity, frequency, fill rate, wait time, or distance measurements
|
||||
- Break-even cost math showing when one option beats another
|
||||
- Layout/navigation details that help someone who has never been there
|
||||
- A recent change not yet reflected on most competing pages
|
||||
|
||||
### C. The "Not For You" Block
|
||||
Every page must include a section honestly telling the reader when this option is a **bad fit**. Name the specific scenario. Include at least one line a competitor would never say because it might scare off a lead. This is the ultimate E-E-A-T trust signal.
|
||||
|
||||
### D. The Information Gain Test
|
||||
A page passes when it contains content that cannot be found by reading the top 10 Google results for the same query. Use the research data to identify what competitors cover, then find what they miss.
|
||||
|
||||
---
|
||||
|
||||
## 6. TECHNICAL MARKUP RULES
|
||||
|
||||
### The RDFa Hack
|
||||
LLMs often ignore JSON-LD in the header. Embed semantic data directly inline using RDFa or Microdata (`<span>` tags). This is "alt-text for your text" -- label entities, costs, and services explicitly within paragraph code so LLMs extract it effortlessly.
|
||||
|
||||
### Required Schema Per Page Type:
|
||||
- **FAQPage:** Wrap every question-based H2 + answer pair
|
||||
- **HowTo:** Any step-by-step booking or pickup process
|
||||
- **Product/Offer:** Pricing tables and service options
|
||||
- **LocalBusiness:** For facilities or lots listed
|
||||
- **BreadcrumbList:** Site navigation context
|
||||
|
||||
See `{{SKILL_PATH}}/references/schema-patterns.md` for JSON-LD templates.
|
||||
|
||||
### Schema Serves 3 Independent Functions:
|
||||
|
||||
| Function | What It Does | Why It Matters |
|
||||
|----------|-------------|----------------|
|
||||
| Searchable (recall) | Can AI find you? | FAQPage surfaces Q&A in rich results and AI Overviews |
|
||||
| Indexable (filtering) | How you rank in structured results | Product/Offer enables price/rating filtering |
|
||||
| Retrievable (citation) | What AI can directly quote or display | Tables, FAQ markup, HowTo steps become citable |
|
||||
|
||||
---
|
||||
|
||||
## 7. VERIFICATION & TAGGING SYSTEM
|
||||
|
||||
You are forbidden from inventing fake studies, statistics, or pricing. Use auditable tags for human editors.
|
||||
|
||||
| Tag | When to Use | Format |
|
||||
|-----|-------------|--------|
|
||||
| `{{VERIFY}}` | Any specific price, rate, capacity, schedule, distance, or operational claim | `{{VERIFY: Garage daily rate $20 \| County Parking Rates PDF}}` |
|
||||
| `{{RESEARCH NEEDED}}` | A section that needs hard data you could not find or confirm | `{{RESEARCH NEEDED: Garage total capacity \| check master plan PDF}}` |
|
||||
| `{{SOURCE NEEDED}}` | A claim that needs a traceable citation before publish | `{{SOURCE NEEDED: shuttle frequency \| check ground transportation page}}` |
|
||||
|
||||
### Source Citation Rules:
|
||||
**Do not cite vaguely.** Never write "official airport website" or "government data."
|
||||
|
||||
Instead cite specifically:
|
||||
- "Broward County Aviation Department -- FLL Parking Rates (broward.org/airport/parking)"
|
||||
- "FLL Airport Master Plan, 2024 update, Section 4.2"
|
||||
- "FDOT Traffic Count Station 0934, I-595 at US-1 interchange"
|
||||
|
||||
---
|
||||
|
||||
## 8. REQUIRED PAGE STRUCTURE
|
||||
|
||||
Use this structure unless the brief explicitly requires something else.
|
||||
|
||||
### 1. Title
|
||||
Clear, includes the main topic naturally, not overstuffed, promises a concrete outcome.
|
||||
|
||||
### 2. Opening Answer Block (first 100-150 words)
|
||||
Answer the main query directly. Explain what makes this page useful or different. Preview the most important distinctions.
|
||||
|
||||
### 3. Fast-Scan Summary (immediately after opening)
|
||||
One of: bullet summary (3-5 bullets max, each with a concrete fact), key takeaways box, comparison table, or quick decision matrix. **Not optional.** Every page needs a scannable extraction target near the top.
|
||||
|
||||
### 4. Main Body with Distinct Sections
|
||||
Every section must do one unique job: explain, compare, quantify, define, rank, warn, price, or instruct. No filler sections. Use research data to determine which sections competitors cover and where the gaps are.
|
||||
|
||||
### 5. Comparison Table
|
||||
Real HTML `<table>` with columns that do real work. Prefer: "Best For" (who should choose), "Main Tradeoff" (what you give up), "Why It Matters" (implication, not just fact), "Typical Cost" with `{{VERIFY}}` tags.
|
||||
|
||||
### 6. Prove-It Section (Information Gain)
|
||||
The material that passes the Reddit Test. At minimum two hard operational facts with traceable citations.
|
||||
|
||||
### 7. Not For You Block
|
||||
Specific scenarios where this is the wrong choice. At least one line a competitor would never publish.
|
||||
|
||||
### 8. Conclusion / Next Step
|
||||
Direct. Summarize the decision and next action. Do not restate the entire page.
|
||||
|
||||
---
|
||||
|
||||
## 9. ABSOLUTE WRITING RULES
|
||||
|
||||
### Never Do:
|
||||
- Generic intros or definitional preambles
|
||||
- "In today's fast-paced world" or any variant
|
||||
- "Whether you're a ... or a ..." constructions
|
||||
- The word "nestled"
|
||||
- Em dashes
|
||||
- Repetitive FAQ fluff
|
||||
- Bulleted lists pretending to be tables
|
||||
- Near-identical sections with only wording changes
|
||||
- Empty headings without content
|
||||
- Generic praise repeated across all items in a listicle
|
||||
- Keyword stuffing
|
||||
- Jump-link TOC patterns that create weak fragment URLs
|
||||
|
||||
### Always Do:
|
||||
- Short to medium sentences, concrete nouns, explicit comparisons
|
||||
- Numbers and specifics over adjectives
|
||||
- Entity-rich language (real product names, locations, service names)
|
||||
- Honest negative recommendations alongside positive ones
|
||||
- Front-load the strongest material
|
||||
|
||||
---
|
||||
|
||||
## 10. VERTICAL-SPECIFIC INSTRUCTIONS
|
||||
|
||||
### Airport / Parking / Transportation Pages
|
||||
1. Terminal-to-facility map or guide. List which airlines operate from which terminals and which parking option serves each best.
|
||||
2. Capacity or availability context. How many spaces? When does it fill? What happens when full?
|
||||
3. Rideshare/transit comparison math. Break-even calculation: at how many days does parking cost more than two Uber rides?
|
||||
4. Pickup/dropoff operational details. Where exactly is rideshare pickup? Cell phone lot? What confuses first-timers?
|
||||
5. Shuttle details. Frequency, hours, known reliability issues.
|
||||
6. Peak-day warning. Name specific days or events that cause fill-ups. Not "busy periods" -- "cruise ship Saturdays," "Thanksgiving Wednesday."
|
||||
|
||||
### Local Service Pages
|
||||
- City/area naturally in title and opening
|
||||
- Cost or pricing expectations with ranges
|
||||
- Practical comparison table (service type vs. cost, emergency vs. standard, residential vs. commercial)
|
||||
- Buyer questions people actually ask
|
||||
|
||||
### Listicles
|
||||
- Each item must be substantively different
|
||||
- Format per item: name, who it's best for, why it made the list, one differentiator, one tradeoff
|
||||
- Strongest items first. Do not pad to reach a number.
|
||||
- Segment large lists into logical groups (best budget, best premium, best for beginners)
|
||||
|
||||
### Comparison / Pricing Pages
|
||||
- Comparison table high on the page
|
||||
- Clear selection criteria and who each option is best for
|
||||
- Real tradeoffs, not just feature checkmarks
|
||||
- Simple recommendation logic
|
||||
|
||||
---
|
||||
|
||||
## 11. LLM / AEO CITATION STRATEGY
|
||||
|
||||
LLMs pull from positions 51-100, not just page 1. Being the most structured and honest comparison page can earn AI citations even without traditional page 1 rankings.
|
||||
|
||||
### To become citable by AI answer engines:
|
||||
- Have the most complete pricing comparison (all modes, all price points)
|
||||
- Have the clearest structured data (tables with labeled columns)
|
||||
- Have unique operational content (terminal pickup, process steps)
|
||||
- Be the page that tells the truth when competitors don't
|
||||
|
||||
### Entity Consensus Generation:
|
||||
When prompted for broader strategy, output variations of core 500-token chunks formatted for cross-posting on LinkedIn, Medium, Reddit, and Vocal Media to build brand authority where LLMs scrape.
|
||||
|
||||
---
|
||||
|
||||
## 12. HUB & SPOKE INTERNAL LINKING
|
||||
|
||||
- **Hub page** = main topic page (e.g., "ATL Airport Parking")
|
||||
- **Spoke pages** = detail pages, hotel pages, destination pages, supplier profiles, terminal guides
|
||||
- Every spoke links back to its hub
|
||||
- Hub links to its most important spokes
|
||||
- Dead-end content (flat lists with no links) wastes crawl equity
|
||||
- Use research data to identify which hub/spoke pages competitors link between
|
||||
|
||||
---
|
||||
|
||||
## 13. EXECUTION PROTOCOL
|
||||
|
||||
When the user provides a target keyword and brief:
|
||||
|
||||
1. **Research**: Run the data layer
|
||||
```bash
|
||||
python3 {{SKILL_PATH}}/scripts/research.py "<keyword>" --output=json
|
||||
```
|
||||
If DataForSEO creds are unavailable, use Ahrefs/SEMRush MCP tools, then WebSearch as fallback.
|
||||
Also search for official source pages, operational documents, recent changes, layout details, comparable cost math, and community feedback.
|
||||
|
||||
2. **Brief**: If the user did not provide a brief, build one:
|
||||
```
|
||||
Topic: [inferred from keyword]
|
||||
Primary Keyword: [target keyword]
|
||||
Search Intent: [from research: informational / commercial / local / comparison / transactional]
|
||||
Audience: [inferred]
|
||||
Geography: [if relevant]
|
||||
Page Type: [from research: service page / listicle / comparison / pricing / local page / guide]
|
||||
Vertical: [airport parking / local service / SaaS / medical / legal / etc.]
|
||||
Information Gain Target: [what should this page add that the top 10 do not?]
|
||||
Reddit Test Target: [which subreddit? what would a knowledgeable commenter expect?]
|
||||
Word Count Target: [from research: recommended_min to recommended_max]
|
||||
H2 Target: [from research: median H2 count]
|
||||
PAA Questions to Answer: [from research]
|
||||
```
|
||||
Confirm with user before writing unless they said "just write it."
|
||||
|
||||
3. **Write**: Front-load the fast-scan summary matrix in the first 200 words. Build 500-token chunks using the Snippet Answer rule. Integrate the "Not For You" block.
|
||||
|
||||
4. **Reddit Test**: If the content would get called "AI slop" on the relevant subreddit, rewrite before delivering.
|
||||
|
||||
5. **Tag**: Insert all `{{VERIFY}}`, `{{RESEARCH NEEDED}}`, and `{{SOURCE NEEDED}}` tags on every specific claim.
|
||||
|
||||
6. **Markup**: Output final markdown with clean `<table>` structures and JSON-LD schema.
|
||||
|
||||
7. **Quality Checklist**: Run the checklist (Section 14) before delivery. If any item fails, revise.
|
||||
|
||||
8. **Save**: Output to `~/Documents/SEO-AGI/pages/` (new pages) or `~/Documents/SEO-AGI/rewrites/` (rewrites).
|
||||
|
||||
### Rewrite Protocol
|
||||
|
||||
When rewriting an existing page:
|
||||
1. Fetch URL (WebFetch) or read local file
|
||||
2. Identify target keyword from title/H1 or ask user
|
||||
3. Run research against the keyword
|
||||
4. Run GSC data if available: `python3 {{SKILL_PATH}}/scripts/gsc_pull.py "<site_url>" --keyword="<keyword>"`
|
||||
5. Gap analysis: compare existing page vs research data. What's missing? What's thin? What fails the Reddit Test?
|
||||
6. Rewrite following gap report
|
||||
7. Output rewritten page + change summary (what changed and why)
|
||||
|
||||
### Batch Mode
|
||||
|
||||
For batch requests ("write 5 location pages for [service]"), decompose into parallel sub-agents:
|
||||
- **Research agent**: Run research per keyword variant
|
||||
- **GSC agent**: Pull performance data if creds available
|
||||
- **Writer agent**: Generate each page from its brief, following full execution protocol
|
||||
- **QA agent**: Run quality checklist on each page
|
||||
|
||||
---
|
||||
|
||||
## 14. QUALITY CHECKLIST
|
||||
|
||||
Run before every delivery. If any answer is NO, revise before delivering.
|
||||
|
||||
| Check | Required |
|
||||
|-------|----------|
|
||||
| Does the page contain information gain over the top 10 Google results? | YES |
|
||||
| Would a knowledgeable Reddit commenter upvote this? | YES |
|
||||
| Is the core answer in the first 150 words? | YES |
|
||||
| Is there a fast-scan summary within the first 200 words? | YES |
|
||||
| Are there 2+ hard operational Prove-It facts? | YES |
|
||||
| Is there at least one real HTML/Markdown table? | YES |
|
||||
| Is every section doing a unique job (no repetition)? | YES |
|
||||
| Are all specific numbers tagged with `{{VERIFY}}`? | YES |
|
||||
| Are all citations specific and traceable? | YES |
|
||||
| Is there a "Not For You" block? | YES |
|
||||
| Is the content structured for LLM extraction (500-token chunks)? | YES |
|
||||
| Does the page avoid all banned phrases and patterns? | YES |
|
||||
| Word count within competitive range (from research data)? | YES |
|
||||
| JSON-LD schema included and matches page type? | YES |
|
||||
| Title tag <60 chars with target keyword? | YES |
|
||||
| Meta description <155 chars with value prop? | YES |
|
||||
|
||||
---
|
||||
|
||||
## 15. OUTPUT FORMAT
|
||||
|
||||
All pages output as Markdown with YAML frontmatter:
|
||||
|
||||
```yaml
|
||||
---
|
||||
title: "Airport Parking at JFK: Rates, Lots & Shuttle Guide [2026]"
|
||||
meta_description: "Compare JFK airport parking from $8/day. Official lots, off-site savings, shuttle times, and tips for every terminal."
|
||||
target_keyword: "airport parking JFK"
|
||||
secondary_keywords: ["JFK long term parking", "cheap parking near JFK"]
|
||||
search_intent: "commercial"
|
||||
page_type: "service-location"
|
||||
schema_type: "FAQPage, LocalBusiness, BreadcrumbList"
|
||||
word_count: 2200
|
||||
reddit_test: "r/travel -- would pass: includes break-even math, terminal-specific tips, real pricing"
|
||||
information_gain: "EV charging availability, cell phone lot capacity, terminal 7 construction impact"
|
||||
created: "2026-03-18"
|
||||
research_file: "~/.local/share/seo-agi/research/airport-parking-jfk-20260318.json"
|
||||
---
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## PAGE BRIEF TEMPLATE
|
||||
|
||||
When the user provides a page assignment, gather or request:
|
||||
|
||||
```
|
||||
Topic: [target topic]
|
||||
Primary Keyword: [target keyword]
|
||||
Search Intent: [informational / commercial / local / comparison / transactional]
|
||||
Audience: [who is reading this]
|
||||
Geography: [location if relevant]
|
||||
Page Type: [service page / listicle / comparison / pricing / local page / guide]
|
||||
Vertical: [airport parking / local service / SaaS / medical / legal / etc.]
|
||||
Information Gain Target: [what should this page add that generic pages do not?]
|
||||
Reddit Test Target: [which subreddit? what would a knowledgeable commenter expect?]
|
||||
```
|
||||
|
||||
If the user provides only a keyword, infer the rest and confirm before writing.
|
||||
|
||||
---
|
||||
|
||||
## REFERENCE FILES
|
||||
|
||||
Load on demand when writing:
|
||||
- `{{SKILL_PATH}}/references/schema-patterns.md` -- JSON-LD templates by page type
|
||||
- `{{SKILL_PATH}}/references/page-templates.md` -- structural templates (supplement, not override, the 500-token chunk architecture)
|
||||
|
||||
## DEPENDENCIES
|
||||
|
||||
```bash
|
||||
pip install requests
|
||||
# For GSC (optional):
|
||||
pip install google-auth google-api-python-client
|
||||
```
|
||||
@@ -0,0 +1,199 @@
|
||||
# SEO-AGI Technical Specification
|
||||
|
||||
## Overview
|
||||
|
||||
seo-agi is a Claude Code skill that writes and rewrites SEO-optimized pages
|
||||
using live competitive data from DataForSEO and Google Search Console.
|
||||
It bridges the gap between "SEO audit tools" and "content generation" by
|
||||
making the data the input to the writing, not a separate workflow.
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
User prompt ("write a page for airport parking JFK")
|
||||
│
|
||||
▼
|
||||
SKILL.md (orchestrator)
|
||||
│
|
||||
├── scripts/research.py ← DataForSEO SERP + keyword data
|
||||
│ └── lib/dataforseo.py ← API client
|
||||
│ └── lib/serp_analyze.py ← Content extraction + gap analysis
|
||||
│ └── lib/paa.py ← People Also Ask extraction
|
||||
│
|
||||
├── scripts/gsc_pull.py ← Google Search Console data
|
||||
│ └── lib/gsc_client.py ← GSC API client
|
||||
│ └── lib/cannibalization.py ← Query overlap detection
|
||||
│
|
||||
├── scripts/setup.py ← First-run config + dependency install
|
||||
│
|
||||
└── references/
|
||||
├── page-templates.md ← Structural templates by page type
|
||||
├── schema-patterns.md ← JSON-LD patterns by content type
|
||||
└── quality-checklist.md ← Scoring rubric for output validation
|
||||
```
|
||||
|
||||
## Data Flow
|
||||
|
||||
### Write Flow
|
||||
|
||||
```
|
||||
1. RESEARCH
|
||||
Input: keyword string, optional location/language
|
||||
Action: DataForSEO SERP API → top 10 results
|
||||
DataForSEO Keywords API → related keywords + volumes
|
||||
DataForSEO PAA → People Also Ask questions
|
||||
(optional) GSC → existing pages for cannibalization check
|
||||
Output: research.json saved to ~/.local/share/seo-agi/research/
|
||||
|
||||
2. ANALYZE
|
||||
Input: research.json
|
||||
Action: Extract competitor content structure (headings, word count, topics)
|
||||
Identify content gaps (topics competitors cover that you'd miss)
|
||||
Score keyword difficulty vs. content depth required
|
||||
Detect search intent (informational, commercial, transactional, navigational)
|
||||
Output: analysis object passed to Claude context
|
||||
|
||||
3. BRIEF
|
||||
Input: analysis object
|
||||
Action: Claude generates content brief using analysis + page template
|
||||
Output: brief.md saved to ~/Documents/SEO-AGI/briefs/
|
||||
|
||||
4. WRITE
|
||||
Input: brief + analysis + research data
|
||||
Action: Claude writes full page following brief constraints
|
||||
Validates against quality checklist
|
||||
Generates schema markup
|
||||
Output: page.md saved to ~/Documents/SEO-AGI/pages/
|
||||
```
|
||||
|
||||
### Rewrite Flow
|
||||
|
||||
```
|
||||
1. INGEST
|
||||
Input: URL or file path
|
||||
Action: Fetch/read existing content
|
||||
Extract: title, meta, headings, word count, structure
|
||||
Identify target keyword (from title/H1 or user input)
|
||||
Output: existing_page object
|
||||
|
||||
2. RESEARCH (same as Write flow, using detected keyword)
|
||||
|
||||
3. GAP ANALYSIS
|
||||
Input: existing_page + research data
|
||||
Action: Compare existing page against top 3 SERP competitors
|
||||
Identify: missing sections, thin areas, outdated claims
|
||||
(optional) GSC: actual query performance, CTR, position trends
|
||||
Output: gap_report object
|
||||
|
||||
4. REWRITE
|
||||
Input: existing_page + gap_report + research data
|
||||
Action: Claude rewrites following gap report
|
||||
Produces change summary (what changed and why)
|
||||
Updates schema markup
|
||||
Output: rewritten_page.md + changes.md saved to ~/Documents/SEO-AGI/rewrites/
|
||||
```
|
||||
|
||||
## DataForSEO API Endpoints Used
|
||||
|
||||
| Endpoint | Purpose | Used In |
|
||||
|---|---|---|
|
||||
| `/v3/serp/google/organic/live/advanced` | Live SERP results | research.py |
|
||||
| `/v3/serp/google/organic/task_post` | Async SERP (batch) | research.py (future) |
|
||||
| `/v3/dataforseo_labs/google/related_keywords/live` | Related keywords | research.py |
|
||||
| `/v3/dataforseo_labs/google/keyword_suggestions/live` | Keyword ideas | research.py |
|
||||
| `/v3/serp/google/organic/live/advanced` (with PAA) | People Also Ask | research.py |
|
||||
| `/v3/on_page/content_parsing/live` | Competitor content | serp_analyze.py |
|
||||
|
||||
## Google Search Console API Usage
|
||||
|
||||
| Method | Purpose | Used In |
|
||||
|---|---|---|
|
||||
| `searchanalytics.query()` | Query performance data | gsc_pull.py |
|
||||
| `sitemaps.list()` | Indexed pages | gsc_pull.py (future) |
|
||||
|
||||
## Configuration
|
||||
|
||||
All config stored in `~/.config/seo-agi/`:
|
||||
|
||||
```
|
||||
~/.config/seo-agi/
|
||||
├── .env # API credentials
|
||||
├── config.json # User preferences (default location, language, site URL)
|
||||
└── sites.json # GSC verified sites cache
|
||||
```
|
||||
|
||||
### config.json schema
|
||||
|
||||
```json
|
||||
{
|
||||
"default_location": 2840,
|
||||
"default_language": "en",
|
||||
"default_site": "https://example.com",
|
||||
"serp_depth": 10,
|
||||
"save_research": true,
|
||||
"output_dir": "~/Documents/SEO-AGI"
|
||||
}
|
||||
```
|
||||
|
||||
## Output Schema
|
||||
|
||||
### research.json
|
||||
|
||||
```json
|
||||
{
|
||||
"keyword": "airport parking JFK",
|
||||
"timestamp": "2026-03-17T14:30:00Z",
|
||||
"location": 2840,
|
||||
"serp": [
|
||||
{
|
||||
"position": 1,
|
||||
"url": "https://...",
|
||||
"title": "...",
|
||||
"description": "...",
|
||||
"word_count": 2400,
|
||||
"headings": ["H1: ...", "H2: ...", "H2: ..."],
|
||||
"topics_covered": ["pricing", "shuttle service", "terminal maps"]
|
||||
}
|
||||
],
|
||||
"related_keywords": [
|
||||
{"keyword": "JFK long term parking", "volume": 4400, "difficulty": 45},
|
||||
{"keyword": "cheap parking near JFK", "volume": 2900, "difficulty": 38}
|
||||
],
|
||||
"paa_questions": [
|
||||
"How much does parking cost at JFK?",
|
||||
"Is there free parking at JFK airport?",
|
||||
"What is the cheapest way to park at JFK?"
|
||||
],
|
||||
"intent": "commercial",
|
||||
"avg_word_count": 1850,
|
||||
"content_gaps": ["shuttle frequency", "EV charging", "real-time availability"]
|
||||
}
|
||||
```
|
||||
|
||||
### page.md frontmatter
|
||||
|
||||
```yaml
|
||||
---
|
||||
title: "Airport Parking at JFK: Pricing, Lots & Shuttle Guide [2026]"
|
||||
meta_description: "Compare JFK airport parking options from $8/day. Official lots, off-site alternatives, shuttle times, and terminal-specific tips."
|
||||
target_keyword: "airport parking JFK"
|
||||
secondary_keywords: ["JFK long term parking", "cheap parking near JFK"]
|
||||
word_count: 2200
|
||||
page_type: "service-location"
|
||||
schema_type: "FAQPage, LocalBusiness"
|
||||
created: "2026-03-17"
|
||||
research_file: "~/.local/share/seo-agi/research/airport-parking-jfk-20260317.json"
|
||||
---
|
||||
```
|
||||
|
||||
## Testing
|
||||
|
||||
```bash
|
||||
# Run with mock fixtures (no API calls)
|
||||
python3 scripts/research.py "test keyword" --mock
|
||||
|
||||
# Run tests
|
||||
python3 -m pytest tests/
|
||||
```
|
||||
|
||||
Fixtures in `fixtures/` provide sample API responses for offline development.
|
||||
@@ -0,0 +1,12 @@
|
||||
[
|
||||
{"keyword": "JFK long term parking", "volume": 4400, "cpc": 2.10, "competition": 0.8, "difficulty": 45},
|
||||
{"keyword": "cheap parking near JFK", "volume": 2900, "cpc": 1.85, "competition": 0.7, "difficulty": 38},
|
||||
{"keyword": "JFK airport parking rates", "volume": 2200, "cpc": 1.95, "competition": 0.75, "difficulty": 42},
|
||||
{"keyword": "JFK parking terminal 4", "volume": 1800, "cpc": 1.50, "competition": 0.5, "difficulty": 30},
|
||||
{"keyword": "off site parking JFK", "volume": 1600, "cpc": 2.30, "competition": 0.85, "difficulty": 48},
|
||||
{"keyword": "JFK airport parking coupon", "volume": 1200, "cpc": 1.20, "competition": 0.6, "difficulty": 35},
|
||||
{"keyword": "park and fly JFK", "volume": 1100, "cpc": 2.00, "competition": 0.7, "difficulty": 40},
|
||||
{"keyword": "JFK parking shuttle", "volume": 900, "cpc": 1.10, "competition": 0.4, "difficulty": 28},
|
||||
{"keyword": "overnight parking JFK airport", "volume": 800, "cpc": 1.75, "competition": 0.65, "difficulty": 37},
|
||||
{"keyword": "JFK cell phone lot", "volume": 720, "cpc": 0.50, "competition": 0.2, "difficulty": 15}
|
||||
]
|
||||
@@ -0,0 +1,73 @@
|
||||
{
|
||||
"organic": [
|
||||
{
|
||||
"position": 1,
|
||||
"url": "https://example.com/airport-parking-jfk",
|
||||
"domain": "example.com",
|
||||
"title": "JFK Airport Parking: Rates, Lots & Tips [2026 Guide]",
|
||||
"description": "Compare JFK parking options from $8/day. Official lots, off-site alternatives, and shuttle info.",
|
||||
"word_count": 2400,
|
||||
"headings": [
|
||||
"H1: JFK Airport Parking Guide",
|
||||
"H2: JFK Parking Options",
|
||||
"H3: Official Terminal Parking",
|
||||
"H3: Long-Term Parking Lots",
|
||||
"H3: Off-Site Airport Parking",
|
||||
"H2: JFK Parking Rates & Pricing",
|
||||
"H2: How to Get to JFK from Parking",
|
||||
"H2: Tips for Saving on JFK Parking",
|
||||
"H2: EV Charging at JFK",
|
||||
"H2: Frequently Asked Questions"
|
||||
]
|
||||
},
|
||||
{
|
||||
"position": 2,
|
||||
"url": "https://competitor1.com/jfk-parking",
|
||||
"domain": "competitor1.com",
|
||||
"title": "JFK Parking - Compare & Save Up to 60%",
|
||||
"description": "Book JFK airport parking in advance. Free shuttles, covered options available.",
|
||||
"word_count": 1800,
|
||||
"headings": [
|
||||
"H1: JFK Airport Parking",
|
||||
"H2: Compare JFK Parking Lots",
|
||||
"H2: Parking Rates at JFK",
|
||||
"H3: Short-Term Rates",
|
||||
"H3: Long-Term Rates",
|
||||
"H2: Free Shuttle Service",
|
||||
"H2: How to Book",
|
||||
"H2: Cancellation Policy",
|
||||
"H2: FAQ"
|
||||
]
|
||||
},
|
||||
{
|
||||
"position": 3,
|
||||
"url": "https://competitor2.com/airports/jfk/parking",
|
||||
"domain": "competitor2.com",
|
||||
"title": "JFK Airport Parking: All Options, Rates, and Shuttle Info",
|
||||
"description": "Everything you need to know about parking at John F. Kennedy International Airport.",
|
||||
"word_count": 3100,
|
||||
"headings": [
|
||||
"H1: Complete Guide to JFK Airport Parking",
|
||||
"H2: Terminal Parking Map",
|
||||
"H2: Official JFK Parking Rates",
|
||||
"H2: Off-Airport Parking Near JFK",
|
||||
"H3: Newark vs JFK Parking Comparison",
|
||||
"H2: AirTrain Connection",
|
||||
"H2: Rideshare vs Parking Cost Comparison",
|
||||
"H2: Accessibility Parking",
|
||||
"H2: Real-Time Availability",
|
||||
"H2: Seasonal Pricing Changes",
|
||||
"H2: FAQ"
|
||||
]
|
||||
}
|
||||
],
|
||||
"paa": [
|
||||
"How much does parking cost at JFK?",
|
||||
"Is there free parking at JFK airport?",
|
||||
"What is the cheapest way to park at JFK?",
|
||||
"How far in advance should I book JFK parking?",
|
||||
"Does JFK have covered parking?"
|
||||
],
|
||||
"featured_snippet": null,
|
||||
"total_results": 12400000
|
||||
}
|
||||
@@ -0,0 +1,138 @@
|
||||
# Page Templates Reference
|
||||
|
||||
Each page type has a specific structure that performs well in search.
|
||||
The skill auto-detects page type from the keyword intent, but users can override.
|
||||
|
||||
## Service Page
|
||||
|
||||
**Triggers:** "[service] in [location]", "[service] company", "[service] near me"
|
||||
|
||||
**Structure:**
|
||||
```
|
||||
H1: [Service] in [Location] - [Value Prop]
|
||||
H2: What Is [Service] / How [Service] Works
|
||||
H2: [Service] Options / Types of [Service]
|
||||
H3: [Option 1]
|
||||
H3: [Option 2]
|
||||
H3: [Option 3]
|
||||
H2: Pricing / What Does [Service] Cost
|
||||
H2: How to Choose [a/the right] [Service Provider]
|
||||
H2: Why [Company/Location] for [Service]
|
||||
H2: Frequently Asked Questions
|
||||
H3: [PAA Question 1]
|
||||
H3: [PAA Question 2]
|
||||
H3: [PAA Question 3]
|
||||
```
|
||||
|
||||
**Schema:** LocalBusiness + FAQPage
|
||||
**Word count target:** 1500-2500
|
||||
**Must include:** Pricing (even ranges), specific location references, social proof
|
||||
|
||||
## Comparison Page
|
||||
|
||||
**Triggers:** "[X] vs [Y]", "best [category]", "[product] alternatives", "top [N] [category]"
|
||||
|
||||
**Structure:**
|
||||
```
|
||||
H1: [X] vs [Y]: [Differentiating Angle] ([Year])
|
||||
H2: Quick Comparison / TL;DR
|
||||
[Comparison table]
|
||||
H2: What Is [X]
|
||||
H3: Key Features
|
||||
H3: Pricing
|
||||
H3: Best For
|
||||
H2: What Is [Y]
|
||||
H3: Key Features
|
||||
H3: Pricing
|
||||
H3: Best For
|
||||
H2: [X] vs [Y]: [Criteria 1]
|
||||
H2: [X] vs [Y]: [Criteria 2]
|
||||
H2: [X] vs [Y]: [Criteria 3]
|
||||
H2: Which Should You Choose?
|
||||
H2: FAQ
|
||||
```
|
||||
|
||||
**Schema:** FAQPage (+ Product if applicable)
|
||||
**Word count target:** 2000-3500
|
||||
**Must include:** Comparison table near top, specific criteria-based sections, clear recommendation
|
||||
|
||||
## How-To / Guide
|
||||
|
||||
**Triggers:** "how to [action]", "[topic] guide", "[topic] tutorial"
|
||||
|
||||
**Structure:**
|
||||
```
|
||||
H1: How to [Action]: [Qualifier] Guide ([Year])
|
||||
H2: What You'll Need / Prerequisites
|
||||
H2: Step 1: [Action]
|
||||
H3: [Sub-step if complex]
|
||||
H2: Step 2: [Action]
|
||||
H2: Step 3: [Action]
|
||||
H2: [N] more steps...
|
||||
H2: Common Mistakes / Troubleshooting
|
||||
H2: Tips from [Experts/Experience]
|
||||
H2: FAQ
|
||||
```
|
||||
|
||||
**Schema:** HowTo + FAQPage
|
||||
**Word count target:** 1500-3000
|
||||
**Must include:** Numbered steps, specific examples, troubleshooting section
|
||||
|
||||
## Location Page
|
||||
|
||||
**Triggers:** "[service] [city]", "[business type] [neighborhood]", "[thing to do] in [location]"
|
||||
|
||||
**Structure:**
|
||||
```
|
||||
H1: [Service/Topic] in [City, State] - [Value Prop]
|
||||
H2: [Service] Options in [City]
|
||||
H3: [Option/Area 1]
|
||||
H3: [Option/Area 2]
|
||||
H2: [City]-Specific Information
|
||||
[Local details: regulations, seasonality, landmarks, neighborhoods]
|
||||
H2: Pricing in [City]
|
||||
H2: How to [Get Started/Book/Find] [Service] in [City]
|
||||
H2: Tips for [Service] in [City]
|
||||
H2: Nearby Alternatives
|
||||
[Adjacent cities/neighborhoods]
|
||||
H2: FAQ
|
||||
```
|
||||
|
||||
**Schema:** LocalBusiness + FAQPage + BreadcrumbList
|
||||
**Word count target:** 1200-2000
|
||||
**Must include:** Specific local details (not generic), nearby alternatives for internal linking, local pricing
|
||||
|
||||
## Product / Feature Page
|
||||
|
||||
**Triggers:** "[product] features", "[product] review", "[product] pricing"
|
||||
|
||||
**Structure:**
|
||||
```
|
||||
H1: [Product Name]: [Key Benefit] ([Year])
|
||||
H2: What Is [Product]
|
||||
H2: Key Features
|
||||
H3: [Feature 1] - [Benefit]
|
||||
H3: [Feature 2] - [Benefit]
|
||||
H3: [Feature 3] - [Benefit]
|
||||
H2: Pricing
|
||||
[Plans table or breakdown]
|
||||
H2: Pros and Cons
|
||||
H2: Who Is [Product] Best For
|
||||
H2: [Product] vs Alternatives
|
||||
H2: How to Get Started
|
||||
H2: FAQ
|
||||
```
|
||||
|
||||
**Schema:** Product + FAQPage + Review (if review angle)
|
||||
**Word count target:** 1500-2500
|
||||
**Must include:** Pricing specifics, honest pros/cons, clear "best for" segment
|
||||
|
||||
## General Quality Rules (All Templates)
|
||||
|
||||
1. Every H2 section should have at least 150 words of substantive content
|
||||
2. FAQ sections use the exact PAA questions (or close variants) as H3s
|
||||
3. Include at least one data point, stat, or specific example per major section
|
||||
4. Internal links should go in contextually relevant spots, not dumped at the bottom
|
||||
5. Schema markup matches the page type (see schema-patterns.md)
|
||||
6. Year in title only if the content is genuinely time-sensitive
|
||||
7. No filler paragraphs. If a section doesn't add value, cut it.
|
||||
@@ -0,0 +1,83 @@
|
||||
# Content Quality Checklist
|
||||
|
||||
SEO-AGI validates every page against this checklist before final output.
|
||||
Each item is scored pass/fail. Pages scoring below 80% get flagged with
|
||||
specific items to fix.
|
||||
|
||||
## Title Tag (10 points)
|
||||
|
||||
- [ ] Contains target keyword (or close variant)
|
||||
- [ ] Under 60 characters
|
||||
- [ ] Unique and compelling (not just "[Keyword] | [Brand]")
|
||||
- [ ] Includes a differentiating element (year, number, qualifier)
|
||||
- [ ] Not duplicating another page's title on the same site
|
||||
|
||||
## Meta Description (10 points)
|
||||
|
||||
- [ ] Under 155 characters
|
||||
- [ ] Contains target keyword naturally
|
||||
- [ ] Includes a call-to-action or value proposition
|
||||
- [ ] Not a copy of the first paragraph
|
||||
- [ ] Would make someone click vs competitors in the SERP
|
||||
|
||||
## Heading Structure (15 points)
|
||||
|
||||
- [ ] Exactly one H1
|
||||
- [ ] H1 closely matches or mirrors title tag
|
||||
- [ ] Logical H2 > H3 hierarchy (no skipped levels)
|
||||
- [ ] H2 count within competitive range (see analysis)
|
||||
- [ ] Headings are descriptive (not "Section 1" or "More Info")
|
||||
|
||||
## Content Depth (25 points)
|
||||
|
||||
- [ ] Word count within competitive range (not arbitrarily long or short)
|
||||
- [ ] Answers at least 3 People Also Ask questions
|
||||
- [ ] Includes specific data, statistics, or concrete examples
|
||||
- [ ] Covers topics that appear in 2+ competitor pages
|
||||
- [ ] No thin sections (every H2 has 150+ words of substance)
|
||||
|
||||
## Search Intent Match (15 points)
|
||||
|
||||
- [ ] Page type matches detected intent (informational/commercial/transactional)
|
||||
- [ ] Content format matches SERP expectations (list vs guide vs comparison)
|
||||
- [ ] Addresses the primary user need within the first 200 words
|
||||
- [ ] If commercial intent: includes pricing or comparison elements
|
||||
- [ ] If informational intent: includes step-by-step or explanatory depth
|
||||
|
||||
## Technical SEO (15 points)
|
||||
|
||||
- [ ] JSON-LD schema markup included and matches page type
|
||||
- [ ] Schema uses correct types (see schema-patterns.md)
|
||||
- [ ] Image alt text suggestions included
|
||||
- [ ] At least 2 internal link suggestions with context
|
||||
- [ ] No orphaned sections (every section connects to the page's topic)
|
||||
|
||||
## Readability (10 points)
|
||||
|
||||
- [ ] No keyword stuffing (target keyword appears naturally, not forced)
|
||||
- [ ] Paragraphs are scannable (no walls of text)
|
||||
- [ ] Uses formatting aids where appropriate (bold key terms, tables for comparisons)
|
||||
- [ ] Transitions between sections are logical
|
||||
- [ ] Reads like it was written by a subject matter expert, not a content mill
|
||||
|
||||
## Scoring
|
||||
|
||||
| Score | Rating | Action |
|
||||
|---|---|---|
|
||||
| 90-100 | Exceptional | Ship it |
|
||||
| 80-89 | Strong | Minor tweaks optional |
|
||||
| 70-79 | Acceptable | Fix flagged items before publishing |
|
||||
| 60-69 | Below standard | Significant revision needed |
|
||||
| <60 | Rewrite | Start over with revised brief |
|
||||
|
||||
## Red Flags (automatic fail)
|
||||
|
||||
These issues override the score and require fixing regardless:
|
||||
|
||||
- Duplicate title tag matching another page on the site
|
||||
- Missing H1 or multiple H1 tags
|
||||
- Zero data/statistics in the entire page
|
||||
- Word count more than 50% below competitive median
|
||||
- No FAQ or PAA coverage
|
||||
- Missing schema markup entirely
|
||||
- Keyword density above 3% (stuffing)
|
||||
@@ -0,0 +1,141 @@
|
||||
# Schema Markup Patterns
|
||||
|
||||
JSON-LD schema markup templates for each page type. The skill generates
|
||||
these at the end of each page as a code block.
|
||||
|
||||
## FAQPage (use on any page with an FAQ section)
|
||||
|
||||
```json
|
||||
{
|
||||
"@context": "https://schema.org",
|
||||
"@type": "FAQPage",
|
||||
"mainEntity": [
|
||||
{
|
||||
"@type": "Question",
|
||||
"name": "Question text here?",
|
||||
"acceptedAnswer": {
|
||||
"@type": "Answer",
|
||||
"text": "Answer text here."
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
## LocalBusiness (service + location pages)
|
||||
|
||||
```json
|
||||
{
|
||||
"@context": "https://schema.org",
|
||||
"@type": "LocalBusiness",
|
||||
"name": "Business Name",
|
||||
"description": "What the business does",
|
||||
"url": "https://example.com",
|
||||
"address": {
|
||||
"@type": "PostalAddress",
|
||||
"streetAddress": "123 Main St",
|
||||
"addressLocality": "City",
|
||||
"addressRegion": "ST",
|
||||
"postalCode": "12345",
|
||||
"addressCountry": "US"
|
||||
},
|
||||
"geo": {
|
||||
"@type": "GeoCoordinates",
|
||||
"latitude": 40.7128,
|
||||
"longitude": -74.0060
|
||||
},
|
||||
"priceRange": "$$"
|
||||
}
|
||||
```
|
||||
|
||||
## HowTo (guide/tutorial pages)
|
||||
|
||||
```json
|
||||
{
|
||||
"@context": "https://schema.org",
|
||||
"@type": "HowTo",
|
||||
"name": "How to [action]",
|
||||
"description": "Brief description",
|
||||
"step": [
|
||||
{
|
||||
"@type": "HowToStep",
|
||||
"name": "Step title",
|
||||
"text": "Step description"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
## Product (product/feature pages)
|
||||
|
||||
```json
|
||||
{
|
||||
"@context": "https://schema.org",
|
||||
"@type": "Product",
|
||||
"name": "Product Name",
|
||||
"description": "Product description",
|
||||
"brand": {
|
||||
"@type": "Brand",
|
||||
"name": "Brand Name"
|
||||
},
|
||||
"offers": {
|
||||
"@type": "AggregateOffer",
|
||||
"lowPrice": "9.99",
|
||||
"highPrice": "99.99",
|
||||
"priceCurrency": "USD"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## BreadcrumbList (all pages, especially location hierarchies)
|
||||
|
||||
```json
|
||||
{
|
||||
"@context": "https://schema.org",
|
||||
"@type": "BreadcrumbList",
|
||||
"itemListElement": [
|
||||
{
|
||||
"@type": "ListItem",
|
||||
"position": 1,
|
||||
"name": "Home",
|
||||
"item": "https://example.com/"
|
||||
},
|
||||
{
|
||||
"@type": "ListItem",
|
||||
"position": 2,
|
||||
"name": "Category",
|
||||
"item": "https://example.com/category/"
|
||||
},
|
||||
{
|
||||
"@type": "ListItem",
|
||||
"position": 3,
|
||||
"name": "Current Page"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
## Combination Patterns
|
||||
|
||||
Most pages should combine multiple schemas. Common combos:
|
||||
|
||||
- **Service page:** LocalBusiness + FAQPage + BreadcrumbList
|
||||
- **Comparison page:** FAQPage + BreadcrumbList
|
||||
- **How-to page:** HowTo + FAQPage + BreadcrumbList
|
||||
- **Location page:** LocalBusiness + FAQPage + BreadcrumbList
|
||||
- **Product page:** Product + FAQPage + BreadcrumbList
|
||||
|
||||
When combining, wrap in an array:
|
||||
|
||||
```json
|
||||
[
|
||||
{ "@context": "https://schema.org", "@type": "FAQPage", ... },
|
||||
{ "@context": "https://schema.org", "@type": "BreadcrumbList", ... }
|
||||
]
|
||||
```
|
||||
|
||||
## Validation
|
||||
|
||||
Always validate generated schema at:
|
||||
- https://search.google.com/test/rich-results
|
||||
- https://validator.schema.org/
|
||||
@@ -0,0 +1,96 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
SEO-AGI Google Search Console data puller.
|
||||
Retrieves query performance data and detects cannibalization.
|
||||
|
||||
Usage:
|
||||
python3 gsc_pull.py "<site_url>" [options]
|
||||
|
||||
Options:
|
||||
--keyword=KEYWORD Filter queries containing this keyword
|
||||
--days=N Lookback period (default: 90)
|
||||
--min-impressions=N Minimum impressions threshold (default: 10)
|
||||
--output=FORMAT Output: json|compact (default: compact)
|
||||
--cannibalization Run cannibalization detection for the keyword
|
||||
"""
|
||||
|
||||
import sys
|
||||
import os
|
||||
import json
|
||||
import argparse
|
||||
|
||||
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
||||
|
||||
from lib.env import get_credentials
|
||||
from lib.gsc_client import GSCClient
|
||||
|
||||
|
||||
def parse_args():
|
||||
parser = argparse.ArgumentParser(description="SEO-AGI GSC Pull")
|
||||
parser.add_argument("site_url", help="GSC site URL (e.g., https://example.com)")
|
||||
parser.add_argument("--keyword", default=None, help="Keyword filter")
|
||||
parser.add_argument("--days", type=int, default=90, help="Lookback days")
|
||||
parser.add_argument("--min-impressions", type=int, default=10, help="Min impressions")
|
||||
parser.add_argument("--output", choices=["json", "compact"], default="compact")
|
||||
parser.add_argument("--cannibalization", action="store_true", help="Detect cannibalization")
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def format_compact(data: list[dict], mode: str = "performance") -> str:
|
||||
lines = []
|
||||
|
||||
if mode == "cannibalization":
|
||||
lines.append("# Cannibalization Report")
|
||||
for item in data:
|
||||
lines.append(f"\nQuery: {item['query']} ({item['page_count']} pages, {item['total_impressions']} impressions)")
|
||||
for page in item["pages"]:
|
||||
lines.append(f" pos {page['position']}: {page['page']} ({page['clicks']} clicks, {page['ctr']}% CTR)")
|
||||
else:
|
||||
lines.append("# Query Performance")
|
||||
lines.append(f"{'Query':<40} {'Clicks':>7} {'Impr':>7} {'CTR':>7} {'Pos':>5} Page")
|
||||
lines.append("-" * 110)
|
||||
for row in data[:50]:
|
||||
lines.append(
|
||||
f"{row['query'][:39]:<40} {row['clicks']:>7} {row['impressions']:>7} "
|
||||
f"{row['ctr']:>6.1f}% {row['position']:>5.1f} {row['page'][:50]}"
|
||||
)
|
||||
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def main():
|
||||
args = parse_args()
|
||||
creds = get_credentials()
|
||||
|
||||
if not creds["has_gsc"]:
|
||||
print("ERROR: Google Search Console credentials not found.", file=sys.stderr)
|
||||
print("Add GSC_SERVICE_ACCOUNT_PATH to ~/.config/seo-agi/.env", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
client = GSCClient(credentials_path=creds["gsc_service_account_path"])
|
||||
|
||||
if args.cannibalization and args.keyword:
|
||||
data = client.detect_cannibalization(
|
||||
site_url=args.site_url,
|
||||
keyword=args.keyword,
|
||||
days=args.days,
|
||||
)
|
||||
if args.output == "json":
|
||||
print(json.dumps(data, indent=2))
|
||||
else:
|
||||
print(format_compact(data, mode="cannibalization"))
|
||||
else:
|
||||
data = client.query_performance(
|
||||
site_url=args.site_url,
|
||||
keyword=args.keyword,
|
||||
days=args.days,
|
||||
min_impressions=args.min_impressions,
|
||||
)
|
||||
if args.output == "json":
|
||||
print(json.dumps(data, indent=2))
|
||||
else:
|
||||
print(format_compact(data))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1 @@
|
||||
# seo-agi scripts library
|
||||
@@ -0,0 +1,244 @@
|
||||
"""
|
||||
DataForSEO API client for SEO-AGI.
|
||||
Handles SERP results, keyword data, People Also Ask, and content parsing.
|
||||
"""
|
||||
|
||||
import json
|
||||
import base64
|
||||
import urllib.request
|
||||
import urllib.error
|
||||
from typing import Optional
|
||||
|
||||
|
||||
class DataForSEOClient:
|
||||
"""Client for DataForSEO REST API v3."""
|
||||
|
||||
BASE_URL = "https://api.dataforseo.com/v3"
|
||||
|
||||
def __init__(self, login: str, password: str):
|
||||
self.login = login
|
||||
self.password = password
|
||||
self._auth_header = self._make_auth_header(login, password)
|
||||
|
||||
@staticmethod
|
||||
def _make_auth_header(login: str, password: str) -> str:
|
||||
token = base64.b64encode(f"{login}:{password}".encode()).decode()
|
||||
return f"Basic {token}"
|
||||
|
||||
def _request(self, endpoint: str, payload: list[dict]) -> dict:
|
||||
"""Make a POST request to DataForSEO API."""
|
||||
url = f"{self.BASE_URL}{endpoint}"
|
||||
data = json.dumps(payload).encode("utf-8")
|
||||
|
||||
req = urllib.request.Request(
|
||||
url,
|
||||
data=data,
|
||||
headers={
|
||||
"Authorization": self._auth_header,
|
||||
"Content-Type": "application/json",
|
||||
},
|
||||
method="POST",
|
||||
)
|
||||
|
||||
try:
|
||||
with urllib.request.urlopen(req, timeout=30) as resp:
|
||||
return json.loads(resp.read().decode())
|
||||
except urllib.error.HTTPError as e:
|
||||
body = e.read().decode() if e.fp else ""
|
||||
raise RuntimeError(
|
||||
f"DataForSEO API error {e.code}: {body}"
|
||||
) from e
|
||||
except urllib.error.URLError as e:
|
||||
raise RuntimeError(
|
||||
f"DataForSEO connection error: {e.reason}"
|
||||
) from e
|
||||
|
||||
def serp_live(
|
||||
self,
|
||||
keyword: str,
|
||||
location_code: int = 2840,
|
||||
language_code: str = "en",
|
||||
depth: int = 10,
|
||||
) -> dict:
|
||||
"""
|
||||
Get live SERP results for a keyword.
|
||||
Returns organic results with position, URL, title, description.
|
||||
"""
|
||||
payload = [
|
||||
{
|
||||
"keyword": keyword,
|
||||
"location_code": location_code,
|
||||
"language_code": language_code,
|
||||
"depth": depth,
|
||||
"se_type": "organic",
|
||||
}
|
||||
]
|
||||
result = self._request(
|
||||
"/serp/google/organic/live/advanced", payload
|
||||
)
|
||||
return self._extract_serp(result)
|
||||
|
||||
def related_keywords(
|
||||
self,
|
||||
keyword: str,
|
||||
location_code: int = 2840,
|
||||
language_code: str = "en",
|
||||
limit: int = 30,
|
||||
) -> list[dict]:
|
||||
"""Get related keywords with search volume and difficulty."""
|
||||
payload = [
|
||||
{
|
||||
"keyword": keyword,
|
||||
"location_code": location_code,
|
||||
"language_code": language_code,
|
||||
"limit": limit,
|
||||
}
|
||||
]
|
||||
result = self._request(
|
||||
"/dataforseo_labs/google/related_keywords/live", payload
|
||||
)
|
||||
return self._extract_keywords(result)
|
||||
|
||||
def keyword_suggestions(
|
||||
self,
|
||||
keyword: str,
|
||||
location_code: int = 2840,
|
||||
language_code: str = "en",
|
||||
limit: int = 30,
|
||||
) -> list[dict]:
|
||||
"""Get keyword suggestions (broader ideation)."""
|
||||
payload = [
|
||||
{
|
||||
"keyword": keyword,
|
||||
"location_code": location_code,
|
||||
"language_code": language_code,
|
||||
"limit": limit,
|
||||
}
|
||||
]
|
||||
result = self._request(
|
||||
"/dataforseo_labs/google/keyword_suggestions/live", payload
|
||||
)
|
||||
return self._extract_keywords(result)
|
||||
|
||||
def content_parse(self, url: str) -> Optional[dict]:
|
||||
"""Parse content from a URL (headings, word count, structure)."""
|
||||
payload = [{"url": url}]
|
||||
try:
|
||||
result = self._request(
|
||||
"/on_page/content_parsing/live", payload
|
||||
)
|
||||
return self._extract_content(result)
|
||||
except RuntimeError:
|
||||
return None
|
||||
|
||||
def _extract_serp(self, raw: dict) -> dict:
|
||||
"""Extract clean SERP data from API response."""
|
||||
tasks = raw.get("tasks", [])
|
||||
if not tasks:
|
||||
return {"organic": [], "paa": [], "featured_snippet": None}
|
||||
|
||||
result = tasks[0].get("result", [])
|
||||
if not result:
|
||||
return {"organic": [], "paa": [], "featured_snippet": None}
|
||||
|
||||
items = result[0].get("items", [])
|
||||
|
||||
organic = []
|
||||
paa_questions = []
|
||||
featured_snippet = None
|
||||
|
||||
for item in items:
|
||||
item_type = item.get("type", "")
|
||||
|
||||
if item_type == "organic":
|
||||
organic.append(
|
||||
{
|
||||
"position": item.get("rank_absolute", 0),
|
||||
"url": item.get("url", ""),
|
||||
"domain": item.get("domain", ""),
|
||||
"title": item.get("title", ""),
|
||||
"description": item.get("description", ""),
|
||||
}
|
||||
)
|
||||
|
||||
elif item_type == "people_also_ask":
|
||||
for paa_item in item.get("items", []):
|
||||
q = paa_item.get("title", "")
|
||||
if q:
|
||||
paa_questions.append(q)
|
||||
|
||||
elif item_type == "featured_snippet":
|
||||
featured_snippet = {
|
||||
"url": item.get("url", ""),
|
||||
"title": item.get("title", ""),
|
||||
"description": item.get("description", ""),
|
||||
}
|
||||
|
||||
return {
|
||||
"organic": organic,
|
||||
"paa": paa_questions,
|
||||
"featured_snippet": featured_snippet,
|
||||
"total_results": result[0].get("se_results_count", 0),
|
||||
}
|
||||
|
||||
def _extract_keywords(self, raw: dict) -> list[dict]:
|
||||
"""Extract keyword data from labs API response."""
|
||||
tasks = raw.get("tasks", [])
|
||||
if not tasks:
|
||||
return []
|
||||
|
||||
result = tasks[0].get("result", [])
|
||||
if not result:
|
||||
return []
|
||||
|
||||
items = result[0].get("items", [])
|
||||
keywords = []
|
||||
|
||||
for item in items:
|
||||
kw_data = item.get("keyword_data", item)
|
||||
keyword_info = kw_data.get("keyword_info", {})
|
||||
keywords.append(
|
||||
{
|
||||
"keyword": kw_data.get("keyword", ""),
|
||||
"volume": keyword_info.get("search_volume", 0),
|
||||
"cpc": keyword_info.get("cpc", 0),
|
||||
"competition": keyword_info.get("competition", 0),
|
||||
"difficulty": kw_data.get(
|
||||
"keyword_properties", {}
|
||||
).get("keyword_difficulty", 0),
|
||||
}
|
||||
)
|
||||
|
||||
return sorted(keywords, key=lambda x: x["volume"], reverse=True)
|
||||
|
||||
def _extract_content(self, raw: dict) -> Optional[dict]:
|
||||
"""Extract content structure from on-page parsing."""
|
||||
tasks = raw.get("tasks", [])
|
||||
if not tasks:
|
||||
return None
|
||||
|
||||
result = tasks[0].get("result", [])
|
||||
if not result:
|
||||
return None
|
||||
|
||||
items = result[0].get("items", [])
|
||||
if not items:
|
||||
return None
|
||||
|
||||
page = items[0].get("page_content", {})
|
||||
|
||||
return {
|
||||
"title": page.get("header", {}).get("title", ""),
|
||||
"word_count": page.get("plain_text_word_count", 0),
|
||||
"headings": self._extract_headings(page),
|
||||
"plain_text_size": page.get("plain_text_size", 0),
|
||||
}
|
||||
|
||||
@staticmethod
|
||||
def _extract_headings(page_content: dict) -> list[str]:
|
||||
"""Pull heading tags from parsed content."""
|
||||
headings = []
|
||||
for level in ["h1", "h2", "h3"]:
|
||||
for heading in page_content.get(level, []):
|
||||
headings.append(f"{level.upper()}: {heading}")
|
||||
return headings
|
||||
@@ -0,0 +1,147 @@
|
||||
"""
|
||||
seo-agi environment and configuration loader.
|
||||
Reads API keys from ~/.config/seo-agi/.env or os.environ.
|
||||
Resolves paths relative to the skill installation directory.
|
||||
"""
|
||||
|
||||
import os
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
# Skill root is two levels up from this file (scripts/lib/env.py -> .)
|
||||
SKILL_DIR = Path(__file__).resolve().parent.parent.parent
|
||||
|
||||
OUTPUT_DIR = Path.home() / "Documents" / "SEO-AGI"
|
||||
DATA_DIR = Path.home() / ".local" / "share" / "seo-agi"
|
||||
CONFIG_DIR = Path.home() / ".config" / "seo-agi"
|
||||
ENV_FILE = CONFIG_DIR / ".env"
|
||||
|
||||
DEFAULT_CONFIG = {
|
||||
"default_location": 2840,
|
||||
"default_language": "en",
|
||||
"default_site": "",
|
||||
"serp_depth": 10,
|
||||
"save_research": True,
|
||||
"output_dir": str(OUTPUT_DIR),
|
||||
}
|
||||
|
||||
|
||||
def load_env() -> dict:
|
||||
"""
|
||||
Load environment variables.
|
||||
Reads from ~/.config/seo-agi/.env first, then overlays os.environ.
|
||||
"""
|
||||
env = {}
|
||||
|
||||
# First, try the config file
|
||||
if ENV_FILE.exists():
|
||||
with open(ENV_FILE, "r") as f:
|
||||
for line in f:
|
||||
line = line.strip()
|
||||
if not line or line.startswith("#"):
|
||||
continue
|
||||
if "=" in line:
|
||||
key, _, value = line.partition("=")
|
||||
key = key.strip()
|
||||
value = value.strip().strip('"').strip("'")
|
||||
if value:
|
||||
env[key] = value
|
||||
|
||||
# Then overlay with os.environ
|
||||
for key in [
|
||||
"DATAFORSEO_LOGIN", "DATAFORSEO_PASSWORD",
|
||||
"GSC_SERVICE_ACCOUNT_PATH", "GSC_CLIENT_ID",
|
||||
"GSC_CLIENT_SECRET", "GSC_REFRESH_TOKEN",
|
||||
"AHREFS_API_KEY", "SEMRUSH_API_KEY",
|
||||
]:
|
||||
val = os.environ.get(key)
|
||||
if val:
|
||||
env[key] = val
|
||||
|
||||
return env
|
||||
|
||||
|
||||
def load_config() -> dict:
|
||||
"""Load user config, merged with defaults."""
|
||||
config = DEFAULT_CONFIG.copy()
|
||||
config_file = CONFIG_DIR / "config.json"
|
||||
if config_file.exists():
|
||||
try:
|
||||
with open(config_file, "r") as f:
|
||||
user_config = json.load(f)
|
||||
config.update(user_config)
|
||||
except (json.JSONDecodeError, IOError):
|
||||
pass
|
||||
return config
|
||||
|
||||
|
||||
def get_credentials() -> dict:
|
||||
"""Get API credentials with availability flags."""
|
||||
env = load_env()
|
||||
|
||||
creds = {
|
||||
"dataforseo_login": env.get("DATAFORSEO_LOGIN", ""),
|
||||
"dataforseo_password": env.get("DATAFORSEO_PASSWORD", ""),
|
||||
"gsc_service_account_path": env.get("GSC_SERVICE_ACCOUNT_PATH", ""),
|
||||
"gsc_client_id": env.get("GSC_CLIENT_ID", ""),
|
||||
"gsc_client_secret": env.get("GSC_CLIENT_SECRET", ""),
|
||||
"gsc_refresh_token": env.get("GSC_REFRESH_TOKEN", ""),
|
||||
"ahrefs_api_key": env.get("AHREFS_API_KEY", ""),
|
||||
"semrush_api_key": env.get("SEMRUSH_API_KEY", ""),
|
||||
}
|
||||
|
||||
creds["has_dataforseo"] = bool(
|
||||
creds["dataforseo_login"] and creds["dataforseo_password"]
|
||||
)
|
||||
creds["has_gsc"] = bool(
|
||||
creds["gsc_service_account_path"]
|
||||
or (creds["gsc_client_id"] and creds["gsc_client_secret"])
|
||||
)
|
||||
creds["has_ahrefs"] = bool(creds["ahrefs_api_key"])
|
||||
creds["has_semrush"] = bool(creds["semrush_api_key"])
|
||||
|
||||
return creds
|
||||
|
||||
|
||||
def ensure_dirs():
|
||||
"""Create output directories if they don't exist."""
|
||||
config = load_config()
|
||||
output_dir = Path(config["output_dir"]).expanduser()
|
||||
|
||||
for subdir in ["research", "briefs", "pages", "rewrites"]:
|
||||
(output_dir / subdir).mkdir(parents=True, exist_ok=True)
|
||||
|
||||
(DATA_DIR / "research").mkdir(parents=True, exist_ok=True)
|
||||
(DATA_DIR / "cache").mkdir(parents=True, exist_ok=True)
|
||||
|
||||
|
||||
def check_setup() -> dict:
|
||||
"""Check setup status and return a summary."""
|
||||
creds = get_credentials()
|
||||
config = load_config()
|
||||
|
||||
return {
|
||||
"runtime": "claude-code",
|
||||
"skill_dir": str(SKILL_DIR),
|
||||
"config_dir_exists": CONFIG_DIR.exists(),
|
||||
"env_file_exists": ENV_FILE.exists(),
|
||||
"has_dataforseo": creds["has_dataforseo"],
|
||||
"has_gsc": creds["has_gsc"],
|
||||
"has_ahrefs": creds["has_ahrefs"],
|
||||
"has_semrush": creds["has_semrush"],
|
||||
"default_location": config["default_location"],
|
||||
"default_language": config["default_language"],
|
||||
"mode": _determine_mode(creds),
|
||||
}
|
||||
|
||||
|
||||
def _determine_mode(creds: dict) -> str:
|
||||
"""Determine operational mode based on available credentials."""
|
||||
if creds["has_dataforseo"] and creds["has_gsc"]:
|
||||
return "full"
|
||||
elif creds["has_dataforseo"]:
|
||||
return "dataforseo-only"
|
||||
elif creds["has_gsc"]:
|
||||
return "gsc-only"
|
||||
else:
|
||||
return "fallback"
|
||||
@@ -0,0 +1,179 @@
|
||||
"""
|
||||
Google Search Console API client for SEO-AGI.
|
||||
Pulls query performance data, cannibalization detection, and indexing status.
|
||||
"""
|
||||
|
||||
import json
|
||||
from typing import Optional
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
class GSCClient:
|
||||
"""Client for Google Search Console API."""
|
||||
|
||||
def __init__(self, credentials_path: str = None, oauth_creds: dict = None):
|
||||
"""
|
||||
Initialize GSC client.
|
||||
|
||||
Args:
|
||||
credentials_path: Path to service account JSON file
|
||||
oauth_creds: Dict with client_id, client_secret, refresh_token
|
||||
"""
|
||||
self.credentials_path = credentials_path
|
||||
self.oauth_creds = oauth_creds
|
||||
self._service = None
|
||||
|
||||
def _get_service(self):
|
||||
"""Lazy-initialize the GSC API service."""
|
||||
if self._service is not None:
|
||||
return self._service
|
||||
|
||||
try:
|
||||
from google.oauth2 import service_account
|
||||
from googleapiclient.discovery import build
|
||||
except ImportError:
|
||||
raise RuntimeError(
|
||||
"GSC requires google-auth and google-api-python-client. "
|
||||
"Install with: pip install google-auth google-api-python-client"
|
||||
)
|
||||
|
||||
if self.credentials_path:
|
||||
creds = service_account.Credentials.from_service_account_file(
|
||||
self.credentials_path,
|
||||
scopes=["https://www.googleapis.com/auth/webmasters.readonly"],
|
||||
)
|
||||
else:
|
||||
raise RuntimeError(
|
||||
"OAuth2 flow not yet implemented. Use a service account."
|
||||
)
|
||||
|
||||
self._service = build("searchconsole", "v1", credentials=creds)
|
||||
return self._service
|
||||
|
||||
def query_performance(
|
||||
self,
|
||||
site_url: str,
|
||||
keyword: str = None,
|
||||
days: int = 90,
|
||||
min_impressions: int = 10,
|
||||
row_limit: int = 100,
|
||||
) -> list[dict]:
|
||||
"""
|
||||
Pull query performance data from GSC.
|
||||
|
||||
Args:
|
||||
site_url: The GSC property URL (e.g., "https://example.com")
|
||||
keyword: Optional keyword filter (partial match)
|
||||
days: Lookback period in days
|
||||
min_impressions: Minimum impressions threshold
|
||||
row_limit: Max rows to return
|
||||
|
||||
Returns:
|
||||
List of dicts with query, page, clicks, impressions, ctr, position
|
||||
"""
|
||||
from datetime import datetime, timedelta
|
||||
|
||||
service = self._get_service()
|
||||
|
||||
end_date = datetime.now().strftime("%Y-%m-%d")
|
||||
start_date = (datetime.now() - timedelta(days=days)).strftime("%Y-%m-%d")
|
||||
|
||||
request_body = {
|
||||
"startDate": start_date,
|
||||
"endDate": end_date,
|
||||
"dimensions": ["query", "page"],
|
||||
"rowLimit": row_limit,
|
||||
"dimensionFilterGroups": [],
|
||||
}
|
||||
|
||||
if keyword:
|
||||
request_body["dimensionFilterGroups"].append(
|
||||
{
|
||||
"filters": [
|
||||
{
|
||||
"dimension": "query",
|
||||
"operator": "contains",
|
||||
"expression": keyword,
|
||||
}
|
||||
]
|
||||
}
|
||||
)
|
||||
|
||||
response = (
|
||||
service.searchanalytics()
|
||||
.query(siteUrl=site_url, body=request_body)
|
||||
.execute()
|
||||
)
|
||||
|
||||
rows = response.get("rows", [])
|
||||
results = []
|
||||
|
||||
for row in rows:
|
||||
impressions = row.get("impressions", 0)
|
||||
if impressions < min_impressions:
|
||||
continue
|
||||
|
||||
results.append(
|
||||
{
|
||||
"query": row["keys"][0],
|
||||
"page": row["keys"][1],
|
||||
"clicks": row.get("clicks", 0),
|
||||
"impressions": impressions,
|
||||
"ctr": round(row.get("ctr", 0) * 100, 2),
|
||||
"position": round(row.get("position", 0), 1),
|
||||
}
|
||||
)
|
||||
|
||||
return sorted(results, key=lambda x: x["impressions"], reverse=True)
|
||||
|
||||
def detect_cannibalization(
|
||||
self,
|
||||
site_url: str,
|
||||
keyword: str,
|
||||
days: int = 90,
|
||||
) -> list[dict]:
|
||||
"""
|
||||
Detect keyword cannibalization: multiple pages ranking for same query.
|
||||
|
||||
Returns:
|
||||
List of queries where 2+ pages from the site appear,
|
||||
sorted by total impressions.
|
||||
"""
|
||||
results = self.query_performance(
|
||||
site_url=site_url,
|
||||
keyword=keyword,
|
||||
days=days,
|
||||
min_impressions=5,
|
||||
row_limit=500,
|
||||
)
|
||||
|
||||
# Group by query
|
||||
query_pages = {}
|
||||
for row in results:
|
||||
q = row["query"]
|
||||
if q not in query_pages:
|
||||
query_pages[q] = []
|
||||
query_pages[q].append(row)
|
||||
|
||||
# Find queries with multiple pages
|
||||
cannibalized = []
|
||||
for query, pages in query_pages.items():
|
||||
if len(pages) > 1:
|
||||
cannibalized.append(
|
||||
{
|
||||
"query": query,
|
||||
"page_count": len(pages),
|
||||
"pages": sorted(
|
||||
pages, key=lambda x: x["position"]
|
||||
),
|
||||
"total_impressions": sum(
|
||||
p["impressions"] for p in pages
|
||||
),
|
||||
}
|
||||
)
|
||||
|
||||
return sorted(
|
||||
cannibalized,
|
||||
key=lambda x: x["total_impressions"],
|
||||
reverse=True,
|
||||
)
|
||||
@@ -0,0 +1,261 @@
|
||||
"""
|
||||
SERP content analyzer for SEO-AGI.
|
||||
Analyzes competitor content structure and identifies content gaps.
|
||||
"""
|
||||
|
||||
import statistics
|
||||
from typing import Optional
|
||||
|
||||
|
||||
def analyze_serp(
|
||||
serp_data: dict,
|
||||
content_data: list[Optional[dict]],
|
||||
target_keyword: str,
|
||||
) -> dict:
|
||||
"""
|
||||
Analyze SERP results + parsed competitor content.
|
||||
|
||||
Args:
|
||||
serp_data: Output from DataForSEOClient.serp_live()
|
||||
content_data: List of DataForSEOClient.content_parse() results
|
||||
(one per organic result, may contain None for failed parses)
|
||||
target_keyword: The keyword being targeted
|
||||
|
||||
Returns:
|
||||
Analysis dict with intent, word count stats, topic coverage, gaps.
|
||||
"""
|
||||
organic = serp_data.get("organic", [])
|
||||
paa = serp_data.get("paa", [])
|
||||
|
||||
# Word count analysis
|
||||
word_counts = []
|
||||
all_headings = []
|
||||
all_topics = set()
|
||||
|
||||
for i, content in enumerate(content_data):
|
||||
if content is None:
|
||||
continue
|
||||
|
||||
wc = content.get("word_count", 0)
|
||||
if wc > 0:
|
||||
word_counts.append(wc)
|
||||
|
||||
headings = content.get("headings", [])
|
||||
all_headings.extend(headings)
|
||||
|
||||
# Extract topics from H2/H3 headings
|
||||
for h in headings:
|
||||
if h.startswith(("H2:", "H3:")):
|
||||
topic = h.split(":", 1)[1].strip().lower()
|
||||
all_topics.add(topic)
|
||||
|
||||
# Word count statistics
|
||||
wc_stats = {}
|
||||
if word_counts:
|
||||
wc_stats = {
|
||||
"min": min(word_counts),
|
||||
"max": max(word_counts),
|
||||
"median": int(statistics.median(word_counts)),
|
||||
"mean": int(statistics.mean(word_counts)),
|
||||
"recommended_min": int(statistics.median(word_counts) * 0.8),
|
||||
"recommended_max": int(statistics.median(word_counts) * 1.3),
|
||||
"count_analyzed": len(word_counts),
|
||||
}
|
||||
|
||||
# Intent detection
|
||||
intent = detect_intent(target_keyword, organic, serp_data)
|
||||
|
||||
# Topic frequency (which topics appear in multiple competitors)
|
||||
topic_freq = _count_topic_frequency(content_data)
|
||||
|
||||
# Heading patterns
|
||||
heading_patterns = _analyze_heading_patterns(content_data)
|
||||
|
||||
return {
|
||||
"keyword": target_keyword,
|
||||
"intent": intent,
|
||||
"word_count_stats": wc_stats,
|
||||
"paa_questions": paa,
|
||||
"topic_frequency": topic_freq,
|
||||
"heading_patterns": heading_patterns,
|
||||
"competitors_analyzed": len(
|
||||
[c for c in content_data if c is not None]
|
||||
),
|
||||
"total_organic_results": len(organic),
|
||||
"featured_snippet": serp_data.get("featured_snippet"),
|
||||
}
|
||||
|
||||
|
||||
def detect_intent(
|
||||
keyword: str, organic: list[dict], serp_data: dict
|
||||
) -> str:
|
||||
"""
|
||||
Detect search intent from keyword and SERP features.
|
||||
|
||||
Returns: informational, commercial, transactional, or navigational
|
||||
"""
|
||||
kw_lower = keyword.lower()
|
||||
|
||||
# Navigational signals
|
||||
nav_signals = [
|
||||
"login",
|
||||
"sign in",
|
||||
"website",
|
||||
"official",
|
||||
".com",
|
||||
".org",
|
||||
]
|
||||
if any(s in kw_lower for s in nav_signals):
|
||||
return "navigational"
|
||||
|
||||
# Transactional signals
|
||||
transactional_signals = [
|
||||
"buy",
|
||||
"purchase",
|
||||
"order",
|
||||
"download",
|
||||
"subscribe",
|
||||
"deal",
|
||||
"discount",
|
||||
"coupon",
|
||||
"price",
|
||||
"pricing",
|
||||
"cost",
|
||||
"cheap",
|
||||
"free trial",
|
||||
]
|
||||
if any(s in kw_lower for s in transactional_signals):
|
||||
return "transactional"
|
||||
|
||||
# Commercial investigation signals
|
||||
commercial_signals = [
|
||||
"best",
|
||||
"top",
|
||||
"review",
|
||||
"comparison",
|
||||
"vs",
|
||||
"versus",
|
||||
"alternative",
|
||||
"vs.",
|
||||
"compared to",
|
||||
"pros and cons",
|
||||
]
|
||||
if any(s in kw_lower for s in commercial_signals):
|
||||
return "commercial"
|
||||
|
||||
# Informational signals
|
||||
informational_signals = [
|
||||
"how to",
|
||||
"what is",
|
||||
"what are",
|
||||
"why",
|
||||
"guide",
|
||||
"tutorial",
|
||||
"learn",
|
||||
"example",
|
||||
"definition",
|
||||
"meaning",
|
||||
"explain",
|
||||
]
|
||||
if any(s in kw_lower for s in informational_signals):
|
||||
return "informational"
|
||||
|
||||
# Default: check SERP patterns
|
||||
if serp_data.get("featured_snippet"):
|
||||
return "informational"
|
||||
|
||||
# If titles contain pricing/comparison language
|
||||
titles = [r.get("title", "").lower() for r in organic[:5]]
|
||||
title_text = " ".join(titles)
|
||||
|
||||
if any(s in title_text for s in ["best", "top", "review", "vs"]):
|
||||
return "commercial"
|
||||
if any(
|
||||
s in title_text for s in ["how to", "guide", "what", "tutorial"]
|
||||
):
|
||||
return "informational"
|
||||
|
||||
return "commercial" # default for ambiguous
|
||||
|
||||
|
||||
def _count_topic_frequency(
|
||||
content_data: list[Optional[dict]],
|
||||
) -> list[dict]:
|
||||
"""Count how often topics (H2/H3 headings) appear across competitors."""
|
||||
topic_counts = {}
|
||||
|
||||
for content in content_data:
|
||||
if content is None:
|
||||
continue
|
||||
|
||||
seen_in_page = set()
|
||||
for heading in content.get("headings", []):
|
||||
if heading.startswith(("H2:", "H3:")):
|
||||
topic = heading.split(":", 1)[1].strip().lower()
|
||||
# Normalize common variations
|
||||
topic = _normalize_topic(topic)
|
||||
if topic and topic not in seen_in_page:
|
||||
seen_in_page.add(topic)
|
||||
topic_counts[topic] = topic_counts.get(topic, 0) + 1
|
||||
|
||||
# Sort by frequency
|
||||
sorted_topics = sorted(
|
||||
topic_counts.items(), key=lambda x: x[1], reverse=True
|
||||
)
|
||||
return [
|
||||
{"topic": t, "competitor_count": c} for t, c in sorted_topics[:30]
|
||||
]
|
||||
|
||||
|
||||
def _normalize_topic(topic: str) -> str:
|
||||
"""Basic topic normalization."""
|
||||
# Remove common filler words at start
|
||||
for prefix in [
|
||||
"the ",
|
||||
"a ",
|
||||
"an ",
|
||||
"our ",
|
||||
"your ",
|
||||
"my ",
|
||||
"about ",
|
||||
]:
|
||||
if topic.startswith(prefix):
|
||||
topic = topic[len(prefix) :]
|
||||
|
||||
# Strip trailing punctuation
|
||||
topic = topic.rstrip(".:!?")
|
||||
|
||||
return topic.strip()
|
||||
|
||||
|
||||
def _analyze_heading_patterns(
|
||||
content_data: list[Optional[dict]],
|
||||
) -> dict:
|
||||
"""Analyze heading patterns across competitors."""
|
||||
h2_counts = []
|
||||
h3_counts = []
|
||||
|
||||
for content in content_data:
|
||||
if content is None:
|
||||
continue
|
||||
|
||||
headings = content.get("headings", [])
|
||||
h2_count = sum(1 for h in headings if h.startswith("H2:"))
|
||||
h3_count = sum(1 for h in headings if h.startswith("H3:"))
|
||||
h2_counts.append(h2_count)
|
||||
h3_counts.append(h3_count)
|
||||
|
||||
return {
|
||||
"avg_h2_count": (
|
||||
round(statistics.mean(h2_counts), 1) if h2_counts else 0
|
||||
),
|
||||
"avg_h3_count": (
|
||||
round(statistics.mean(h3_counts), 1) if h3_counts else 0
|
||||
),
|
||||
"median_h2_count": (
|
||||
int(statistics.median(h2_counts)) if h2_counts else 0
|
||||
),
|
||||
"median_h3_count": (
|
||||
int(statistics.median(h3_counts)) if h3_counts else 0
|
||||
),
|
||||
}
|
||||
@@ -0,0 +1,334 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
seo-agi research orchestrator.
|
||||
Pulls SERP data, keyword data, PAA questions, and competitor content analysis
|
||||
into a single research.json file that feeds content generation.
|
||||
|
||||
Usage:
|
||||
python3 research.py "<keyword>" [options]
|
||||
|
||||
Options:
|
||||
--serp-depth=N Number of SERP results to analyze (default: 10)
|
||||
--include-paa Include People Also Ask extraction (default: true)
|
||||
--location=CODE DataForSEO location code (default: 2840 = US)
|
||||
--language=CODE Language code (default: en)
|
||||
--output=FORMAT Output: json|compact|brief (default: compact)
|
||||
--save-dir=PATH Save raw data (default: ~/.local/share/seo-agi/research/)
|
||||
--content-depth=N Number of top results to parse for content (default: 5)
|
||||
--mock Use fixture data instead of live API calls
|
||||
"""
|
||||
|
||||
import sys
|
||||
import os
|
||||
import json
|
||||
import argparse
|
||||
from datetime import datetime, timezone
|
||||
from pathlib import Path
|
||||
|
||||
# Add parent dir to path for lib imports
|
||||
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
||||
|
||||
from lib.env import load_env, load_config, get_credentials, ensure_dirs
|
||||
from lib.dataforseo import DataForSEOClient
|
||||
from lib.serp_analyze import analyze_serp
|
||||
|
||||
|
||||
def parse_args():
|
||||
parser = argparse.ArgumentParser(description="SEO-AGI Research")
|
||||
parser.add_argument("keyword", help="Target keyword or topic")
|
||||
parser.add_argument(
|
||||
"--serp-depth", type=int, default=10, help="SERP depth"
|
||||
)
|
||||
parser.add_argument(
|
||||
"--content-depth",
|
||||
type=int,
|
||||
default=5,
|
||||
help="Number of competitors to parse content from",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--location", type=int, default=None, help="Location code"
|
||||
)
|
||||
parser.add_argument(
|
||||
"--language", default=None, help="Language code"
|
||||
)
|
||||
parser.add_argument(
|
||||
"--output",
|
||||
choices=["json", "compact", "brief"],
|
||||
default="compact",
|
||||
help="Output format",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--save-dir", default=None, help="Directory to save research data"
|
||||
)
|
||||
parser.add_argument(
|
||||
"--mock",
|
||||
action="store_true",
|
||||
help="Use fixture data (no API calls)",
|
||||
)
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def load_mock_data(keyword: str) -> dict:
|
||||
"""Load fixture data for testing without API calls."""
|
||||
fixtures_dir = (
|
||||
Path(__file__).parent.parent / "fixtures"
|
||||
)
|
||||
|
||||
serp_fixture = fixtures_dir / "serp_sample.json"
|
||||
keywords_fixture = fixtures_dir / "keywords_sample.json"
|
||||
|
||||
serp_data = {"organic": [], "paa": [], "featured_snippet": None}
|
||||
related_kw = []
|
||||
|
||||
if serp_fixture.exists():
|
||||
with open(serp_fixture) as f:
|
||||
serp_data = json.load(f)
|
||||
|
||||
if keywords_fixture.exists():
|
||||
with open(keywords_fixture) as f:
|
||||
related_kw = json.load(f)
|
||||
|
||||
return {
|
||||
"keyword": keyword,
|
||||
"timestamp": datetime.now(timezone.utc).isoformat(),
|
||||
"source": "mock",
|
||||
"serp": serp_data,
|
||||
"related_keywords": related_kw,
|
||||
"analysis": {
|
||||
"intent": "commercial",
|
||||
"word_count_stats": {
|
||||
"min": 800,
|
||||
"max": 3200,
|
||||
"median": 1800,
|
||||
"recommended_min": 1440,
|
||||
"recommended_max": 2340,
|
||||
},
|
||||
"paa_questions": serp_data.get("paa", []),
|
||||
"topic_frequency": [],
|
||||
"heading_patterns": {
|
||||
"avg_h2_count": 6,
|
||||
"avg_h3_count": 8,
|
||||
"median_h2_count": 5,
|
||||
"median_h3_count": 7,
|
||||
},
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def run_research(args) -> dict:
|
||||
"""Execute the full research pipeline."""
|
||||
creds = get_credentials()
|
||||
config = load_config()
|
||||
|
||||
location = args.location or config["default_location"]
|
||||
language = args.language or config["default_language"]
|
||||
|
||||
if args.mock:
|
||||
return load_mock_data(args.keyword)
|
||||
|
||||
if not creds["has_dataforseo"]:
|
||||
print(
|
||||
"ERROR: DataForSEO credentials not found.",
|
||||
file=sys.stderr,
|
||||
)
|
||||
print(
|
||||
"Run: python3 scripts/setup.py",
|
||||
file=sys.stderr,
|
||||
)
|
||||
print(
|
||||
"Or add DATAFORSEO_LOGIN and DATAFORSEO_PASSWORD to "
|
||||
"~/.config/seo-agi/.env",
|
||||
file=sys.stderr,
|
||||
)
|
||||
sys.exit(1)
|
||||
|
||||
client = DataForSEOClient(
|
||||
creds["dataforseo_login"], creds["dataforseo_password"]
|
||||
)
|
||||
|
||||
# Step 1: SERP results
|
||||
print(f"Fetching SERP for: {args.keyword}", file=sys.stderr)
|
||||
serp_data = client.serp_live(
|
||||
args.keyword, location, language, args.serp_depth
|
||||
)
|
||||
|
||||
# Step 2: Related keywords
|
||||
print("Fetching related keywords...", file=sys.stderr)
|
||||
related_kw = client.related_keywords(args.keyword, location, language)
|
||||
|
||||
# Step 3: Parse competitor content (top N)
|
||||
content_data = []
|
||||
organic = serp_data.get("organic", [])
|
||||
parse_count = min(args.content_depth, len(organic))
|
||||
|
||||
for i in range(parse_count):
|
||||
url = organic[i].get("url", "")
|
||||
if url:
|
||||
print(
|
||||
f"Parsing content ({i+1}/{parse_count}): {url[:80]}...",
|
||||
file=sys.stderr,
|
||||
)
|
||||
content = client.content_parse(url)
|
||||
content_data.append(content)
|
||||
|
||||
# Merge content data back into organic result
|
||||
if content:
|
||||
organic[i]["word_count"] = content.get("word_count", 0)
|
||||
organic[i]["headings"] = content.get("headings", [])
|
||||
else:
|
||||
content_data.append(None)
|
||||
|
||||
# Step 4: Analyze
|
||||
print("Analyzing competitive landscape...", file=sys.stderr)
|
||||
analysis = analyze_serp(serp_data, content_data, args.keyword)
|
||||
|
||||
# Assemble research output
|
||||
research = {
|
||||
"keyword": args.keyword,
|
||||
"timestamp": datetime.now(timezone.utc).isoformat(),
|
||||
"location": location,
|
||||
"language": language,
|
||||
"source": "dataforseo",
|
||||
"serp": serp_data,
|
||||
"related_keywords": related_kw[:20],
|
||||
"analysis": analysis,
|
||||
}
|
||||
|
||||
return research
|
||||
|
||||
|
||||
def save_research(research: dict, save_dir: str = None):
|
||||
"""Save research data to disk."""
|
||||
ensure_dirs()
|
||||
config = load_config()
|
||||
|
||||
if save_dir:
|
||||
out_dir = Path(save_dir).expanduser()
|
||||
else:
|
||||
out_dir = Path.home() / ".local" / "share" / "seo-agi" / "research"
|
||||
|
||||
out_dir.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
# Filename from keyword + date
|
||||
slug = (
|
||||
research["keyword"]
|
||||
.lower()
|
||||
.replace(" ", "-")
|
||||
.replace("/", "-")[:50]
|
||||
)
|
||||
date_str = datetime.now().strftime("%Y%m%d")
|
||||
filename = f"{slug}-{date_str}.json"
|
||||
|
||||
filepath = out_dir / filename
|
||||
with open(filepath, "w") as f:
|
||||
json.dump(research, f, indent=2)
|
||||
|
||||
print(f"Research saved: {filepath}", file=sys.stderr)
|
||||
return str(filepath)
|
||||
|
||||
|
||||
def format_compact(research: dict) -> str:
|
||||
"""Format research as compact human-readable output."""
|
||||
lines = []
|
||||
kw = research["keyword"]
|
||||
analysis = research.get("analysis", {})
|
||||
serp = research.get("serp", {})
|
||||
organic = serp.get("organic", [])
|
||||
|
||||
lines.append(f"# Research: {kw}")
|
||||
lines.append(f"Intent: {analysis.get('intent', 'unknown')}")
|
||||
|
||||
# Word count
|
||||
wc = analysis.get("word_count_stats", {})
|
||||
if wc:
|
||||
lines.append(
|
||||
f"Competitor word count: {wc.get('min', '?')}-{wc.get('max', '?')} "
|
||||
f"(median: {wc.get('median', '?')})"
|
||||
)
|
||||
lines.append(
|
||||
f"Recommended range: {wc.get('recommended_min', '?')}-"
|
||||
f"{wc.get('recommended_max', '?')} words"
|
||||
)
|
||||
|
||||
# Top results
|
||||
lines.append(f"\n## Top {len(organic)} Results")
|
||||
for r in organic[:10]:
|
||||
wc_str = (
|
||||
f" ({r.get('word_count', '?')} words)"
|
||||
if r.get("word_count")
|
||||
else ""
|
||||
)
|
||||
lines.append(f" {r['position']}. {r['title']}{wc_str}")
|
||||
lines.append(f" {r['url']}")
|
||||
|
||||
# PAA
|
||||
paa = analysis.get("paa_questions", serp.get("paa", []))
|
||||
if paa:
|
||||
lines.append(f"\n## People Also Ask ({len(paa)})")
|
||||
for q in paa:
|
||||
lines.append(f" - {q}")
|
||||
|
||||
# Related keywords
|
||||
related = research.get("related_keywords", [])
|
||||
if related:
|
||||
lines.append(f"\n## Related Keywords (top 10)")
|
||||
for kw_data in related[:10]:
|
||||
lines.append(
|
||||
f" - {kw_data['keyword']} "
|
||||
f"(vol: {kw_data['volume']}, "
|
||||
f"diff: {kw_data.get('difficulty', '?')})"
|
||||
)
|
||||
|
||||
# Topics
|
||||
topics = analysis.get("topic_frequency", [])
|
||||
if topics:
|
||||
lines.append(f"\n## Common Topics Across Competitors")
|
||||
for t in topics[:15]:
|
||||
lines.append(
|
||||
f" - {t['topic']} (in {t['competitor_count']} pages)"
|
||||
)
|
||||
|
||||
# Heading patterns
|
||||
hp = analysis.get("heading_patterns", {})
|
||||
if hp:
|
||||
lines.append(f"\n## Heading Structure")
|
||||
lines.append(
|
||||
f" Avg H2s: {hp.get('avg_h2_count', '?')}, "
|
||||
f"Avg H3s: {hp.get('avg_h3_count', '?')}"
|
||||
)
|
||||
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def main():
|
||||
args = parse_args()
|
||||
research = run_research(args)
|
||||
|
||||
# Save
|
||||
filepath = save_research(research, args.save_dir)
|
||||
|
||||
# Output
|
||||
if args.output == "json":
|
||||
print(json.dumps(research, indent=2))
|
||||
elif args.output == "brief":
|
||||
# Minimal output for piping into content generation
|
||||
analysis = research.get("analysis", {})
|
||||
brief_data = {
|
||||
"keyword": research["keyword"],
|
||||
"intent": analysis.get("intent"),
|
||||
"word_count_stats": analysis.get("word_count_stats"),
|
||||
"paa_questions": analysis.get(
|
||||
"paa_questions",
|
||||
research.get("serp", {}).get("paa", []),
|
||||
),
|
||||
"topic_frequency": analysis.get("topic_frequency", [])[:10],
|
||||
"heading_patterns": analysis.get("heading_patterns"),
|
||||
"research_file": filepath,
|
||||
}
|
||||
print(json.dumps(brief_data, indent=2))
|
||||
else:
|
||||
print(format_compact(research))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,160 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
SEO-AGI setup script.
|
||||
Creates config directory, prompts for API keys, installs dependencies.
|
||||
|
||||
Usage:
|
||||
python3 setup.py
|
||||
"""
|
||||
|
||||
import os
|
||||
import sys
|
||||
import json
|
||||
import subprocess
|
||||
from pathlib import Path
|
||||
|
||||
CONFIG_DIR = Path.home() / ".config" / "seo-agi"
|
||||
ENV_FILE = CONFIG_DIR / ".env"
|
||||
CONFIG_FILE = CONFIG_DIR / "config.json"
|
||||
OUTPUT_DIR = Path.home() / "Documents" / "SEO-AGI"
|
||||
|
||||
|
||||
def main():
|
||||
print("=" * 50)
|
||||
print(" SEO-AGI Setup")
|
||||
print("=" * 50)
|
||||
print()
|
||||
|
||||
# Create directories
|
||||
CONFIG_DIR.mkdir(parents=True, exist_ok=True)
|
||||
for subdir in ["research", "briefs", "pages", "rewrites"]:
|
||||
(OUTPUT_DIR / subdir).mkdir(parents=True, exist_ok=True)
|
||||
|
||||
data_dir = Path.home() / ".local" / "share" / "seo-agi"
|
||||
for subdir in ["research", "cache"]:
|
||||
(data_dir / subdir).mkdir(parents=True, exist_ok=True)
|
||||
|
||||
print("[OK] Directories created")
|
||||
|
||||
# Install Python dependencies
|
||||
print("\nInstalling dependencies...")
|
||||
deps = ["requests"]
|
||||
try:
|
||||
subprocess.check_call(
|
||||
[sys.executable, "-m", "pip", "install", "--quiet"]
|
||||
+ deps
|
||||
+ ["--break-system-packages"],
|
||||
stdout=subprocess.DEVNULL,
|
||||
stderr=subprocess.DEVNULL,
|
||||
)
|
||||
print("[OK] Core dependencies installed (requests)")
|
||||
except subprocess.CalledProcessError:
|
||||
try:
|
||||
subprocess.check_call(
|
||||
[sys.executable, "-m", "pip", "install", "--quiet"] + deps,
|
||||
stdout=subprocess.DEVNULL,
|
||||
stderr=subprocess.DEVNULL,
|
||||
)
|
||||
print("[OK] Core dependencies installed (requests)")
|
||||
except subprocess.CalledProcessError:
|
||||
print("[WARN] Could not install requests. Install manually:")
|
||||
print(" pip install requests")
|
||||
|
||||
# API Keys
|
||||
print("\n--- API Keys ---")
|
||||
print("DataForSEO is REQUIRED. Get credentials at:")
|
||||
print(" https://app.dataforseo.com/api-dashboard\n")
|
||||
|
||||
existing_env = {}
|
||||
if ENV_FILE.exists():
|
||||
with open(ENV_FILE) as f:
|
||||
for line in f:
|
||||
line = line.strip()
|
||||
if line and not line.startswith("#") and "=" in line:
|
||||
k, _, v = line.partition("=")
|
||||
existing_env[k.strip()] = v.strip()
|
||||
|
||||
# DataForSEO
|
||||
dfs_login = existing_env.get("DATAFORSEO_LOGIN", "")
|
||||
dfs_pass = existing_env.get("DATAFORSEO_PASSWORD", "")
|
||||
|
||||
if dfs_login:
|
||||
print(f"DataForSEO login found: {dfs_login[:4]}***")
|
||||
change = input("Update DataForSEO credentials? [y/N]: ").strip().lower()
|
||||
if change == "y":
|
||||
dfs_login = input("DataForSEO login (email): ").strip()
|
||||
dfs_pass = input("DataForSEO password: ").strip()
|
||||
else:
|
||||
dfs_login = input("DataForSEO login (email): ").strip()
|
||||
dfs_pass = input("DataForSEO password: ").strip()
|
||||
|
||||
# GSC (optional)
|
||||
print("\nGoogle Search Console (OPTIONAL - press Enter to skip)")
|
||||
gsc_path = existing_env.get("GSC_SERVICE_ACCOUNT_PATH", "")
|
||||
if gsc_path:
|
||||
print(f"GSC service account found: {gsc_path}")
|
||||
change = input("Update? [y/N]: ").strip().lower()
|
||||
if change == "y":
|
||||
gsc_path = input("Path to service account JSON: ").strip()
|
||||
else:
|
||||
gsc_path = input("Path to GSC service account JSON (or Enter to skip): ").strip()
|
||||
|
||||
# Write .env
|
||||
env_lines = [
|
||||
"# SEO-AGI Configuration",
|
||||
f"# Generated by setup.py on {__import__('datetime').datetime.now().isoformat()[:10]}",
|
||||
"",
|
||||
"# DataForSEO",
|
||||
f"DATAFORSEO_LOGIN={dfs_login}",
|
||||
f"DATAFORSEO_PASSWORD={dfs_pass}",
|
||||
"",
|
||||
"# Google Search Console",
|
||||
f"GSC_SERVICE_ACCOUNT_PATH={gsc_path}",
|
||||
"",
|
||||
"# Future integrations",
|
||||
f"AHREFS_API_KEY={existing_env.get('AHREFS_API_KEY', '')}",
|
||||
f"SEMRUSH_API_KEY={existing_env.get('SEMRUSH_API_KEY', '')}",
|
||||
"",
|
||||
"# Defaults",
|
||||
f"DEFAULT_LOCATION={existing_env.get('DEFAULT_LOCATION', '2840')}",
|
||||
f"DEFAULT_LANGUAGE={existing_env.get('DEFAULT_LANGUAGE', 'en')}",
|
||||
]
|
||||
|
||||
with open(ENV_FILE, "w") as f:
|
||||
f.write("\n".join(env_lines) + "\n")
|
||||
|
||||
print(f"\n[OK] Credentials saved to {ENV_FILE}")
|
||||
|
||||
# Write default config.json if not exists
|
||||
if not CONFIG_FILE.exists():
|
||||
default_config = {
|
||||
"default_location": 2840,
|
||||
"default_language": "en",
|
||||
"default_site": "",
|
||||
"serp_depth": 10,
|
||||
"save_research": True,
|
||||
"output_dir": str(OUTPUT_DIR),
|
||||
}
|
||||
with open(CONFIG_FILE, "w") as f:
|
||||
json.dump(default_config, f, indent=2)
|
||||
print(f"[OK] Config saved to {CONFIG_FILE}")
|
||||
|
||||
# Verify
|
||||
print("\n--- Setup Summary ---")
|
||||
print(f" DataForSEO: {'configured' if dfs_login else 'NOT SET'}")
|
||||
print(f" GSC: {'configured' if gsc_path else 'not configured (optional)'}")
|
||||
print(f" Output dir: {OUTPUT_DIR}")
|
||||
print(f" Config dir: {CONFIG_DIR}")
|
||||
|
||||
if dfs_login:
|
||||
print("\n[READY] Run a test:")
|
||||
print(f" python3 {Path(__file__).parent}/research.py \"test keyword\"")
|
||||
else:
|
||||
print("\n[WARN] DataForSEO credentials missing. The skill will fall")
|
||||
print(" back to Claude's web search for basic research.")
|
||||
|
||||
print()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,149 @@
|
||||
"""Tests for DataForSEO client response parsing."""
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(
|
||||
0, str(Path(__file__).parent.parent / "scripts")
|
||||
)
|
||||
|
||||
from lib.dataforseo import DataForSEOClient
|
||||
|
||||
|
||||
def test_extract_serp_empty():
|
||||
client = DataForSEOClient("test", "test")
|
||||
result = client._extract_serp({"tasks": []})
|
||||
assert result["organic"] == []
|
||||
assert result["paa"] == []
|
||||
assert result["featured_snippet"] is None
|
||||
|
||||
|
||||
def test_extract_serp_with_organic():
|
||||
raw = {
|
||||
"tasks": [{
|
||||
"result": [{
|
||||
"se_results_count": 1000000,
|
||||
"items": [
|
||||
{
|
||||
"type": "organic",
|
||||
"rank_absolute": 1,
|
||||
"url": "https://example.com/page",
|
||||
"domain": "example.com",
|
||||
"title": "Test Page",
|
||||
"description": "A test description",
|
||||
},
|
||||
{
|
||||
"type": "organic",
|
||||
"rank_absolute": 2,
|
||||
"url": "https://other.com/page",
|
||||
"domain": "other.com",
|
||||
"title": "Other Page",
|
||||
"description": "Another description",
|
||||
},
|
||||
{
|
||||
"type": "people_also_ask",
|
||||
"items": [
|
||||
{"title": "What is a test?"},
|
||||
{"title": "How do tests work?"},
|
||||
],
|
||||
},
|
||||
]
|
||||
}]
|
||||
}]
|
||||
}
|
||||
|
||||
client = DataForSEOClient("test", "test")
|
||||
result = client._extract_serp(raw)
|
||||
|
||||
assert len(result["organic"]) == 2
|
||||
assert result["organic"][0]["position"] == 1
|
||||
assert result["organic"][0]["url"] == "https://example.com/page"
|
||||
assert result["organic"][1]["title"] == "Other Page"
|
||||
|
||||
assert len(result["paa"]) == 2
|
||||
assert result["paa"][0] == "What is a test?"
|
||||
|
||||
assert result["total_results"] == 1000000
|
||||
|
||||
|
||||
def test_extract_keywords_empty():
|
||||
client = DataForSEOClient("test", "test")
|
||||
result = client._extract_keywords({"tasks": []})
|
||||
assert result == []
|
||||
|
||||
|
||||
def test_extract_keywords_with_data():
|
||||
raw = {
|
||||
"tasks": [{
|
||||
"result": [{
|
||||
"items": [
|
||||
{
|
||||
"keyword_data": {
|
||||
"keyword": "test keyword one",
|
||||
"keyword_info": {
|
||||
"search_volume": 5000,
|
||||
"cpc": 1.50,
|
||||
"competition": 0.7,
|
||||
},
|
||||
"keyword_properties": {
|
||||
"keyword_difficulty": 42,
|
||||
},
|
||||
}
|
||||
},
|
||||
{
|
||||
"keyword_data": {
|
||||
"keyword": "test keyword two",
|
||||
"keyword_info": {
|
||||
"search_volume": 8000,
|
||||
"cpc": 2.10,
|
||||
"competition": 0.85,
|
||||
},
|
||||
"keyword_properties": {
|
||||
"keyword_difficulty": 55,
|
||||
},
|
||||
}
|
||||
},
|
||||
]
|
||||
}]
|
||||
}]
|
||||
}
|
||||
|
||||
client = DataForSEOClient("test", "test")
|
||||
result = client._extract_keywords(raw)
|
||||
|
||||
# Should be sorted by volume descending
|
||||
assert len(result) == 2
|
||||
assert result[0]["keyword"] == "test keyword two"
|
||||
assert result[0]["volume"] == 8000
|
||||
assert result[1]["keyword"] == "test keyword one"
|
||||
assert result[1]["volume"] == 5000
|
||||
assert result[1]["difficulty"] == 42
|
||||
|
||||
|
||||
def test_extract_headings():
|
||||
page_content = {
|
||||
"h1": ["Main Title"],
|
||||
"h2": ["Section One", "Section Two"],
|
||||
"h3": ["Subsection A", "Subsection B"],
|
||||
}
|
||||
headings = DataForSEOClient._extract_headings(page_content)
|
||||
assert "H1: Main Title" in headings
|
||||
assert "H2: Section One" in headings
|
||||
assert "H3: Subsection A" in headings
|
||||
assert len(headings) == 5
|
||||
|
||||
|
||||
def test_auth_header():
|
||||
client = DataForSEOClient("user@test.com", "mypassword")
|
||||
assert client._auth_header.startswith("Basic ")
|
||||
assert len(client._auth_header) > 10
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
test_extract_serp_empty()
|
||||
test_extract_serp_with_organic()
|
||||
test_extract_keywords_empty()
|
||||
test_extract_keywords_with_data()
|
||||
test_extract_headings()
|
||||
test_auth_header()
|
||||
print("All tests passed.")
|
||||
@@ -0,0 +1,50 @@
|
||||
"""Tests for env module."""
|
||||
|
||||
import sys
|
||||
import os
|
||||
import json
|
||||
import tempfile
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(
|
||||
0, str(Path(__file__).parent.parent / "scripts")
|
||||
)
|
||||
|
||||
from lib.env import load_config, _determine_mode, DEFAULT_CONFIG
|
||||
|
||||
|
||||
def test_default_config():
|
||||
config = DEFAULT_CONFIG.copy()
|
||||
assert config["default_location"] == 2840
|
||||
assert config["default_language"] == "en"
|
||||
assert config["serp_depth"] == 10
|
||||
assert config["save_research"] is True
|
||||
|
||||
|
||||
def test_determine_mode_full():
|
||||
creds = {"has_dataforseo": True, "has_gsc": True, "has_ahrefs": False, "has_semrush": False}
|
||||
assert _determine_mode(creds) == "full"
|
||||
|
||||
|
||||
def test_determine_mode_dataforseo_only():
|
||||
creds = {"has_dataforseo": True, "has_gsc": False, "has_ahrefs": False, "has_semrush": False}
|
||||
assert _determine_mode(creds) == "dataforseo-only"
|
||||
|
||||
|
||||
def test_determine_mode_gsc_only():
|
||||
creds = {"has_dataforseo": False, "has_gsc": True, "has_ahrefs": False, "has_semrush": False}
|
||||
assert _determine_mode(creds) == "gsc-only"
|
||||
|
||||
|
||||
def test_determine_mode_fallback():
|
||||
creds = {"has_dataforseo": False, "has_gsc": False, "has_ahrefs": False, "has_semrush": False}
|
||||
assert _determine_mode(creds) == "fallback"
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
test_default_config()
|
||||
test_determine_mode_full()
|
||||
test_determine_mode_dataforseo_only()
|
||||
test_determine_mode_gsc_only()
|
||||
test_determine_mode_fallback()
|
||||
print("All tests passed.")
|
||||
@@ -0,0 +1,125 @@
|
||||
"""Tests for serp_analyze module."""
|
||||
|
||||
import sys
|
||||
import os
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
# Add scripts dir to path
|
||||
sys.path.insert(
|
||||
0, str(Path(__file__).parent.parent / "scripts")
|
||||
)
|
||||
|
||||
from lib.serp_analyze import analyze_serp, detect_intent
|
||||
|
||||
|
||||
def load_fixture(name: str):
|
||||
fixtures_dir = Path(__file__).parent.parent / "fixtures"
|
||||
with open(fixtures_dir / name) as f:
|
||||
return json.load(f)
|
||||
|
||||
|
||||
def test_detect_intent_commercial():
|
||||
assert detect_intent("best project management tools", [], {}) == "commercial"
|
||||
assert detect_intent("X vs Y comparison", [], {}) == "commercial"
|
||||
assert detect_intent("slack alternatives", [], {}) == "commercial"
|
||||
|
||||
|
||||
def test_detect_intent_informational():
|
||||
assert detect_intent("how to park at JFK", [], {}) == "informational"
|
||||
assert detect_intent("what is SEO", [], {}) == "informational"
|
||||
assert detect_intent("python tutorial for beginners", [], {}) == "informational"
|
||||
|
||||
|
||||
def test_detect_intent_transactional():
|
||||
assert detect_intent("buy parking pass JFK", [], {}) == "transactional"
|
||||
assert detect_intent("JFK parking coupon discount", [], {}) == "transactional"
|
||||
assert detect_intent("cheap parking near me", [], {}) == "transactional"
|
||||
|
||||
|
||||
def test_detect_intent_navigational():
|
||||
assert detect_intent("JFK airport official website", [], {}) == "navigational"
|
||||
assert detect_intent("parkingaccess.com login", [], {}) == "navigational"
|
||||
|
||||
|
||||
def test_analyze_serp_with_fixtures():
|
||||
serp_data = load_fixture("serp_sample.json")
|
||||
|
||||
# Build content_data from the fixture's embedded word_count/headings
|
||||
content_data = []
|
||||
for item in serp_data.get("organic", []):
|
||||
if item.get("word_count"):
|
||||
content_data.append({
|
||||
"word_count": item["word_count"],
|
||||
"headings": item.get("headings", []),
|
||||
})
|
||||
else:
|
||||
content_data.append(None)
|
||||
|
||||
analysis = analyze_serp(serp_data, content_data, "airport parking JFK")
|
||||
|
||||
assert analysis["keyword"] == "airport parking JFK"
|
||||
assert analysis["intent"] in (
|
||||
"informational", "commercial", "transactional", "navigational"
|
||||
)
|
||||
|
||||
# Word count stats should reflect fixture data
|
||||
wc = analysis["word_count_stats"]
|
||||
assert wc["min"] > 0
|
||||
assert wc["max"] >= wc["min"]
|
||||
assert wc["median"] > 0
|
||||
assert wc["recommended_min"] > 0
|
||||
assert wc["recommended_max"] > wc["recommended_min"]
|
||||
|
||||
# PAA should come through
|
||||
assert len(analysis["paa_questions"]) > 0
|
||||
|
||||
# Topic frequency should have entries from headings
|
||||
assert len(analysis["topic_frequency"]) > 0
|
||||
|
||||
# Heading patterns should be populated
|
||||
hp = analysis["heading_patterns"]
|
||||
assert hp["avg_h2_count"] > 0
|
||||
|
||||
|
||||
def test_analyze_serp_empty():
|
||||
"""Handle empty SERP data gracefully."""
|
||||
analysis = analyze_serp(
|
||||
{"organic": [], "paa": [], "featured_snippet": None},
|
||||
[],
|
||||
"nonexistent keyword",
|
||||
)
|
||||
assert analysis["keyword"] == "nonexistent keyword"
|
||||
assert analysis["word_count_stats"] == {}
|
||||
assert analysis["competitors_analyzed"] == 0
|
||||
|
||||
|
||||
def test_analyze_serp_partial_content():
|
||||
"""Handle mix of parsed and failed content parses."""
|
||||
serp_data = {
|
||||
"organic": [
|
||||
{"position": 1, "url": "https://a.com", "title": "A", "description": ""},
|
||||
{"position": 2, "url": "https://b.com", "title": "B", "description": ""},
|
||||
],
|
||||
"paa": ["Question 1?"],
|
||||
"featured_snippet": None,
|
||||
}
|
||||
content_data = [
|
||||
{"word_count": 2000, "headings": ["H1: Title", "H2: Section One", "H2: Section Two"]},
|
||||
None, # failed parse
|
||||
]
|
||||
|
||||
analysis = analyze_serp(serp_data, content_data, "test keyword")
|
||||
assert analysis["competitors_analyzed"] == 1
|
||||
assert analysis["word_count_stats"]["median"] == 2000
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
test_detect_intent_commercial()
|
||||
test_detect_intent_informational()
|
||||
test_detect_intent_transactional()
|
||||
test_detect_intent_navigational()
|
||||
test_analyze_serp_with_fixtures()
|
||||
test_analyze_serp_empty()
|
||||
test_analyze_serp_partial_content()
|
||||
print("All tests passed.")
|
||||
Reference in New Issue
Block a user