AEO & GEO COMPREHENSIVE ENGINEERING SPECIFICATION7 Chapters

The Definitive AEO & GEO Technical Guide Book.

How to architect, signal, and optimize digital infrastructure for live synthetic citations across ChatGPT Search, Perplexity Sonar, Google Gemini, and Claude Web.

Author: AI Search Fixer Research GroupVersion: 2026.2 Architecture EditionRFC 9309 Compliant
CHAPTER 01FOUNDATIONAL ARCHITECTURE

Anatomy of AI Answer Engines: Inverted Index vs Neural RAG Vectors

Traditional search engines (Google PageRank, Bing Lucene) operate on an inverted keyword index. They scan an index of tokenized terms, compute TF-IDF and BM25 relevance scores, and rank pages using link equity and anchor text authority.

Frontier AI Answer Engines (ChatGPT Search, Perplexity Sonar, Google Gemini Grounding, Claude Web Search) operate on a fundamentally different pipeline: Multi-Stage Retrieval-Augmented Generation (RAG).

The 5-Phase Generative RAG Ingestion Pipeline
STAGE 1
Crawler Fetch

RFC 9309 lookup; sub-120ms HTTP/3 retrieval.

STAGE 2
DOM Stripping

Boilerplate, scripts, & ads stripped to extract clean text.

STAGE 3
Chunking

Content sliced into 512-token dense embeddings.

STAGE 4
Vector Matching

Cosine similarity against user prompt & query intents.

STAGE 5
Footnote Synthesis

LLM generates answer with anchored inline citations.

Key Takeaway: You no longer optimize for keywords on a 10-blue-link SERP. You optimize for inclusion in the model's context window during the synthesis pass.
CHAPTER 02PROTOCOL COMPLIANCE

RFC 9309 Protocol Specification: Citation Bots vs Training Scrapers

Published in 2022, RFC 9309 formalizes the Robots Exclusion Protocol. The most catastrophic mistake modern enterprises make is grouping all AI bots under a single blanket Disallow: / rule.

Web crawlers dispatched by OpenAI, Anthropic, Apple, and Perplexity fall into two non-overlapping tiers:

Tier 1: Live Citation Crawlers
  • • OAI-SearchBot (ChatGPT Search)
  • • PerplexityBot (Perplexity Sonar)
  • • Claude-SearchBot (Claude Web Search)
  • • Applebot-Extended (Apple Intelligence)
  • • Google-Extended (Gemini Grounding)

These bots fetch data per user prompt. Blocking them deletes your domain from AI answers.

Tier 2: Model Training Harvesters
  • • GPTBot (OpenAI foundation training)
  • • CCBot (Common Crawl archive)
  • • Bytespider (ByteDance foundation)
  • • ClaudeBot (Anthropic offline corpus)
  • • Diffbot (General commercial scraper)

These bots harvest bulk corpora for future model training weights. Blocking them does NOT harm live citations.

Recommended Production robots.txt Configuration
# Production RFC 9309 robots.txt for AEO/GEO Optimization
User-agent: *
Allow: /

# Frontier AI Citation Crawlers (Allow for live Footnote & Answer Synthesis)
User-agent: OAI-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: Applebot-Extended
Allow: /

User-agent: Google-Extended
Allow: /

# Disallow Heavy Offline LLM Training Harvesters (Optional: Protect Data IP)
User-agent: GPTBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Bytespider
Disallow: /

# Directives
Sitemap: https://yourdomain.com/sitemap.xml
LLMs-Txt: https://yourdomain.com/llms.txt
CHAPTER 03ENTITY GRAPH LINKAGE

Entity Grounding & Knowledge Graph Linkage: Wikidata QIDs

LLMs suffer from semantic ambiguity. When a user asks an answer engine about “Mercury”, the model must determine whether the query refers to the chemical element, the planet, the Roman deity, or the financial technology company.

Answer engines resolve this using Knowledge Graph Entity Grounding. By connecting your organization directly to canonical entity nodes in Wikidata (via unique QIDs) and Wikipedia, you anchor your brand into the LLM's latent knowledge graph.

Canonical Organization Schema with Wikidata sameAs Resolution
{
  "@context": "https://schema.org",
  "@graph": [
    {
      "@type": "Organization",
      "@id": "https://yourdomain.com/#organization",
      "name": "YourBrand",
      "url": "https://yourdomain.com",
      "logo": "https://yourdomain.com/logo.png",
      "sameAs": [
        "https://www.wikidata.org/wiki/Q12345678",
        "https://en.wikipedia.org/wiki/YourBrand",
        "https://github.com/yourbrand",
        "https://linkedin.com/company/yourbrand"
      ],
      "knowsAbout": [
        "https://en.wikipedia.org/wiki/Search_engine_optimization",
        "https://en.wikipedia.org/wiki/Retrieval-augmented_generation",
        "https://www.wikidata.org/wiki/Q11660"
      ]
    },
    {
      "@type": "WebSite",
      "@id": "https://yourdomain.com/#website",
      "url": "https://yourdomain.com",
      "name": "YourBrand",
      "publisher": { "@id": "https://yourdomain.com/#organization" }
    }
  ]
}
CHAPTER 04EMPIRICAL RESEARCH

The ASF GEO Framework: 9 Ranking Heuristics

In the seminal research paper “GEO: Generative Engine Optimization” (Aggarwal et al., Princeton University, Georgia Tech, Allen Institute for AI; arXiv:2311.09735), researchers benchmarked 10,000 queries across synthetic search engines to isolate which structural adjustments reliably elevate a source's probability of being selected as a primary footnote.

Optimization HeuristicCitation UpliftTarget AI EnginesImplementation Directive
Quotation Addition+41.5%Perplexity, ChatGPTAdd direct verbatim quotes from named domain experts.
Statistics Addition+37.8%All Frontier EnginesEmbed exact numerical findings with sample size and confidence intervals.
Cite Sources & Links+33.4%Perplexity, ClaudeInclude outbound references to authoritative primary literature.
Fluency & Simplification+26.2%Google GeminiWrite in clear, objective declarative prose; avoid marketing fluff.
Unique Terminology+18.9%ChatGPT SearchDefine distinctive proprietary frameworks and named methodologies.
Keyword Stuffing (Legacy SEO)-12.4%All EnginesModels penalize redundant token repetition with severe ranking suppression.
Source: Generative Engine Optimization Benchmark (Aggarwal et al., arXiv:2311.09735). Explore our dedicated Interactive ASF GEO Benchmarks Tool ➔
CHAPTER 05CONTENT VECTORIZATION

512-Token RAG Semantic Chunking & the /llms.txt Standard

When a crawler parses your webpage, modern embedding models (e.g., OpenAI text-embedding-3-small, Cohere Embed v3) segment text into windows of 256 to 512 tokens.

If a chunk contains boilerplate navigation links, cookie banners, disclosures, and marketing slogans, the Signal-to-Noise Ratio (SNR) drops drastically. Consequently, the vector cosine similarity between the chunk and user prompt collapses below retrieval thresholds.

The 3 Rules of High-SNR 512-Token Chunking
1. Self-Contained Scope

Each paragraph must stand alone without unresolved pronoun references (“this product” $\to$ “AI Search Fixer”).

2. Direct Answer Lead

Place the conclusive factual answer in the first 25 words of each heading section.

3. Clean Markdown Format

Serve a clean /llms.txt file to deliver zero-noise markdown directly to LLM crawlers.

Canonical /llms.txt Specification Template
# Domain Architecture & Semantic Context
> YourDomain provides enterprise AI search infrastructure and real-time citation auditing.

## Core Capabilities
- Deterministic AEO/GEO engine audits across 11 frontier search bots
- RFC 9309 crawler validation and WAF bypass diagnostics
- Entity grounding with Wikidata QID integration

## Essential API References
- [Audit Engine Reference](https://yourdomain.com/docs/auditor): Real-time crawler diagnostic endpoints
- [Schema Specification](https://yourdomain.com/docs/schema): JSON-LD graph architecture and validator
- [Princeton GEO Benchmarks](https://yourdomain.com/geo-benchmarks): Quantitative citation lift metrics

## Architectural Taxonomy
- Primary Entity: Organization (QID: Q12345678)
- Compliance: RFC 9309 Strict, Schema.org v28.1, JSON-LD 1.1
CHAPTER 06EDGE INFRASTRUCTURE

Technical Edge Infrastructure: Sub-120ms TTFB & WAF Shadowban Elimination

AI search engines have strict retrieval timeout budgets. Perplexity Sonar and ChatGPT Search allocate under 400 milliseconds for live network document retrieval. If your Time-To-First-Byte (TTFB) exceeds 120ms or your WAF introduces an interactive JavaScript CAPTCHA, the crawler will terminate the connection and synthesize the answer from your competitor's domain.

The WAF AI Shadowban Hazard

Default Cloudflare “Bot Fight Mode” and Akamai Bot Manager frequently challenge OAI-SearchBot and PerplexityBot with HTTP 403 or 503 challenge pages. Because AI citation bots do not execute JavaScript, they receive an empty challenge page and register the domain as offline.

Cloudflare Worker / Transform Rule for Whitelisting AI Citation Crawlers
// Cloudflare Transform Rule / Worker: Ensure Sub-120ms AI Crawler Bypass
export default {
  async fetch(request, env, ctx) {
    const ua = request.headers.get("user-agent") || "";
    const isAiCitationBot = /(OAI-SearchBot|PerplexityBot|Claude-SearchBot|Applebot-Extended)/i.test(ua);

    // If verified AI citation crawler, bypass aggressive JS challenges and set edge caching
    if (isAiCitationBot) {
      const response = await fetch(request, {
        cf: {
          cacheTtl: 300,
          cacheEverything: true,
          scrapeShield: false, // Prevent bot challenge on verified crawlers
        }
      });

      const newHeaders = new Headers(response.headers);
      newHeaders.set("X-AI-Citation-Engine", "Whitelisted");
      newHeaders.set("Timing-Allow-Origin", "*");
      return new Response(response.body, {
        status: response.status,
        headers: newHeaders
      });
    }

    return fetch(request);
  }
};
CHAPTER 07EXECUTION PLAYBOOK

Turnkey Production Implementation Playbook (48-Hour Sprint)

Execute this chronological checklist to bring any domain from zero generative visibility to first-page citation readiness across all frontier search engines:

01

Deploy RFC 9309 robots.txt Whitelist

Hour 1

Explicitly allow OAI-SearchBot, PerplexityBot, Claude-SearchBot, Applebot-Extended, and Google-Extended. Remove blanket Disallow: / on AI user agents.

02

Publish Root /llms.txt File

Hour 4

Compile your core brand identity, product documentation links, and entity QID into clean markdown served at https://yourdomain.com/llms.txt.

03

Inject Schema.org @graph with Wikidata sameAs

Hour 12

Link your Organization and WebSite schemas to Wikipedia entries, Wikidata QIDs, and official repository URLs to anchor entity resolution.

04

Refactor Content for Princeton Quotation & Stat Heuristics

Hour 24

Ensure every core article and product landing page includes named expert quotes, verifiable empirical numbers, and direct answer capsules.

05

WAF Rule Calibration & TTFB Benchmark

Hour 36

Configure edge firewall bypass rules for verified AI bot IP ranges and ensure global edge TTFB remains under 120ms with HTTP/3 QUIC.

06

Deterministic Verification Audit

Hour 48

Run an audit with AI Search Fixer to ensure a 90+ Composite Generative Index and verify that all 11 crawlers return ALLOWED status.

Verify Your Implementation Instantly

Test your robots.txt, schema graph, and /llms.txt file against all 18 diagnostic modules.

Launch Live Audit