This Week in Brief
Semrush's expanded 2026 AI Visibility Index — built on 126 million U.S. AI search prompts — has landed as the most data-dense benchmark the GEO market has seen, while Seer Interactive's v3 CTR study signals that AI Overview click suppression may be stabilising after a year of decline. Separately, Microsoft's MAI-Thinking-1 preprint confirms Common Crawl was used in training despite public marketing claims to the contrary, a disclosure with direct implications for how practitioners understand AI knowledge provenance.
Market Analysis — GEO & ASO
2026 AI Visibility Index: 126 Million U.S. AI Search Prompts Analysed
Per a study from Semrush, the 2026 AI Visibility Index scaled from an initial 2,500-prompt baseline (September 2025) to 126 million U.S. AI search prompts captured January–April 2026, making it one of the largest AI search citation datasets published to date. The study maps how brands are mentioned, cited, and surfaced across major AI search platforms. For GEO practitioners, the scale of the dataset provides a more statistically robust foundation for benchmarking citation share than prior single-platform or small-sample analyses.
AEO Periodic Table V4: Framework Drawn From 1.13 Million Prompts Across Six AI Engines
Per a study from Goodie, the fourth edition of its AEO Periodic Table was derived from 1.13 million prompts run across ChatGPT, Claude, Perplexity, Grok, Gemini, and Google AI Mode. The framework identifies that brand mentions across third-party surfaces — Wikipedia, Reddit, YouTube, review sites, and forums — correlate with AI citation outcomes as significantly as, and often more than, a brand's own website content. Practitioners should note the study's finding that no single team typically owns the full surface area that AI engines draw from when forming a brand answer.
AIO Impact on Google CTR: V3 Study Covering 5.47 Million Queries Shows Decline Stabilising
Per a study from Seer Interactive, analysis of 53 brands, 5.47 million tracked queries, and 2.43 billion organic impressions across full-year 2025 and Q1 2026 actuals found that the CTR decline attributable to AI Overviews has not continued to deepen as Q1 2026 projections anticipated — the trend has levelled off. This is directionally significant for practitioners modelling traffic impact: the acute suppression phase may be plateauing, though the study notes the data is directional rather than causal.
AI Search & ASO
Conductor 7-Month Study: Each AI Engine Has a Distinct 'Editorial Identity' for Citations
A seven-month analysis by Conductor tracking 1,056 data points across ChatGPT, ChatGPT Search, Perplexity, Google AI Overviews, Google AI Mode, Gemini, and Claude (September 2025–March 2026) found that each engine exhibits a persistent, intent-specific source preference: ChatGPT and ChatGPT Search are the only engines that regularly surface Wikipedia; Perplexity and Google Gemini both heavily favour YouTube across most intents; Google AI Mode routes users back to Google properties exclusively. The core practitioner finding is that a single AEO content strategy cannot cover the full AI search ecosystem — engine-specific citation behaviour requires engine-specific content and authority-signal mapping.
Ahrefs Updates AI Overview Citation Study: 38% of Citations Pull From Top-10 Organic Results
Ahrefs' updated citation analysis — covering 863,000 keyword SERPs and 4 million AI Overview URLs, more than double the previous study — finds that 38% of AI Overview citations are drawn from pages already ranking in the organic top 10, now that AI Overviews are powered by Gemini 3 (as of January 2026). The finding reinforces that organic search authority remains a prerequisite for AI Overview citation rather than an alternative path; practitioners cannot decouple GEO from foundational SEO ranking work.
AI Lab Signals
At Build 2026 on June 2, Microsoft AI CEO Mustafa Suleyman described MAI-Thinking-1 as trained exclusively on 'enterprise grade, clean and commercially licensed data.' The model's own published technical preprint describes a pipeline that includes Common Crawl — the unlicensed open web repository — with 24.2 billion pages retained after filtering and deduplication, a discrepancy first flagged by developer Simon Willison and confirmed by The Decoder. For GEO practitioners, the episode is a reminder that vendor training-data disclosures require cross-referencing with primary technical documents before informing content or compliance strategy.
Perplexity Integrates Deep Research Into 'Computer' Mode, Routing Queries Across 20+ Models
Perplexity has moved its Deep Research capability into its 'Computer' product, routing tasks dynamically across more than 20 AI models depending on the query type. The multi-model architecture means that citation and retrieval behaviour within Perplexity Deep Research may differ substantially from standard Perplexity Search, as different underlying models apply different retrieval and ranking heuristics. Practitioners optimising specifically for Perplexity citation should account for this routing layer when evaluating which content formats and authority signals perform.
Epoch AI Estimates High-Quality Public Training Data Exhaustion by 2028
Epoch AI researchers (Villalobos et al.) estimate the public internet holds roughly 15–20 trillion tokens of high-quality human-written text, with median exhaustion of that supply projected around 2028 at current frontier training rates; a single GPT-4-scale run may have already consumed 60–80% of the available stock. For GEO practitioners, this accelerates the strategic importance of retrieval-augmented generation (RAG) as the primary mechanism by which AI systems access current information — shifting the optimisation target firmly from training-time inclusion to real-time retrievability.
Training Data & Crawl
Microsoft MAI Preprint: 24.2 Billion Common Crawl Pages in Training Pipeline
The MAI-Thinking-1 technical preprint details a training data pipeline in which Common Crawl — after filtering and deduplication — contributed 24.2 billion pages, contradicting Microsoft's public-facing claim of exclusive reliance on commercially licensed sources. The disclosure confirms that web-crawled content, including content published under no explicit AI-training licence, continues to underpin frontier model training at scale. Content publishers relying on robots.txt directives or site terms to exclude AI training should treat such signals as probabilistic rather than guaranteed.
LLM Training Eligibility vs. RAG Retrievability: A Practitioner Distinction
A practitioner analysis from DerivateX maps the five-stage filtering pipeline — quality scoring, deduplication, toxicity screening, crawler access verification, and domain authority checks — that determines whether content enters a training corpus before weights are finalised. The piece draws an explicit distinction between training-time inclusion and real-time RAG retrievability, arguing that most AI responses today are driven by retrieval rather than parametric memory. For GEO practitioners, the implication is direct: optimising for crawler access and structured content extractability is a higher-leverage activity than attempting to influence training-data inclusion.
Research Radar (arXiv)
KARLA: Knowledge-base Augmented Retrieval for Language Models
KARLA proposes training LLMs to emit special tokens during generation that trigger live queries to an external knowledge base, enabling factual updates without model retraining and providing full traceability of facts to their source. Experiments show improvements in factual grounding for both short- and long-form generation, and the authors demonstrate that smaller models can match larger models' factual accuracy when equipped with knowledge-base access. For GEO practitioners, the architecture underlines the trajectory toward structured, queryable knowledge sources — well-maintained, entity-rich, and citable content — as the primary surface AI systems will draw facts from, independent of training cycles.
Retrieval-Augmented Generation for Natural Language Processing: A Survey
This open-access systematic review in Artificial Intelligence Review introduces a novel taxonomy of retrieval fusion methods — query-based, logits-based, latent, and parametric — and provides structured comparisons of retriever architectures across NLP tasks, while identifying hallucination and knowledge-currency as the core failure modes RAG is designed to address. For GEO practitioners, the survey is a useful primary reference confirming that RAG — not static parametric memory — is the dominant mechanism governing what AI systems cite in real-time responses, making content retrievability the central optimisation target.
Practitioner Takeaway
Audit your content for engine-specific citation readiness this week, not for 'AI search' as a monolith. The Conductor 7-month study confirms that ChatGPT, Perplexity, Google AI Overviews, and Gemini each apply distinct source preferences by intent type. Map your three highest-priority query intents against each engine's editorial identity — for example, ensure YouTube presence if Perplexity or Gemini citation is a target — and cross-reference against the Ahrefs finding that 38% of AI Overview citations come from organic top-10 pages. If your pages are not ranking in the top 10 organically for a given query, AI Overview citation for that query is statistically unlikely regardless of GEO-specific optimisation applied on-page.
The 6-phase framework used to structure this newsletter is available as a complete methodology guide — including audit tools, templates, and implementation checklists.
Get Access — free trial, then $19.99/mo or $200/yrNew to AI knowledge publication? Download the free briefing flyer — the data case for why your organisation cannot wait.