This Month in Brief
Two legal developments — a Munich court ruling treating Google AI Overviews as a media publisher and a U.S. district court partially upholding Reddit's DMCA anti-circumvention claims against Perplexity and SerpApi — signal that the regulatory perimeter around AI search is tightening faster than platform policy alone. Meanwhile, Goodie's 31-million-citation study and Semrush's 126-million-prompt AI Visibility Index together supply the most granular public benchmarks yet on which surfaces cite which sources, giving practitioners a clearer empirical basis for prioritisation decisions.
Market Analysis — GEO & ASO
Per a study from Goodie, analysis of 1.13 million prompts across ChatGPT, Claude, Perplexity, Grok, Gemini, and Google AI Mode finds that a brand's AI citation profile is shaped as much by third-party surfaces — Wikipedia, Reddit, YouTube, review sites, and forums — as by its own indexed pages. The study notes that SEO, PR, and social functions in most organisations share no common analytics layer, leaving no single team with visibility over the full signal set that determines AI answer inclusion. Practitioners should audit third-party brand-mention surfaces independently of their owned-content optimisation work.
Across 31 million AI citations and robots.txt audits of 105 US and UK news publishers, Goodie found that Grok, Google AI Overviews, and DeepSeek collectively account for roughly half of all news citations in the sample while providing no functioning opt-out mechanism that honours publisher blocks. Of 37 major news domains tracked, only 34 recorded any citations at all. For GEO practitioners advising publisher clients, this data indicates that access-control strategy must be modelled engine-by-engine rather than assumed to operate uniformly across the AI search ecosystem.
Semrush 2026 AI Visibility Index: 126 Million U.S. AI Search Prompts Analysed Across Major Platforms
Semrush expanded its AI Visibility Index from an initial 2,500 prompts to 126 million U.S. AI search prompts from January through April 2026, examining how brands are mentioned, cited, and surfaced across major AI search platforms. The scale of the dataset makes it the largest single public benchmark of AI brand-mention behaviour to date, though the study's methodology and full variable definitions remain behind a registration gate. Practitioners should treat the index as a market-sizing reference while cross-checking its category-level findings against engine-specific citation data from primary sources.
AI Search & ASO
The Munich Regional Court I issued a ruling on 28 May 2026 treating Google AI Overviews as a media publisher for liability purposes, a legal classification that — if upheld and replicated — would require AI-generated search summaries to meet editorial standards applied to conventional publishers. According to Lex Machina Review's docket mapping (https://legalaiinsights.com/risk-digest/google-ai-overviews-publisher-legal-risk), Google AI Overviews and Gemini now face confirmed legal challenges across at least four jurisdictions: Germany (Munich Regional Court I), the United States (federal courts), the European Union (European Commission), and the United Kingdom (CMA). Practitioners optimising for AI Overview citations should monitor whether these proceedings produce compliance changes to how Google attributes and displays source content, as publisher-liability rulings historically alter how platforms handle source attribution and opt-out mechanisms.
Google has introduced an embeddable 'Preferred Sources' button that allows readers to designate a publication as a preferred source directly from the publisher's own page; designated sources receive increased visibility in Top Stories, AI Overviews, and AI Mode for that user. This is a first-party mechanism connecting explicit reader preference signals to AI surface inclusion, bypassing the indirect signals — backlinks, engagement metrics, E-E-A-T — that have been the primary proxies for source authority to date. Publishers and GEO practitioners should implement the button immediately and monitor whether aggregate Preferred Source designations correlate with measurable citation-frequency changes in AI Overviews, as this could represent a new direct lever for AI visibility.
AI Lab Signals
OpenAI's updated Model Spec formalises the chain of command governing how its models balance usefulness, safety, and developer/user instructions, noting explicitly that production models do not yet fully reflect the spec but are being iteratively aligned toward it. For GEO practitioners, the spec is the closest public document to an editorial policy for ChatGPT: it governs when the model defers to operators, when it exercises independent judgement, and implicitly which content signals it is trained to treat as authoritative. Monitoring version changes to the spec — it is dated and versioned — is a lightweight way to track upstream shifts in model behaviour that may precede observable citation-pattern changes.
Google Releases Gemini 3.7 Flash, Improving Agent Web Retrieval and Code-Generation Accuracy
Gemini 3.7 Flash succeeds 3.6 Flash with gains in software engineering benchmarks (FrontierCode 1.1: 43.6% vs 34.4%; DeepSWE v1.1: 65.3% vs 49.0%) and improved first-pass code accuracy, launched at half the per-token cost of 3.6 Flash. The model is the default for Managed Agents in the Gemini API, meaning web-retrieval and agentic workflows built on the Gemini stack will now run on 3.7 Flash automatically. Practitioners building content pipelines or citation-monitoring tools on the Gemini API should retest retrieval fidelity and structured-data parsing behaviour after the model transition, as capability jumps of this magnitude can shift which content formats are preferred during synthesis.
(Pre-publication) The DataComp for VLMs benchmark, covering 160 datasets and 6 trillion multimodal tokens evaluated across models from 1B to 8B parameters, finds that instruction-heavy data mixtures scale better than caption-heavy ones and that data mixing strategy outweighs filtering as the primary driver of downstream task performance. While the research focuses on vision-language rather than text-only LLMs, its finding that mixture composition — not scale alone — determines model capability is consistent with the broader industry shift toward precision curation. GEO practitioners should note that multimodal content (images with structured captions, alt-text, and schema) is increasingly part of training pipelines and retrieval surfaces, not just text.
Training Data & Crawl
A U.S. district court largely denied the motion to dismiss Reddit's DMCA anti-circumvention claims against SerpApi and Perplexity, finding that CAPTCHA and bot-detection systems qualify as technological access-control measures under the DMCA even when the underlying content is accessible to human users. Reddit operates at over 100 million unique daily users and holds licensing agreements with AI companies for access to its corpus of conversational data, making this ruling directly relevant to teams whose training pipelines or retrieval systems bypass platform rate-limiting or bot-detection infrastructure. Practitioners advising on data-acquisition strategy for RAG or fine-tuning pipelines should treat this ruling as confirmation that circumventing access controls — even on publicly visible content — carries material legal exposure under current U.S. copyright law.
LLM Pretraining Pipeline Guide: Industry Shift From Quantity to Quality Curation Documented for 2026
A practitioner-facing pipeline guide documents the industry's confirmed move from brute-force web-scale ingestion toward precision curation, citing Apple's November 2024 Benchmark-Targeted Ranking research as evidence that cleaned, targeted datasets outperform raw-scale corpora on downstream task benchmarks. The guide notes that data quality — not architectural innovation — is now the primary competitive variable in pretraining, a framing consistent with the DataComp-VLM findings on mixture composition. For GEO practitioners, this trend reinforces the case for publishing clean, well-structured, machine-readable content: as labs apply stricter quality filters to training and retrieval corpora, content that passes automated quality gates is more likely to be retained in the data mix.
Practitioner Takeaway
Implement Google's 'Preferred Sources' publisher button on all client properties this month and instrument a citation-frequency baseline in AI Overviews before and after rollout: it is the first mechanism that creates a direct, auditable signal path from reader behaviour to AI surface inclusion, and the window to establish an early-mover baseline closes as adoption becomes standard practice.
The 6-phase framework used to structure this newsletter is available as a complete methodology guide — including audit tools, templates, and implementation checklists.
Get Access — free trial, then $19.99/mo or $200/yrNew to AI knowledge publication? Download the free briefing flyer — the data case for why your organisation cannot wait.