This Week in Brief
A German court ruling holds Google liable for AI Overview outputs under media law, while a U.S. district court finds Perplexity faces DMCA anti-circumvention claims for scraping Reddit content — two legal signals that materially change the compliance and citation landscape for AI search platforms. Meanwhile, Goodie's 31-million-citation study of news publishers finds that blocking AI crawlers is ineffective against the engines (Grok, AI Overviews, DeepSeek) that generate roughly half of all news citations.
Market Analysis — GEO & ASO
Semrush 2026 AI Visibility Index: 126 Million U.S. AI Search Prompts Analyzed
Per a study from Semrush, analysis of 126 million U.S. AI search prompts across January–April 2026 maps how brands are mentioned, cited, and surfaced across major AI search platforms — an expansion from the 2,500-prompt baseline of the September 2025 edition. The dataset offers practitioners one of the larger empirical views of AI-powered brand discovery to date. Teams benchmarking citation share across platforms should note the scale of the underlying prompt corpus when comparing this index to smaller-sample studies.
Goodie AEO Periodic Table V4: Findings from 1.13 Million Prompts Across Six AI Engines
Per a framework study from Goodie, analysis of 1.13 million prompts across ChatGPT, Claude, Perplexity, Grok, Gemini, and Google AI Mode finds that a brand's AI visibility is shaped as much by third-party surfaces — Wikipedia, Reddit, review sites, news coverage — as by its own web content. The study argues that the gap between SEO, PR, and social ownership within most organisations leaves no single team with a complete view of how a brand appears in AI answers. Practitioners building AEO programmes should assess whether their measurement layer consolidates owned, earned, and third-party signals.
Goodie: Blocking AI Crawlers Fails Against the Engines That Drive Half of All News Citations
Per a study from Goodie, 31 million AI citations across 11 AI surfaces and a robots.txt audit of 105 news publishers found that Grok, AI Overviews, and DeepSeek together account for roughly half of all news citations in the sample and do not honour robots.txt blocks — meaning that publisher access policy built around crawler blocking addresses only a subset of the citation ecosystem. For GEO practitioners advising publishers or content-heavy brands, the finding indicates that robots.txt alone is an insufficient access control strategy and that licensing and legal levers must be evaluated separately by engine.
AI Search & ASO
A Munich court has ruled that Google AI Overview outputs are subject to media liability law, a legal classification that treats generative search answers as publisher output rather than neutral intermediary retrieval. The ruling — analysed by media law specialists at MediaLaws — has direct implications for how AI Overviews must handle factual accuracy, correction obligations, and potential liability for harmful or misleading summaries. For GEO practitioners, the decision introduces a new compliance dimension: brands cited inaccurately in AI Overviews in German-jurisdiction markets may now have clearer legal routes to correction, and Google may face pressure to adjust how AI Overview content is generated or attributed in the EU.
Per a seven-month citation-behaviour analysis from Conductor tracking ChatGPT, ChatGPT Search, Perplexity, Google AI Overviews, Google AI Mode, Gemini, and Claude from September 2025 through March 2026 (1,056 data points), each engine demonstrates a persistent editorial identity — Perplexity and Gemini favour YouTube across most intents, ChatGPT and ChatGPT Search are the only engines that regularly surface Wikipedia, and Google AI Mode routes users back to Google properties. The data indicate that a single unified AEO content strategy cannot achieve consistent citation across all engines simultaneously, and that per-engine optimisation is necessary.
AI Lab Signals
Google Launches Gemini 3.7 Flash, Cutting Per-Token Price by Half
Google released Gemini 3.7 Flash on 13 August 2026 at half the per-million-token price of Gemini 3.6 Flash, with reported gains in coding, web development, and knowledge-work tasks. The model is positioned as the default workhorse for coding agents and complex workflows. For GEO practitioners, Gemini 3.7 Flash's improved web development and reasoning capabilities may affect how Gemini-surface AI answers are assembled and which content signals the model weights when synthesising citations.
On 24 July 2026, Justice Amit Bansal of the Delhi High Court dismissed Asian News International's application for an interim injunction against OpenAI in India's first copyright infringement action against an LLM developer, addressing territorial jurisdiction, training-data storage, output generation, and the fair-dealing exception under the Copyright Act 1957. The 135-page judgment does not grant ANI the relief sought at this stage, but the suit continues and the court's reasoning on training-data storage and output generation sets an early Indian precedent. Practitioners operating in South Asian markets should note that the legal perimeter around AI training on news content remains actively litigated and unsettled.
MaLA Corpus: 74-Billion-Token Multilingual Dataset Covering 546 Languages Released at COLM 2026
Researchers from multiple institutions presented MaLA, a 74-billion-token open-access corpus engineered for continual pre-training across 546 languages, at COLM 2026. The dataset applies deliberate upsampling of low-resource languages and downsampling of high-resource ones to reduce forgetting in code generation while expanding language coverage. For GEO practitioners targeting non-English markets, MaLA's design signals that the next generation of multilingual LLMs will have meaningfully broader language coverage — potentially improving citation retrieval in markets previously underserved by training data.
Training Data & Crawl
The Southern District of New York largely denied motions to dismiss Reddit's DMCA anti-circumvention claims against SerpApi and Perplexity, holding that CAPTCHA and bot-detection systems preventing automated scraping of Reddit content in Google search results can qualify as technological access control measures under the DMCA — even where the same content is available to human users. Reddit holds licensing agreements with Google and OpenAI for its content; it has not extended equivalent rights to SerpApi or Perplexity. Practitioners building AI training pipelines or retrieval-augmented systems that ingest third-party web content should treat this ruling as a material signal that bot-detection bypass carries DMCA anti-circumvention exposure, not merely terms-of-service risk.
A technical explainer notes that a corpus has no fixed token count until a specific tokenizer is named: the same text can yield figures that differ by 40% or more depending on vocabulary size, meaning published training-data size claims are not directly comparable across labs or papers. Byte and document counts introduce further ambiguity from encoding choices and document length variance. GEO practitioners citing corpus size figures in competitive research or client reporting should specify the tokenizer and whether byte counts are compressed or uncompressed to avoid propagating misleading comparisons.
Practitioner Takeaway
Audit your AI crawler access policy by engine, not by a single robots.txt rule. This week's Goodie study confirms that Grok, AI Overviews, and DeepSeek — together responsible for roughly half of all news citations in a 31-million-citation sample — do not honour robots.txt blocks. Separately, the Munich ruling and the Reddit v. SerpApi DMCA decision signal that legal and compliance exposure around AI content use is hardening. Run a three-column audit: (1) which engines are citing your content, (2) which honoured your robots.txt if you issued one, and (3) whether your content is being ingested via scraping rather than a licensed or direct-index pathway — because the latter now carries documented DMCA risk.
The 6-phase framework used to structure this newsletter is available as a complete methodology guide — including audit tools, templates, and implementation checklists.
Get Access — free trial, then $19.99/mo or $200/yrNew to AI knowledge publication? Download the free briefing flyer — the data case for why your organisation cannot wait.