← All Editions

Signal — AI Knowledge: ASO & GEO Insights

Issue 14 — 07/07/2026

Published by Digital Human Assistants · aiknowledgesignal.io · Weekly practitioner briefing

This Week in Brief

Three large-scale empirical studies published in June–July 2026 converge on a single structural insight: AI engines have distinct, persistent source preferences that make a one-size-fits-all GEO strategy insufficient. Concurrently, a Tsinghua arXiv paper exposes a fundamental flaw in how LLM training data mixtures are optimised, with direct implications for which content domains get amplified at scale. Practitioners should audit both their per-engine citation presence and their content's structural fitness for RAG retrieval pipelines.

AEO Periodic Table V4: Brand Visibility Factors for AI Search — Study of 1.13 Million Prompts

Goodie · 23/06/2026

Per a study from Goodie, an analysis of 1.13 million prompts across ChatGPT, Claude, Perplexity, Grok, Gemini, and Google AI Mode finds that a brand's AI citation footprint is determined by signals outside its own domain — including Wikipedia, Reddit, YouTube, review sites, and forums — which collectively outweigh on-site content in shaping how AI answers represent a brand. The study identifies a structural gap in most organisations: SEO, PR, and social teams operate in silos with no shared analytics layer, leaving no single function with visibility over the full AI citation picture. Practitioners should treat off-site authority surfaces as first-class GEO assets, not secondary channels.

2026 AI Visibility Index: Analysis of 126 Million U.S. AI Search Prompts

Semrush · 26/06/2026

Per a study from Semrush, an expansion of its AI Visibility Index from 2,500 to 126 million U.S. AI search prompts (January–April 2026) provides one of the largest public datasets to date on how brands are mentioned, cited, and surfaced across major AI search platforms. The scale increase allows for statistically meaningful cross-platform comparisons of brand visibility patterns rather than directional estimates. Practitioners can use the underlying methodology as a benchmark framework when designing prompt-sampling strategies for their own AI visibility audits.

How AI Engines Choose and Cite Sources: A 7-Month Analysis Across 7 Engines

Conductor · 07/05/2026

Per a study from Conductor, tracking citation behaviour across ChatGPT, ChatGPT Search, Perplexity, Google AI Overviews, Google AI Mode, Gemini, and Claude from September 2025 through March 2026 (1,056 data points) reveals that each engine exhibits a persistent editorial identity: Perplexity defaults to YouTube across most intents, Google Gemini cites YouTube across every intent tracked, and only ChatGPT and ChatGPT Search surface Wikipedia. The data shows Google AI Mode uniquely routes users back to Google properties rather than third-party sources. A single unified AEO content strategy cannot efficiently cover the full AI search ecosystem; practitioners need per-engine optimisation playbooks differentiated by source type and intent.

Bing Webmaster Tools Adds Intents, Topics, Citation Share and Compare to AI Visibility Reporting

Bing Search Blog (Microsoft) · 16/06/2026

Microsoft expanded Bing Webmaster Tools' AI Performance report with four new capabilities in global preview: Intents (classifies grounding queries by purpose — Informational, Commercial, Navigational, Research, and more), Topics (clusters related grounding queries into thematic groups), Citation Share (the percentage of citations a site earns for a given grounding query, out of all citations shown for that query across all sites), and Compare (overlays a prior period against the current one). Microsoft describes Citation Share as an observational metric that does not expose competitor domains or represent traffic share. This is the first per-query, per-site GEO benchmark Bing has published directly through webmaster tooling; practitioners should use Topics and Intents to identify which query clusters their content already wins and where Citation Share reveals coverage gaps.

Seer Interactive: AIO CTR Decline Has Levelled Off — Full-Year 2025 Data Across 53 Brands and 5.47M Queries

Seer Interactive · 24/04/2026

Seer Interactive's third-edition AIO CTR study — covering 53 brands, 5.47 million tracked queries, and 2.43 billion organic impressions across full-year 2025 plus Q1 2026 actuals — finds that the expected continued CTR decline did not materialise: the downward trend has reversed direction. The study had previously modelled a continuing decline into 2026 based on Q4 2025 data, making the reversal a material update to planning assumptions. Practitioners who built 2026 traffic forecasts on a continuing decline model should revisit those projections with this dataset before Q3 budget reviews.

Ahrefs: ChatGPT Cites Only ~50% of Retrieved URLs — Title and Snippet Are the Primary Gatekeeping Signal

Ahrefs · 07/07/2026

Per Ahrefs research across 1.4 million prompts, ChatGPT retrieves dozens of URLs per query but ultimately cites approximately 50% of them; the selection decision occurs before the model reads page content, based on the page title, URL, and brief snippet returned at retrieval time. This confirms a gatekeeping layer upstream of content quality: a structurally weak title or uninformative meta snippet disqualifies a page from citation regardless of its substantive authority. Practitioners should audit title tags and meta descriptions specifically for snippet informativeness within the retrieval context, not solely for click-through optimisation.

Google: AI Overviews Now Appear on Nearly 65% of Question-Based Searches — Google Publishes Official Optimisation Guide

Seer Interactive (citing Google) · 28/05/2026

According to Seer Interactive's analysis of 8,500 stratified keywords across 30 industries, Google AI Overviews now appear on nearly 65% of question-based searches; Google has also formally published a guide to optimising for generative AI features on Google Search, explicitly affirming that the core SEO playbook remains the baseline. Seer's study — generated from 214,056 candidate keywords, filtered to 18,260 with verified search volume — examined which SEO factors correlate with earning the first-citation slot in AIOs. The first-citation position functions as the new position zero for informational queries; practitioners should verify that their highest-traffic informational pages satisfy both standard technical SEO and AI Overview citation criteria as defined in Google's published guidance.

Perplexity AI Faces Unresolved Publisher Disputes Including CNN Lawsuit and Cloudflare Crawler Allegations in 2026

Perplexity AI Magazine · 02/07/2026

A comparative analysis published by Perplexity AI Magazine notes that Perplexity's citation architecture coexists with unresolved publisher disputes, including a CNN lawsuit filed in 2026 and Cloudflare crawler allegations, flagging these as an active publisher risk category for brands relying on Perplexity citations as a GEO channel. Separately, the same analysis confirms that Google maintained more than 91% global search market share per StatCounter's June 2026 data, contextualising Perplexity's scale relative to the broader search landscape. GEO practitioners building citation strategies on Perplexity should monitor the legal landscape, as adverse rulings could alter the platform's crawling or syndication behaviour.

Google StatCounter Data: Google Holds 91%+ Global Search Market Share as of June 2026

Perplexity AI Magazine (citing StatCounter) · 02/07/2026

Per StatCounter data cited in Perplexity AI Magazine's July 2026 competitive analysis, Google continued to hold more than 91% of global search market share in June 2026, underscoring that AI answer engines remain supplementary surfaces rather than replacement channels at current adoption levels. The same analysis notes that 2026 AI Overview research identified unsupported or mismatched citations in a measurable share of generated answers, flagging citation accuracy as an unresolved quality issue. Practitioners should continue to treat traditional Google ranking as the primary channel while building GEO as an incremental visibility layer, and should audit AI-cited versions of their content for factual accuracy.

Being Crawled by an AI Bot Does Not Mean Your Content Enters Training Data — Segonzac Clarifies the Multi-Step Pipeline

David Epding (citing Segonzac) · 07/07/2026

A practitioner explainer summarising Segonzac's clarification documents that an AI crawler hit initiates a multi-step refinement process — crawl, text extraction, deduplication, filtering, data mix, tokenisation, and training — before any content can influence model behaviour; a crawler visit is a necessary but far from sufficient condition for inclusion. The post also corrects two persistent myths: that website updates propagate to a frozen deployed model (they do not — updates can only enter via the retrieval layer or a future training cycle), and that training data is stored as an HTML database (models are trained on extracted text tokens, not raw markup). Practitioners optimising for training-data inclusion should focus on surviving the filtering and deduplication stages — meaning high-information-density, low-duplicate content — not merely on ensuring crawler accessibility.

CausalMix: Data Mixture as Causal Inference for Language Model Training

Tsinghua University researchers et al. · https://www.techtimes.com/articles/319548/20260702/llm-data-mixture-breaks-when-training-pools-shift-causal-inference-offers-fix.htm · 02/07/2026

(Pre-publication / arXiv) A paper posted to arXiv on 1 July 2026 by Tsinghua University researchers identifies a structural flaw in standard LLM data mixture optimisation: proxy experiments used to find optimal domain ratios assume a fixed data pool, an assumption that breaks in production environments where the pool changes continuously. The proposed fix — treating data mixture selection as a causal inference problem — is designed to remain valid as the underlying data pool evolves. For GEO practitioners, this research signals that the domain-level weighting decisions baked into future model training runs will be less stable than previously assumed, making content positioning across multiple topical domains a more robust strategy than concentrating authority in a single niche.

DuoFlow-KG: A Dual-Modal Evidence Retrieval Framework for High-Density LLM-Augmented Knowledge Graph Question Answering

DuoFlow-KG authors et al. · https://www.nature.com/articles/s41598-026-59526-3 · 07/07/2026

Published in Scientific Reports, this paper proposes a retrieval framework that jointly optimises semantic relevance and structural dependencies in knowledge graph question answering by constructing compact, high-density evidence subgraphs — addressing the failure of existing methods in multi-hop reasoning scenarios where fragmented retrieval or search-space explosion degrades answer quality. The dual-directional knowledge anchoring strategy enriches entity representations by incorporating both incoming and outgoing relational neighbourhoods. For GEO practitioners, the paper reinforces the practical value of building content with explicit entity relationships and structured semantic context: content that mirrors knowledge-graph-style entity connectivity is better positioned to survive the evidence-selection stage in LLM-augmented retrieval systems.

Practitioner Takeaway

Audit your AI citation strategy at the per-engine level before scaling content volume. Conductor's 7-month citation study confirms that Perplexity, Gemini, ChatGPT, and Google AI Mode each favour structurally different source types — YouTube dominates Gemini and Perplexity; Wikipedia is exclusive to ChatGPT surfaces; Google AI Mode self-references. Separately, Ahrefs' 1.4-million-prompt study confirms that the citation gatekeeping decision happens at the title-and-snippet layer, before your content is read. This week's combined signal: map which engine is highest-priority for your category, identify the source types that engine consistently favours, ensure your content appears on those surfaces (not just your own domain), and tighten every title tag and meta description for snippet informativeness rather than keyword density alone.

Sources This Edition

  1. https://higoodie.com/blog/aeo-periodic-table-v4/
  2. https://www.semrush.com/news/463141-semrush-releases-expanded-2026-ai-visibility-index-analyzing-126-million-ai-search-prompts/
  3. https://www.conductor.com/academy/how-ai-citations-differ/
  4. https://www.seerinteractive.com/insights/aio-impact-on-google-ctr-2026-update
  5. https://ahrefs.com/blog/why-chatgpt-cites-pages/
  6. https://www.seerinteractive.com/insights/what-it-takes-to-rank-in-googles-ai-overviews-in-2026-is-not-what-you-think
  7. https://perplexityaimagazine.com/perplexity-hub/perplexity-vs-bing-copilot-research-productivity/
  8. https://perplexityaimagazine.com/perplexity-hub/perplexity-vs-google-research-or-reach/
  9. https://www.david-epding.de/post/segonzac-llm-training-data-vs-llm-bot-crawls-layers-of-ai-chat-answers
  10. https://www.techtimes.com/articles/319548/20260702/llm-data-mixture-breaks-when-training-pools-shift-causal-inference-offers-fix.htm
  11. https://www.nature.com/articles/s41598-026-59526-3
  12. https://blogs.bing.com/search/June-2026/New-AI-Visibility-Insights-in-Bing-Webmaster-Tools-Intents-Topics-Citation-Share-Compare

Get the full AI Knowledge Signal Publication Framework

The 6-phase framework used to structure this newsletter is available as a complete methodology guide — including audit tools, templates, and implementation checklists.

Get Access — free trial, then $19.99/mo or $200/yr

New to AI knowledge publication? Download the free briefing flyer — the data case for why your organisation cannot wait.