Answer in brief

AI chatbots do not publish one stable, universal formula for choosing citations. The defensible way to study source selection is to observe a fixed set of prompts by model and date, record the links shown, and verify whether each linked page supports the claim beside it.

A citation is not the whole source-selection process

A page can be retrieved as a candidate, used to ground part of an answer, or displayed as a source. Those are related events, but they are not interchangeable. A link shown beside an answer proves that the system displayed that reference; it does not expose every candidate considered or every passage used during generation.

OpenAI describes ChatGPT search as returning timely answers with links to relevant web sources and a Sources panel. Google similarly describes AI Overviews as pairing synthesized answers with links for further exploration. Neither description supplies a universal public ranking formula that publishers can simply implement.

Separate observations from hypotheses

The observable layer includes the exact prompt, product, date, answer text, cited URLs, citation position, and whether the linked page supports the nearby claim. These facts can be stored, reviewed, and compared without guessing about hidden system behavior.

Query interpretation, candidate retrieval, reliability scoring, passage extraction, and citation-display decisions are useful conceptual stages. Their precise implementation and relative weight are normally proprietary. Treat proposed ranking factors as hypotheses to test, not as established rules.

Measure a fixed prompt panel by model and date

Start with a stable panel of prompts grouped by audience, intent, topic, and decision stage. Run the same panel across the products you care about on scheduled dates. Preserve the full answer, every visible source URL, the model or product label, and any visible search setting.

Keep model-level results separate. A blended citation rate can conceal the fact that one product cites an owned page while another relies on third-party coverage for the same question. Repeated dated observations also help distinguish normal answer variation from a durable change.

Track the page and the claim, not only the domain

Brand mentions answer whether an entity appears. Linked citations answer which page received visible attribution. Direct-support review answers whether that page actually substantiates the associated claim. Each is a different metric and should have its own field.

Normalize URL variants, map citations to page topics, and compare cited pages with uncited pages. Look for differences in answer clarity, explicit definitions, dates, primary-source links, information structure, entity coverage, and the specificity of evidence. These comparisons produce testable publishing hypotheses.

Make evidence easier to extract and verify

Useful experiments include a concise answer near the top, descriptive headings, semantic HTML, dated claims, data tables, visible authorship, primary-source links, and clear definitions. These choices improve inspectability for people and machines, but none guarantees a citation.

Change one meaningful element at a time, retain a control set of prompts, and measure several scheduled runs before attributing an outcome to the edit. Citation behavior is an observed result under stated conditions, not a universal ranking score.

Data behind the finding

An operational model for studying source selection without claiming access to proprietary ranking systems
MeasureResultContext
RetrievedCandidate sourceA system may find a page without using or displaying it; usually not externally observable
GroundedEvidence usedA page may inform an answer without appearing as a visible link; often only partly observable
DisplayedVisible referenceA URL, source card, or unambiguous reference shown with the answer; directly observable
Methodology note

AthenaHQ's content workflow generated the initial brief and draft from four tracked prompts about source selection, citation optimization, citation-rate measurement, and page attribution. The published version was edited to remove promotional claims, replace secondary guidance with primary sources where available, and label observations separately from inferences.

How to cite this article

A canonical source for this finding

Plum, Jenna. “How do AI chatbots choose sources to cite?” The Answer Signal, August 13, 2026.

https://theanswersignal.com/blog/how-ai-chatbots-choose-sources-to-cite

Sources and related research

  1. Introducing ChatGPT searchOpenAI. Primary product documentation for web answers, source links, and the Sources panel
  2. Generative AI in Search: Let Google do the searching for youGoogle. Primary product context for AI Overviews and links to supporting websites
  3. News Source Citing Patterns in AI Search SystemsKai-Cheng Yang, arXiv. An empirical study of more than 366,000 citations across OpenAI, Perplexity, and Google systems
Data disclosure: This article uses aggregated AthenaHQ data and follows our published methodology. Research claims are reviewed against the cited source material before publication.