What to remember
- A visible citation starts the checking process. It does not complete it.
- Document retrieval improves some measured results but does not guarantee that every sentence is supported.
- A reference can exist without containing the stated evidence, or be accurate in a different context.
- Human responsibility remains when the answer informs analysis, recommendations or a decision file.
Showing a source is not enough to demonstrate a claim
A link changes the situation because the reader can attempt to check the answer. Yet control is real only when several successive conditions are met. A well-formatted reference may not exist. An existing page may be inaccessible. An accessible document may not contain the stated passage.
Do the source and its metadata match a real document?
Can the person responsible for checking consult the full text?
Is the relevant passage identified by page, section or verifiable extract?
Does the passage actually support the claim without reversal or overgeneralization?
Do the date, population, context and limits allow the intended use?
These five levels do not make an answer infallible. They identify where uncertainty lies: in the reference, the passage, the interpretation or its transfer to the decision at hand.
What the 2020 RAG study actually shows
Lewis et al. combine a generative model with an external memory made of Wikipedia passages. The system retrieves relevant documents before generating an answer. The authors describe this memory as inspectable and easier to update than knowledge held only in model parameters. The provenance problem is stated on the first page of the full scientific PDF.
On the FEVER dataset, the reference document appears first in 71% of reported cases and in the top ten in 90%. Retrieval is often useful but not systematic. On the evaluated tasks, RAG generations are also judged more factual and specific than those from the compared BART model. Results and examples appear on page 7 of the PDF.
The study supports the value of external documentary memory on specific tasks and datasets. It does not show that an enterprise assistant cites every statement correctly or that the first retrieved source is necessarily the best source for a given decision.
Why citation constraints are not sufficient
Davis and Mahmoud examine 36 academic-style documents produced by four systems and verify references through Crossref, OpenAlex and arXiv. Their study tests different citation instructions to determine whether stricter constraints reduce bibliographic errors.
The authors find that constraints change the distribution of errors without eliminating them. Their automated verification can establish that a reference exists, but not that the source supports the precise claim attached to it. This limit appears on pages 10 and 11 of the full scientific PDF.
The scope is bounded: prompts and models are fixed, each condition has three runs, and the sample concerns academic writing and software engineering with a limited number of systems. The full paper is also available on the PMLR proceedings page. It would be excessive to infer a universal error rate.
A 2023 EMNLP study separates two controls: how completely cited passages support the answer and how relevant each citation is. A response can therefore cite real documents while leaving claims unsupported or adding unnecessary references.
On the study’s ELI5 dataset, even the best systems tested lacked complete support for their citations in 50% of cases. This result belongs to that protocol and is not a universal rate. The authors also note that their automated evaluation depends on the accuracy of the model used to assess support. Method and limits are detailed on pages 1, 4 and 10 of the full PDF.
Mata v. Avianca: failed verification becomes failed responsibility
In Mata v. Avianca, lawyers filed a brief containing non-existent decisions and citations generated by an artificial-intelligence system. The sanctions order states that using such a system would not itself be improper. The issue was presenting unchecked material as real legal authority and failing to perform the professional verification expected.
The complete court order states the principle on page 1 and details the inadequate inquiry, generated responses and sanctions on pages 29 to 31. A full CBS News article provides a second factual account of the filing and its context.
The case concerns professional duties in a US judicial context. It does not define a general rule for every strategic decision. It does illustrate a transferable mechanism: an apparently documented output may be more misleading than an unsourced answer if an invented reference is granted authority it does not deserve.
Five checks before using a sourced answer
| Check | Question | Evidence to retain |
|---|---|---|
| Existence | Do the title, authors, organization and identifier match? | Direct URL, DOI or official identifier. |
| Direct access | Can the decision-maker open the full text without relying on a generated summary? | Accessible PDF or full article. |
| Location | Where exactly is the cited element? | Page, section, table or paragraph. |
| Support | Does the source say what the answer attributes to it? | Checked passage and faithful reformulation. |
| Scope | Are date, context and limits compatible with the decision? | Conditions of use and explicit caveats. |
A critical claim should be checked against the primary source whenever it is available. Important current events benefit from a second direct or independent source. If the full text is unavailable, the information should be labelled unverified rather than turned into certainty by fluent writing.
What a traceable documentary corpus provides, and where it stops
A documentary analysis can begin with explicit seed sources and a bounded scope. It can examine actors, sources, co-citations and themes while distinguishing calculated elements from estimates. Corpus provenance makes the documentary path more readable than a synthesis detached from its sources.
This scope does not mean that an assistant automatically verifies whether each passage supports every claim. It also does not demonstrate a RAG architecture or an autonomous decision system.
A bounded corpus with visible provenance supports documentary exploration. Judging source quality, fit with a claim and relevance to the decision remains a human operation.
What this analysis does and does not establish
- It establishes that a visible source is necessary for reproducible checking.
- It distinguishes document retrieval, bibliographic validity and actual support for a claim.
- It does not provide a universal reliability score for AI assistants.
- It does not assume that a direct link makes content true, current or relevant.
- It does not present visible sources as automatic verification or proof of a RAG architecture.
An answer useful for decision-making is not merely fluent and sourced. It must allow an identified person to retrieve the evidence, understand its scope, challenge the interpretation and decide whether the remaining uncertainty is acceptable.
Full sources
Scientific research
- Lewis et al. (2020), Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, full PDFProvenance and external memory: p. 1. Retrieval and factuality results: p. 7.
- Gao et al. (2023), Enabling Large Language Models to Generate Text with Citations, full PDFCitation completeness and relevance: p. 4. ELI5 result: p. 1. Evaluation limits: p. 10.
- Davis and Mahmoud (2026), Citation Constraints and Reference Hallucinations in Large Language Models, full PDFMethod: p. 1. Limits, validity and need for verification: pp. 10-11.
Real-world case
- United States District Court, Mata v. Avianca, complete sanctions orderPrinciple and facts: p. 1. Analysis and sanctions: pp. 29-31.
- CBS News, full report dated 29 May 2023Journalistic corroboration of the filing, non-existent references and hearing.