01 · The Question
Can an AI Give You a Citation to a Paper That Does Not Exist?
You ask generative AI for studies supporting a claim. Within seconds, it returns a tidy reference list containing authors, article titles, journal names, publication years, volume and issue numbers, page ranges, and perhaps DOIs.
The references look considerably more scholarly than “I don't know.” Unfortunately, appearance is not proof of existence.
Generative AI can produce citations to papers, books, chapters, conference papers, and other sources that cannot be verified because the works do not exist as described. It can also cite a publication that genuinely exists while getting some of its bibliographic information wrong.
For researchers, these are related but distinct problems. One requires determining whether the source exists at all. The other requires determining whether the citation accurately identifies the real source.
03 · What You Need to Know
What Does an AI-Hallucinated Citation Look Like?
A Hallucinated Citation Can Look Completely Ordinary
Academic references have predictable structures. Depending on the source type and citation style, a journal reference may contain authors, a publication year, article title, journal title, volume, issue, pages or article number, and a DOI.
Language models are very capable of reproducing that structure.
Consequently, a fabricated reference need not look absurd. It may have realistic author names, a title appropriate to those authors' research interests, a genuine journal, plausible volume information, and formatting that conforms neatly to APA or another citation style.
This is one reason AI hallucinations can be so persuasive. Bibliographic realism and bibliographic authenticity are different properties.
Looks like a citation
The entry follows the linguistic and formatting conventions of scholarly references.
Is a verifiable citation
The entry identifies an actual source whose bibliographic record can be independently established.
A Reference Can Fail in Several Different Ways
“Fake citation” is useful shorthand, but it can conceal several distinct error patterns.
| Reference problem |
What it means |
| Entirely fabricated source |
The cited work does not appear to exist as a published or otherwise publicly disseminated source |
| Real journal, nonexistent article |
The journal exists, but the particular article does not |
| Real authors, nonexistent article |
The named researchers exist, perhaps even in the relevant field, but did not publish the cited work |
| Real paper, incorrect metadata |
The work exists, but one or more authors, title elements, dates, journal details, volume, issue, or pages are wrong |
| Real paper, wrong identifier |
The reference is genuine but the supplied DOI or other identifier does not correspond to it |
| Real paper, wrong attribution |
The citation is genuine but does not support the claim attached to it |
The last problem extends beyond bibliographic accuracy into substantive source use. An AI can misrepresent what a real research paper actually reports, so confirming existence is only the first stage of verification.
Fabricated Citations Have Been Demonstrated Empirically
This problem is not based merely on anecdotes.
Walters and Wilder systematically evaluated 636 bibliographic citations generated in short literature reviews by GPT-3.5 and GPT-4. In their experiment, 55% of GPT-3.5 citations and 18% of GPT-4 citations were fabricated. Among citations referring to real works, 43% of those generated by GPT-3.5 and 24% of those generated by GPT-4 contained substantive citation errors.
Those percentages describe particular model versions, prompts, and experimental conditions from 2023. They should not be treated as universal hallucination rates for generative AI, especially as systems continue to change. Their importance lies in demonstrating two separate failure modes: models can generate nonexistent references, and they can incorrectly describe genuine ones.
Research in systematic-review contexts has found similar concerns. Chelli and colleagues evaluated 471 references generated by GPT-3.5, GPT-4, and Bard under their study conditions and reported substantial hallucination rates, while also finding poor precision and recall relative to human-conducted systematic reviews. The authors concluded that references generated by these systems required thorough researcher validation.
Hallucinated References Often Contain Real Components
One reason fabricated references can be difficult to detect is that they need not be composed entirely of fictional elements.
Walters and Wilder observed that many fabricated citations in their study included real journals, publishers, or organizations. More recent research has likewise characterized hallucinated citations as pattern-driven constructions that can combine real authors, journals, and topical language into structurally plausible but nonexistent references.
Imagine a researcher who genuinely publishes on educational technology, a real journal that regularly publishes work on AI in education, and a plausible article title matching both. Combining those three real-looking pieces does not establish that the resulting paper exists.
This is why checking only the journal or author is insufficient.
Real References Can Contain Hallucinated Metadata
Suppose you search the title and discover that the paper actually exists. Problem solved?
Not quite.
The generated citation may contain the wrong year, omit an author, add someone who was not an author, alter the title, provide the wrong journal, or supply incorrect volume, issue, or page information.
Walters and Wilder found substantive errors among references to genuine works, including errors in author names, article titles, dates, journal titles, volume or issue information, page numbers, and publishers.
This distinction matters because reference-management software and citation styles can correct formatting, but they cannot reliably transform an invented bibliographic record into the right publication if the underlying metadata are wrong.
Watch Out
Do not treat “I found something similar” as verification. Confirm that the authors, title, publication venue, year, and identifier all refer to the same source rather than quietly repairing an AI-generated citation by combining details from different works.
A DOI Does Not Automatically Make the Citation Real
A DOI looks authoritative because it is a persistent identifier associated with scholarly objects. But a generative AI system can produce a DOI-shaped string whether or not that identifier belongs to the cited paper.
Sometimes the DOI will resolve to nothing. More confusingly, an incorrect DOI may resolve to a completely different publication.
The existence of a DOI-like string therefore does not verify the citation. Resolve it and compare the resulting record with the authors, title, journal, and other details you were given.
The specific problem of AI-generated fake or incorrect DOIs deserves its own verification process because an identifier is useful only when it actually identifies the claimed work.
A Real Citation Still Does Not Prove the Claim
This is the verification step most likely to be skipped after the relief of discovering that a paper exists.
A citation performs an evidentiary function. It is not enough that the source is genuine. The source must support the proposition for which you cite it.
Step 1: Existence
Does this publication actually exist?
Step 2: Identity
Are the bibliographic details and identifier correct?
Step 3: Relevance
Is this actually the source intended, rather than merely a similarly titled publication?
Step 4: Support
Does the source provide evidence for the claim you are attaching to it?
A perfectly accurate APA reference can still be an inappropriate citation if the paper says something different from the generated claim.
Asking the Same AI to Verify Its Citation Is Not Independent Verification
A common response to a suspicious reference is to ask, “Are you sure this paper exists?”
The model may correct itself. It may also confidently affirm the original fabrication, alter some details while preserving the nonexistent source, or generate a replacement citation that requires verification of its own.
Walters and Wilder specifically noted prior evidence that ChatGPT could provide inaccurate responses when asked to verify the legitimacy of works it cited.
The solution is simple: move outside the unsupported generation. Search an appropriate scholarly source.
Where Should You Verify a Citation?
The best verification source depends on the publication type and discipline. Useful places include the publisher's official article page, Crossref, PubMed, discipline-specific bibliographic databases, library discovery systems, and other authoritative scholarly indexes.
Google Scholar can be useful for discovery, but verification should ideally reach the actual publication record when possible.
For a DOI, resolve or search the identifier and make sure the metadata match. For an article, confirm the journal and publication record. For a book chapter, verify both the chapter and the containing book. For a conference paper, identify the actual proceedings or official record.
If repeated searches produce no trace of the exact work, do not cite it merely because the AI insists that it exists.