03 · What You Need to Know
A Plausible Reconstruction Is Not the Same as a Source-Grounded Description
First Ask What “Has Not Accessed” Actually Means
There is an important difference between not having the full paper and having no information about it whatsoever.
An AI may lack the article's full text while still having access to its title, authors, abstract, DOI, keywords, citation context, press coverage, review articles, or other papers that discuss its findings. Some of those sources can support legitimate statements about the study.
The problem begins when the generated answer does not distinguish information actually supported by those sources from details inferred beyond them.
| Information available to AI |
Reasonable description |
Higher-risk claims |
| Title only |
Topic suggested by the title |
Sample, methods, results, effect sizes, conclusions |
| Metadata |
Authors, journal, year, DOI, publication details |
Substantive study findings not present in the metadata |
| Abstract |
Purpose and findings explicitly reported in the abstract |
Unreported methodological details, secondary analyses, tables, nuanced limitations |
| Secondary source |
What that secondary source says about the study |
Claims presented as though independently verified from the original paper |
| Full paper |
Claims grounded in the accessible article |
Anything the article itself does not support |
Titles Can Make Missing Details Surprisingly Predictable
Research titles often reveal a great deal. A title might identify the population, variables, intervention, study design, location, or broad conclusion. If the topic belongs to a familiar literature, an LLM can also draw on learned patterns about how such studies are commonly conducted.
Suppose a title mentions a randomized trial of an educational intervention and mathematics achievement. A model might reasonably infer that the study compared intervention and control conditions and measured some form of mathematics performance.
But “reasonable inference” is not the same as “reported method.” The actual study may use cluster randomization, an active comparator, delayed treatment, several outcomes, unusual eligibility criteria, or an analytical strategy the title does not reveal.
A model's strength at completing plausible patterns becomes a liability when the user mistakes those completions for document retrieval.
An Abstract Can Support a Partial Description, Not a Full Reconstruction
If the abstract is available, the AI has accessed part of the paper's published content. It can reasonably describe information explicitly reported there.
Abstracts, however, are compressed representations. They may omit details about recruitment, measurement, missing data, statistical models, secondary outcomes, sensitivity analyses, adverse events, limitations, and numerous other features.
Research evaluating AI-generated summaries of abstracts has found that LLMs can produce useful and highly rated summaries under controlled conditions. That supports using available abstracts for appropriate abstract-level tasks. It does not justify filling absent sections of the paper through inference.
Related Literature Can Make a Fabricated Detail Look Convincing
Scientific studies rarely exist in isolation. Papers within the same field often use similar designs, instruments, terminology, analytical procedures, and theoretical frameworks.
An LLM may therefore generate a method that is entirely plausible for the research area. It may even be the method most researchers would expect the authors to have used.
That can make the error difficult to notice. A bizarre fabrication is easy to challenge. A perfectly reasonable but nonexistent methodological detail may survive into notes, presentations, literature reviews, or manuscripts precisely because nothing about it looks suspicious.
The Model May Have Encountered Information About the Paper During Training
An AI can sometimes provide correct information about a paper without retrieving it during the current conversation because relevant information may have been represented in its training data.
The original paper itself might have been present in some form, or its findings may have appeared in reviews, citations, webpages, teaching materials, news coverage, or other documents.
Researchers usually cannot infer the exact provenance of a generated statement from the response itself. Consequently, a correct answer does not prove direct paper access, while lack of direct access does not prove that every generated statement must be wrong.
This is another reason to distinguish whether generative AI actually “knows” an answer from whether that answer has been verified.
Hallucination Makes Unsupported Reconstruction Especially Risky
Large language models can produce fluent false or unsupported outputs. Scientific literature tasks are particularly sensitive to this problem because titles, authors, journals, methods, and results often follow recognizable patterns.
In the OpenScholar evaluation published in Nature in 2026, a general-purpose model without the specialized retrieval system fabricated paper titles at high rates when asked to answer scientific literature questions with citations. Retrieval and attribution mechanisms substantially improved performance.
The exact rates from one benchmark should not be generalized to every model or every prompt. The result nevertheless demonstrates the underlying problem: when a system is asked to supply literature-specific information without adequate retrieval, it may generate plausible scholarly content rather than abstain.
A Correct Citation Does Not Validate the Description
Researchers often verify that a cited paper exists and stop there. That catches fabricated references but not fabricated claims about real references.
A genuine paper can be paired with an inaccurate description. The AI may misstate the sample size, study design, direction of an effect, statistical significance, population, outcome, or conclusion while providing a perfectly valid DOI.
Verification therefore needs two stages: Does the source exist? and Does the source support what the AI says about it?
Summarization Can Introduce Errors Even When the Source Is Available
Direct access does not create a perfect safeguard. Scientific summarization itself can alter meaning.
A 2025 study comparing thousands of LLM-generated summaries with original scientific texts found that models tended to omit details limiting the generalizability of findings, resulting in broader claims than the sources warranted. This is a useful reminder that source access is necessary for source-grounded analysis but not sufficient for accurate interpretation.
If errors can occur when the text is available, confidence should be lower when the relevant text was never available in the first place.
Paper-Specific Questions Require Paper-Specific Evidence
There is a practical difference between asking, “What does research generally say about this topic?” and “What did Garcia et al. report in Table 3?”
The first question may be answerable from broader knowledge and multiple sources. The second requires access to information specific to that paper.
As your question becomes more document-specific, the need for direct document access increases. Exact sample sizes, exclusion criteria, coefficients, confidence intervals, instrument wording, subgroup analyses, and limitations should not be reconstructed from general domain knowledge.
Paywalls Make This Boundary Particularly Easy to Miss
A system may correctly identify a subscription-only article from public metadata and then continue generating as though the entire article were available. Researchers should therefore establish whether the AI actually accessed the paywalled paper or only information surrounding it.
That provenance question is more useful than asking the model whether it has “read” the paper, especially when the interface does not provide traceable retrieval evidence.
Watch Out
Never fill missing sections of a paper by asking AI to “infer what the authors probably did” and then record the result as though it came from the article. Plausible reconstruction may be useful for generating questions to check, but it is not data extraction.