03 · What You Need to Know
Transparency Should Help You Make a Research Decision
Start With Intended Use and Intended Users
A provider should explain what the system is designed to do and, equally important, what it is not designed to do.
NIST's AI Risk Management Framework calls for intended purposes, expected uses, context, assumptions, limitations, and prospective deployment settings to be understood and documented. It also calls for documentation of knowledge limits and how outputs may be used and overseen by humans.
For a research tool, useful documentation might therefore distinguish literature discovery from comprehensive searching, exploratory summarization from evidence synthesis, or coding assistance from autonomous statistical interpretation.
Without a defined use, performance claims become difficult to interpret. “Accurate AI for research” tells you very little because the research task remains unspecified.
The Tool Should Explain What Information It Works From
If a system answers questions using external information, researchers need to understand the relevant evidence base.
For a literature-oriented tool, that may mean identifying databases, indexes, repositories, publisher content, or other collections it can search. Researchers may also need to know whether the system has metadata, abstracts, full text, or another representation of the material.
Coverage information matters because every evidence source has boundaries. Disciplinary, geographic, linguistic, temporal, and document-type gaps can affect what the AI is capable of finding and therefore what conclusions its output can reasonably support.
The deeper question of whether an AI research tool should explain where its information comes from deserves particular attention when generated answers make factual or scholarly claims.
Training Data and Retrieved Sources Are Not the Same Thing
Researchers should distinguish information used to develop a model from information supplied to it at the time of a particular task.
A generative model may have learned statistical patterns from large training datasets, while a research application may separately retrieve current papers or process documents uploaded by the user. These mechanisms have different implications for provenance and verification.
If the provider cannot disclose a complete training dataset, that does not automatically make the tool unusable. But researchers should avoid assuming that a model's internal knowledge constitutes a transparent scholarly database. For source-dependent claims, inspectable retrieval and citation can provide a much stronger basis for verification.
Model training information
Data used in developing or adapting the model, which may not be individually retrievable when an answer is generated.
Retrieved or supplied evidence
Sources, documents, database records, or other information provided to the system for a particular query or task and potentially available for direct verification.
Performance Claims Need Methods, Not Just Percentages
If a provider says a tool is 95 percent accurate, researchers should immediately have several questions. Accurate at what? On which dataset? Compared with what reference? Using which version? Who evaluated the answers? How were ambiguous cases handled?
An evaluation result becomes meaningful only when its method is interpretable.
Useful documentation may describe the evaluation dataset, task definitions, reference standard, metrics, comparison systems, sample characteristics, prompting or configuration where relevant, and important limitations on generalization.
NIST's AI RMF specifically calls for testing, evaluation, verification, and validation considerations to be documented and for deployed systems to be demonstrated as valid and reliable, with limits on generalizability beyond tested conditions identified.
Results Should Include Failure Information
Average performance is rarely the whole story.
A research tool should disclose known weaknesses where they are material to use. Perhaps performance deteriorates on long tables, certain languages, older documents, particular disciplines, scanned PDFs, uncommon terminology, or questions requiring information unavailable in its sources.
NIST's framework calls for knowledge limits, assumptions, limitations, and risks to be documented, while its Playbook encourages documentation of testing used to identify errors and limitations.
This information can be more useful than another promotional benchmark. Knowing where a tool tends to fail tells researchers where additional verification is necessary.
The Tool Should Explain How Outputs Can Be Verified
For evidence-dependent research, transparency should extend to individual outputs where feasible.
If a system summarizes a paper, can you identify the paper? If it makes a claim, can you inspect the supporting source? If it extracts a value, can you locate the relevant passage or table? If it performs a calculation, can you inspect the inputs, procedure, or code?
Traceability does not make an output correct, but it makes correctness testable.
This is why verifiable citations can be particularly valuable in AI research tools. They give the researcher a path back to evidence rather than requiring trust in generated prose alone.
Provenance Should Be Available Where It Matters
Provenance concerns the origin and history of information or content. NIST's Generative AI Profile describes provenance mechanisms as ways to provide information about the origin and history of data and generated content, supporting assessments of authenticity and integrity.
For research tools, useful provenance can take several forms: identifying the source document behind an extraction, recording which model or system generated an output, indicating when an answer was produced, or preserving links between a generated synthesis and the evidence from which it was constructed.
Not every brainstorming response requires elaborate provenance metadata. The need becomes stronger when generated material enters datasets, analyses, evidence syntheses, or other parts of the research record.
Data Handling Is Part of How the Tool Works
Technical performance documentation is incomplete if researchers cannot determine what happens to the information they provide.
Relevant questions may include what data are collected, why they are processed, how long they are retained, whether they may be used to improve systems, whether third parties are involved, what controls users have, and what protections apply to sensitive material.
UNESCO's guidance specifically emphasizes data privacy and institutional validation in the adoption of generative AI for education and research. The European Commission's 2026 Living Guidelines likewise retain accountability, transparency, responsibility, and research integrity as central principles for responsible generative AI use in research.
The detailed legal and contractual questions belong in the tool's privacy policy and terms of service, but the information is still part of a researcher's overall assessment of the system.
The Provider Should Make Material Changes Visible
AI tools can change without researchers deliberately switching methods. A provider may update the underlying model, retrieval system, source coverage, interface, safety controls, file processing, or other components.
For casual use, those changes may be welcome and inconsequential. For an established research workflow, they can alter performance during a project.
Useful documentation should therefore make consequential changes reasonably discoverable. Version information, release notes, dated documentation, model identifiers, or other change records can help researchers understand whether the system used later in a study is meaningfully different from the one originally evaluated.
Independent Evidence Strengthens Provider Documentation
Providers necessarily know their systems better than outside researchers, so first-party documentation is essential. It is also evidence produced by an interested party.
Independent evaluations can test whether performance claims generalize beyond provider-selected datasets and conditions. Peer-reviewed studies, external benchmarks, reproducibility exercises, audits, or evaluations conducted by research groups may therefore add useful evidence.
The strongest basis for a consequential decision may combine provider documentation, independent evidence, and your own task-specific testing rather than relying exclusively on any one source.
Not Every Tool Needs to Reveal Its Source Code
Transparency is sometimes treated as all-or-nothing: either a system is completely open or it is a black box. Research decisions are usually more nuanced.
A proprietary tool can still disclose intended uses, limitations, source coverage, evaluation procedures, data practices, provenance, and change information. Conversely, access to source code alone does not tell a researcher whether a deployed system performs reliably on a particular task.
The relevant question is whether enough information is available to assess the risks and benefits of the proposed use.
| Evidence or documentation |
What researchers should learn from it |
Why it matters |
| Intended-use documentation |
What the system is and is not designed to do |
Prevents inappropriate extrapolation to unsupported tasks |
| Source or data coverage |
What information the system can actually work from |
Reveals potential evidence gaps |
| Evaluation methods |
How performance claims were produced |
Makes accuracy claims interpretable |
| Evaluation results |
How the system performed under defined conditions |
Provides empirical evidence rather than marketing alone |
| Known limitations |
Where performance may deteriorate or fail |
Guides verification and appropriate boundaries |
| Output provenance |
Where important information or generated content came from |
Supports verification and accountability |
| Privacy and data documentation |
What happens to information provided by researchers |
Supports decisions about confidential and sensitive material |
| Version and change information |
Whether the system has materially changed |
Supports reproducibility and re-evaluation |
The Higher the Stakes, the Stronger the Evidence Should Be
A tool used only to suggest synonyms can remain useful despite substantial uncertainty about its internal operation because researchers can inspect the suggestions directly and the consequences of error are modest.
A tool that autonomously screens studies, extracts outcome data, classifies sensitive material, or materially shapes research conclusions presents a different case. Its performance, limitations, provenance, and governance deserve much stronger evidence.
If the information needed for that assessment is unavailable, the problem is no longer merely imperfect documentation. The tool may be too opaque for the proposed research use.