Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Evidence Should an AI Research Tool Provide About How It Works?

Researchers do not need every proprietary detail of an AI system, but they need enough evidence to judge whether it is fit for a particular research use. That includes intended use, sources, evaluation, limitations, data practices, and important system changes.

73
Evidence AI Research Tools Should Provide Guide 73 of 80
01 · The Question

How Much Should an AI Research Tool Tell You About Itself?

Researchers routinely ask methods questions. Where did the data come from? How was performance measured? What assumptions were made? What are the limitations? Under what conditions should the result not be generalized?

An AI tool that participates in research should not make those questions disappear.

Researchers do not necessarily need access to proprietary source code, model weights, or every detail of system architecture. They do need enough information to determine what the system is intended to do, what information it works from, how its performance has been evaluated, where its limits lie, what happens to research data, and whether important outputs can be traced and checked.

02 · The Short Answer

The Evidence Should Be Sufficient to Judge Fitness for Use

In Brief

An AI research tool should provide enough evidence for researchers to understand its intended uses, relevant data or source coverage, evaluation methods and results, known limitations and knowledge boundaries, data-handling practices, important provenance information, and material system changes that could affect performance.

The required level of evidence should be proportionate to the role of the tool. A system that merely helps brainstorm terminology can tolerate more uncertainty than one that screens studies, extracts data, classifies participants, analyzes evidence, or influences research conclusions.

03 · What You Need to Know

Transparency Should Help You Make a Research Decision

Start With Intended Use and Intended Users

A provider should explain what the system is designed to do and, equally important, what it is not designed to do.

NIST's AI Risk Management Framework calls for intended purposes, expected uses, context, assumptions, limitations, and prospective deployment settings to be understood and documented. It also calls for documentation of knowledge limits and how outputs may be used and overseen by humans.

For a research tool, useful documentation might therefore distinguish literature discovery from comprehensive searching, exploratory summarization from evidence synthesis, or coding assistance from autonomous statistical interpretation.

Without a defined use, performance claims become difficult to interpret. “Accurate AI for research” tells you very little because the research task remains unspecified.

The Tool Should Explain What Information It Works From

If a system answers questions using external information, researchers need to understand the relevant evidence base.

For a literature-oriented tool, that may mean identifying databases, indexes, repositories, publisher content, or other collections it can search. Researchers may also need to know whether the system has metadata, abstracts, full text, or another representation of the material.

Coverage information matters because every evidence source has boundaries. Disciplinary, geographic, linguistic, temporal, and document-type gaps can affect what the AI is capable of finding and therefore what conclusions its output can reasonably support.

The deeper question of whether an AI research tool should explain where its information comes from deserves particular attention when generated answers make factual or scholarly claims.

Training Data and Retrieved Sources Are Not the Same Thing

Researchers should distinguish information used to develop a model from information supplied to it at the time of a particular task.

A generative model may have learned statistical patterns from large training datasets, while a research application may separately retrieve current papers or process documents uploaded by the user. These mechanisms have different implications for provenance and verification.

If the provider cannot disclose a complete training dataset, that does not automatically make the tool unusable. But researchers should avoid assuming that a model's internal knowledge constitutes a transparent scholarly database. For source-dependent claims, inspectable retrieval and citation can provide a much stronger basis for verification.

Model training information Data used in developing or adapting the model, which may not be individually retrievable when an answer is generated.
Retrieved or supplied evidence Sources, documents, database records, or other information provided to the system for a particular query or task and potentially available for direct verification.

Performance Claims Need Methods, Not Just Percentages

If a provider says a tool is 95 percent accurate, researchers should immediately have several questions. Accurate at what? On which dataset? Compared with what reference? Using which version? Who evaluated the answers? How were ambiguous cases handled?

An evaluation result becomes meaningful only when its method is interpretable.

Useful documentation may describe the evaluation dataset, task definitions, reference standard, metrics, comparison systems, sample characteristics, prompting or configuration where relevant, and important limitations on generalization.

NIST's AI RMF specifically calls for testing, evaluation, verification, and validation considerations to be documented and for deployed systems to be demonstrated as valid and reliable, with limits on generalizability beyond tested conditions identified.

Results Should Include Failure Information

Average performance is rarely the whole story.

A research tool should disclose known weaknesses where they are material to use. Perhaps performance deteriorates on long tables, certain languages, older documents, particular disciplines, scanned PDFs, uncommon terminology, or questions requiring information unavailable in its sources.

NIST's framework calls for knowledge limits, assumptions, limitations, and risks to be documented, while its Playbook encourages documentation of testing used to identify errors and limitations.

This information can be more useful than another promotional benchmark. Knowing where a tool tends to fail tells researchers where additional verification is necessary.

The Tool Should Explain How Outputs Can Be Verified

For evidence-dependent research, transparency should extend to individual outputs where feasible.

If a system summarizes a paper, can you identify the paper? If it makes a claim, can you inspect the supporting source? If it extracts a value, can you locate the relevant passage or table? If it performs a calculation, can you inspect the inputs, procedure, or code?

Traceability does not make an output correct, but it makes correctness testable.

This is why verifiable citations can be particularly valuable in AI research tools. They give the researcher a path back to evidence rather than requiring trust in generated prose alone.

Provenance Should Be Available Where It Matters

Provenance concerns the origin and history of information or content. NIST's Generative AI Profile describes provenance mechanisms as ways to provide information about the origin and history of data and generated content, supporting assessments of authenticity and integrity.

For research tools, useful provenance can take several forms: identifying the source document behind an extraction, recording which model or system generated an output, indicating when an answer was produced, or preserving links between a generated synthesis and the evidence from which it was constructed.

Not every brainstorming response requires elaborate provenance metadata. The need becomes stronger when generated material enters datasets, analyses, evidence syntheses, or other parts of the research record.

Data Handling Is Part of How the Tool Works

Technical performance documentation is incomplete if researchers cannot determine what happens to the information they provide.

Relevant questions may include what data are collected, why they are processed, how long they are retained, whether they may be used to improve systems, whether third parties are involved, what controls users have, and what protections apply to sensitive material.

UNESCO's guidance specifically emphasizes data privacy and institutional validation in the adoption of generative AI for education and research. The European Commission's 2026 Living Guidelines likewise retain accountability, transparency, responsibility, and research integrity as central principles for responsible generative AI use in research.

The detailed legal and contractual questions belong in the tool's privacy policy and terms of service, but the information is still part of a researcher's overall assessment of the system.

The Provider Should Make Material Changes Visible

AI tools can change without researchers deliberately switching methods. A provider may update the underlying model, retrieval system, source coverage, interface, safety controls, file processing, or other components.

For casual use, those changes may be welcome and inconsequential. For an established research workflow, they can alter performance during a project.

Useful documentation should therefore make consequential changes reasonably discoverable. Version information, release notes, dated documentation, model identifiers, or other change records can help researchers understand whether the system used later in a study is meaningfully different from the one originally evaluated.

Independent Evidence Strengthens Provider Documentation

Providers necessarily know their systems better than outside researchers, so first-party documentation is essential. It is also evidence produced by an interested party.

Independent evaluations can test whether performance claims generalize beyond provider-selected datasets and conditions. Peer-reviewed studies, external benchmarks, reproducibility exercises, audits, or evaluations conducted by research groups may therefore add useful evidence.

The strongest basis for a consequential decision may combine provider documentation, independent evidence, and your own task-specific testing rather than relying exclusively on any one source.

Not Every Tool Needs to Reveal Its Source Code

Transparency is sometimes treated as all-or-nothing: either a system is completely open or it is a black box. Research decisions are usually more nuanced.

A proprietary tool can still disclose intended uses, limitations, source coverage, evaluation procedures, data practices, provenance, and change information. Conversely, access to source code alone does not tell a researcher whether a deployed system performs reliably on a particular task.

The relevant question is whether enough information is available to assess the risks and benefits of the proposed use.

Evidence or documentation What researchers should learn from it Why it matters
Intended-use documentation What the system is and is not designed to do Prevents inappropriate extrapolation to unsupported tasks
Source or data coverage What information the system can actually work from Reveals potential evidence gaps
Evaluation methods How performance claims were produced Makes accuracy claims interpretable
Evaluation results How the system performed under defined conditions Provides empirical evidence rather than marketing alone
Known limitations Where performance may deteriorate or fail Guides verification and appropriate boundaries
Output provenance Where important information or generated content came from Supports verification and accountability
Privacy and data documentation What happens to information provided by researchers Supports decisions about confidential and sensitive material
Version and change information Whether the system has materially changed Supports reproducibility and re-evaluation

The Higher the Stakes, the Stronger the Evidence Should Be

A tool used only to suggest synonyms can remain useful despite substantial uncertainty about its internal operation because researchers can inspect the suggestions directly and the consequences of error are modest.

A tool that autonomously screens studies, extracts outcome data, classifies sensitive material, or materially shapes research conclusions presents a different case. Its performance, limitations, provenance, and governance deserve much stronger evidence.

If the information needed for that assessment is unavailable, the problem is no longer merely imperfect documentation. The tool may be too opaque for the proposed research use.

04 · A Practical Example

Evaluating an Accuracy Claim From an AI Research Tool

Hypothetical Example

A Provider Advertises “95% Accurate Study Extraction”

A research platform says its AI can extract information from scientific papers with 95 percent accuracy. A research team is considering using the feature to populate fields for a review.

Ask what was measured The team looks for the fields included in the evaluation, the definition of a correct extraction, and whether missing and fabricated values counted as errors.
Inspect the evaluation material They determine which disciplines, paper formats, languages, and document types were represented and whether those conditions resemble their own review.
Check the reference standard The team looks for information about who established the correct answers and how disagreements or ambiguous reporting were handled.
Find the limitations Documentation reveals that performance is lower for information contained only in complex tables, which are common in the team's literature.
Run a local evaluation The researchers test representative papers from their own project and compare AI extractions with manually verified values.
Interpret the original claim properly The advertised figure becomes one useful piece of evidence, not a guarantee that 95 percent of the team's own extraction decisions will be correct.

The number was not necessarily misleading. It was simply incomplete without the methodological information required to interpret it. Researchers have encountered this problem before. Usually it arrives wearing a p-value rather than an AI logo.

05 · What Researchers Often Get Wrong

Common Misunderstandings About AI Transparency

Misconception

The Provider Must Reveal Everything About the Model

Complete technical openness is not necessary for every research use. The practical requirement is sufficient information to assess intended use, performance, evidence, limitations, data handling, and relevant risks. More consequential uses justify stronger transparency requirements.

Misconception

A Published Accuracy Percentage Is Enough Evidence

A percentage without task definitions, evaluation data, metrics, reference standards, and limitations can be difficult to interpret. Researchers need the method behind the number.

Misconception

If the Tool Provides Sources, Its Operation Is Transparent

Source traceability addresses one important part of transparency. It does not explain evaluation performance, coverage limitations, data handling, model changes, or other aspects relevant to research use.

Misconception

Open Source Automatically Means Trustworthy

Technical openness can enable inspection and independent evaluation, but trustworthiness still depends on the deployed system, data, performance, security, configuration, and intended use. Openness is useful evidence, not a substitute for evaluation.

Misconception

Researchers Can Compensate for Missing Documentation by Checking the Output

Output checking can manage some risks, particularly when errors are easy to detect. It cannot answer every question about evidence coverage, data handling, system changes, or hidden failure conditions. For consequential use, missing documentation can itself limit defensibility.

06 · What This Means for You

Ask for the Evidence You Would Need to Defend the Use

A simple decision framework

If the AI performs a low-risk, easily inspected task
Basic documentation plus direct researcher verification may be sufficient.
If the system makes source-dependent research claims
Require meaningful information about evidence sources, coverage, retrieval, and output traceability.
If the provider makes quantitative performance claims
Look for enough methodological detail to interpret and, where feasible, independently assess those claims.
If AI influences research data, analysis, screening, classification, or conclusions
Expect stronger evidence about validation, failure modes, limitations, provenance, and appropriate human oversight.
If essential evidence is unavailable
Reduce the tool's role, strengthen independent validation, or select a system that can be assessed more adequately.

One useful test is to imagine explaining the tool's role to a skeptical reviewer. Could you explain what it did, what evidence it worked from, why you considered it sufficiently reliable, what its known limitations were, and how you checked the output? If the answer is largely “the website said it was for researchers,” the evidentiary chain needs work.

07 · A Quick Checklist

What an AI Research Tool Should Let You Establish

Before relying on a research-facing AI tool, look for evidence about:
Purpose: What is the system designed to do, and what uses are outside its intended scope?
Sources: What data, documents, databases, or other evidence can it access for the task?
Coverage: What important disciplinary, linguistic, temporal, geographic, or document-type limitations apply?
Evaluation: How was performance tested, on what material, using which metrics and reference standards?
Failures: What known errors, knowledge limits, or conditions reduce performance?
Verification: Can consequential outputs be traced to inspectable evidence or reproducible procedures?
Data handling: What happens to prompts, files, research data, and other information provided to the service?
Versioning: Can you determine which system or version you used and whether important components have changed?
Independent evidence: Are relevant external evaluations available in addition to provider claims?
08 · Frequently Asked Questions

Questions About Evidence and Transparency in AI Research Tools

Does an AI company need to disclose its complete training dataset?

Complete disclosure may not always be available or necessary for every use. Researchers should nevertheless understand enough about relevant data, sources, coverage, limitations, and retrieval mechanisms to assess the proposed task, particularly when the AI makes evidence-dependent claims.

Should an AI research tool publish benchmark results?

Relevant evaluations can be useful, but the methodology matters as much as the score. Researchers need to know what was tested, on which data, with which metric and reference standard, and how closely those conditions resemble the intended use.

Is provider documentation enough to trust an AI tool?

Provider documentation is important because the provider has information outsiders may not possess. For consequential uses, independent evaluations and your own task-specific testing can provide stronger evidence about whether claims generalize to your context.

Should researchers know which AI model a tool uses?

Model identity can be useful for documentation, comparison, and detecting changes, but the complete deployed system matters too. Retrieval, document processing, databases, prompts, settings, and other components can substantially affect performance.

What if the provider changes models without clearly documenting it?

That uncertainty can affect reproducibility and the relevance of earlier evaluation. For consequential workflows, consider monitoring system documentation, recording dates and observable version information, and re-testing when a material change is suspected or confirmed.

Does providing citations make an AI tool transparent?

It improves one form of traceability when the citations are genuine and appropriately connected to claims. Full research assessment may also require information about coverage, evaluation, limitations, data handling, and system changes.

How much transparency is enough?

Enough to make a defensible decision about the proposed use. The threshold should rise with the sensitivity of the data, difficulty of detecting errors, and consequences the AI output could have for evidence, analysis, participants, or conclusions.

09 · The Bottom Line

A Research Tool Should Provide Evidence You Can Interrogate

The Bottom Line

An AI research tool should provide enough evidence about its purpose, information sources, evaluation, limitations, data practices, provenance, and material changes for researchers to judge whether it is fit for the role they intend to give it.

You do not need to understand every parameter inside the model. You do need enough visibility to ask ordinary research questions about evidence and method. As the AI's role becomes more consequential, “trust us” becomes progressively less adequate as technical documentation.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes