03 · What You Need to Know
Evaluate the Use Case, Not AI in the Abstract
Define Exactly What You Want the Tool to Do
An evaluation cannot begin with “Is this AI accurate?” Accuracy at what?
Define the intended use narrowly enough that performance can be examined. “Help with my literature review” is still broad. “Identify potentially relevant papers from a defined query,” “extract sample size and study design from included papers,” and “summarize the limitations reported by each study” are different tasks and should be evaluated separately.
This follows a broader principle in NIST's AI Risk Management Framework: trustworthy characteristics and appropriate measurements depend on context. A system may be suitable for one use and unsuitable for another.
If you have not yet identified which system fits the task, first narrow the candidate tools according to your research requirements. Evaluation begins once you have something specific to test.
Decide What Successful Performance Looks Like
Before testing the tool, define the outcome you need. Otherwise, an impressive demonstration can quietly become the evaluation criterion.
The relevant measure depends on the task. For extraction, you may care about correctly identified values and omissions. For literature discovery, you may care about relevant retrieval, source authenticity, and coverage. For summarization, you may examine whether central findings, methods, and limitations are represented faithfully. For coding, you may test whether the code implements the intended analysis and produces validated results.
Do not rely only on whether an answer “looks good.” Convert quality into something you can inspect.
Build Representative Test Cases
A useful evaluation should resemble the conditions under which the system will actually be used.
If your project contains long papers with tables, do not test only short abstracts. If the literature contains specialized terminology, include it. If documents vary in format, include that variation. If the project is multilingual, test the relevant languages. If edge cases matter, deliberately include some.
The test set does not need to become a publication-quality benchmark for every low-risk use. It does need to be sufficiently representative that success on the test means something for the planned workflow.
Establish a Reference Where Possible
You need something against which to judge the AI output. For many research tasks, that means creating or identifying reference answers independently of the AI system being evaluated.
For example, manually verify bibliographic details before testing citation extraction. Have a researcher code a sample of documents before evaluating AI-assisted coding. Use an analysis with known results when testing generated statistical code. Compare paper summaries directly with the original papers.
The reference does not need to be infallible. Human coding can disagree too. But without an independent basis for comparison, evaluation risks becoming little more than asking the AI whether the AI did a good job.
Measure the Errors That Matter
An overall success rate can hide the mistakes that are most consequential to a study.
Suppose an extraction system correctly captures 95 percent of fields but systematically misidentifies the primary outcome. That remaining 5 percent may matter far more than many of the correct fields. Likewise, a literature tool that returns mostly real citations but occasionally fabricates one has a qualitatively important failure mode.
Examine not only how often the tool fails, but how it fails.
| Evaluation question |
What to inspect |
Why it matters |
| Does it produce the correct result? |
Accuracy against independently checked answers |
Tests basic task performance |
| What does it miss? |
Omissions, false negatives, incomplete extraction |
Apparently clean output can conceal missing evidence |
| What does it invent? |
Unsupported claims, values, sources, or interpretations |
Fabricated information can enter the research record unnoticed |
| Does performance vary? |
Different topics, documents, languages, prompts, and edge cases |
Average performance may conceal unstable subgroups or conditions |
| Can errors be detected? |
Traceability to source material and inspectable intermediate output |
Detectable mistakes are easier to manage than opaque ones |
| What happens under uncertainty? |
Whether the tool acknowledges missing evidence or produces an answer anyway |
Failure behavior affects the risk of confident misinformation |
Test Source Traceability When the Task Depends on Evidence
If the tool makes claims about published research, evaluation should extend beyond whether citations appear on screen.
Check whether cited sources exist, whether bibliographic details are correct, whether you can access or identify the original material, and whether the source actually supports the generated claim. The distinction matters because source citations do not guarantee that an AI-generated answer is correct.
For source-dependent workflows, provenance is part of performance.
Repeat Some Tests
Generative systems may not produce identical outputs every time. Prompt wording, system updates, model routing, settings, retrieved sources, and stochastic generation can all affect results.
Repeat at least some important tests and vary realistic inputs. You are looking for stability where stability matters and for conditions that produce unexpected changes.
If a workflow requires reproducibility, document enough about the system, prompts, settings, source material, dates, and procedures to understand what was done. Exact reproduction may not always be possible with changing proprietary services, which is itself a limitation worth recognizing.
Evaluate the Complete System, Not Just the Model
An AI application may combine a generative model with search, retrieval, document parsing, databases, code execution, or other components. Failures can arise anywhere in that chain.
A strong underlying model cannot summarize text it never received correctly. A retrieval system can supply irrelevant papers to an excellent model. A document parser can corrupt a table before the model analyzes it.
Likewise, greater underlying model capability does not automatically establish research reliability. Evaluate what you actually interact with.
Examine Data Handling Before Uploading Research Material
Performance testing answers only part of the question. A highly accurate system can still be unsuitable if using it requires inappropriate disclosure of research information.
Before testing with real sensitive data, determine what information the tool receives, how it is processed and retained, what controls apply, and whether the proposed use complies with relevant ethical, contractual, institutional, and legal requirements.
UNESCO's guidance on generative AI in education and research emphasizes protection of data privacy. The European Commission's Living Guidelines on the Responsible Use of Generative AI in Research likewise place responsibility on researchers to protect privacy, confidentiality, intellectual property, and sensitive information.
Use synthetic, public, de-identified, or otherwise appropriate test material where necessary rather than exposing protected information merely to discover whether a tool works.
Check Whether the Tool Is Transparent Enough to Evaluate
You cannot evaluate every internal mechanism of a proprietary AI system. But you should be able to establish enough about the system to make a defensible decision.
Relevant information may include intended functions, limitations, source coverage, data practices, model or system documentation, update behavior, and how important outputs are generated or grounded.
If the information needed to assess a consequential use is unavailable, opacity itself becomes part of the risk. There is a point at which an AI tool can become too opaque for a particular research use.
Decide the Human Oversight Before Deployment
Evaluation should not end with “pass” or “fail.” It should determine the conditions under which the tool may be used.
Perhaps every generated citation must be manually verified. Perhaps a researcher must review all extracted outcome data. Maybe AI-generated code must be inspected and tested before execution. Or perhaps the tool is permitted only for brainstorming and never receives participant data.
These conditions convert evaluation findings into a usable research protocol.
Re-Evaluate When Something Important Changes
AI systems are not necessarily static research instruments. Providers can change models, retrieval systems, interfaces, limits, policies, and features.
Re-evaluation may be warranted when the system changes materially, the research task changes, the type of data changes, an important failure is discovered, or the tool begins playing a more consequential role in the study.
A successful pilot is evidence about the system and conditions you tested, not a permanent certificate of reliability.
04 · A Practical Example
Testing AI Before Using It for Study Data Extraction
Hypothetical Example
Can AI Reliably Extract Information From Research Papers?
A research team wants to use an AI tool to extract sample size, participant characteristics, study design, and outcome information from papers included in a review. Because extraction errors could affect the synthesis, the team evaluates the tool before applying it to the full set.
1. Define the use
The AI will propose structured extractions from included papers. It will not independently decide the final values entered into the review dataset.
2. Create a test set
The team selects a varied sample of papers containing different designs, reporting styles, tables, and levels of complexity.
3. Establish reference data
Researchers manually extract and reconcile the target fields before comparing them with the AI output.
4. Test the system
The team uses the planned prompts and workflow, recording incorrect values, omissions, unsupported extractions, and ambiguous cases.
5. Examine failure patterns
The AI performs well on sample size but occasionally mistakes subgroup counts for total enrollment and struggles when outcomes appear only in tables.
6. Set safeguards
The team decides that AI can assist with initial extraction, but a researcher must verify every value against the original paper, with particular attention to sample counts and tables.
7. Monitor during use
Unexpected failures are recorded. If error patterns change or the tool itself changes materially, the workflow is reassessed.
The evaluation does not establish that the tool is “accurate AI.” It establishes a narrower and more useful conclusion: under the tested conditions, it can assist this particular extraction process if specified verification remains in place.