03 · What You Need to Know
Opacity Is a Matter of Degree and Context
Not Knowing Everything Does Not Make a Tool Unusable
Complete transparency is an unrealistic standard for many modern AI systems. Researchers may not know every training example, parameter, optimization decision, system prompt, retrieval rule, or internal transformation involved in producing an output.
NIST treats transparency as the extent to which appropriate information about an AI system and its outputs is available to people interacting with it. Importantly, NIST describes meaningful transparency as providing information at a level appropriate to the relevant role and stage of the AI lifecycle rather than requiring identical disclosure for everyone.
That provides a useful research principle: the relevant question is not whether everything about the AI is visible, but whether the information necessary for your research decision is visible.
Transparency, Explainability, and Interpretability Are Related but Different
These terms are often collapsed into the vague demand that AI should “explain itself.” They refer to different aspects of understanding a system.
Transparency
Relevant information about the system, its development, intended use, operation, data, outputs, limitations, or governance is available to appropriate users.
Explainability and interpretability
Users can obtain useful information about how a system produces or arrives at an output and what that output means in its context of use.
NIST treats accountability and transparency, as well as explainability and interpretability, as characteristics relevant to trustworthy AI. It also notes that transparency supports accountability and that appropriate information needs vary according to context.
A research tool can therefore be sufficiently transparent for one purpose without making every internal model mechanism interpretable.
Ask What You Actually Need to Know for the Task
The transparency requirement should follow the research use.
If you ask an AI system to suggest alternative wording for a sentence, you can inspect the output directly. You may not need to know what training corpus produced its preference for one adjective over another.
If the AI identifies which studies should be included in a review, the situation changes. You may need information about what evidence it receives, how the decision is produced or at least validated, how accurately it performs, where it fails, and whether excluded records can be audited.
If it extracts outcome values that enter a meta-analysis, source-level traceability may become more important still.
| AI role |
Typical transparency need |
Why |
| Brainstorming keywords |
Relatively low |
The researcher can directly accept, reject, or modify suggestions |
| Language editing |
Relatively low |
Changes are visible and can usually be inspected directly |
| Literature discovery |
Moderate to high |
Source coverage and retrieval limitations affect what evidence can be found |
| Summarizing papers |
High for consequential claims |
The researcher needs a path back to the original evidence |
| Study screening or classification |
High |
Opaque errors can alter which evidence enters the study |
| Research data extraction or analysis |
High |
Outputs may directly affect datasets, calculations, and findings |
| Decisions affecting participants or other people |
Very high |
Accountability, validation, explanation, fairness, and potential harms become more consequential |
Opacity About Information Sources Is Particularly Problematic
When an AI tool makes claims about research evidence, researchers should be able to understand enough about the relevant evidence base to interpret the answer.
If a system says that “the literature shows” something but does not identify what literature it can access, what sources it retrieved, or how its answer can be checked, the researcher cannot readily distinguish evidence-based synthesis from plausible generation.
This does not require a complete reconstruction of everything the underlying model encountered during training. It does require meaningful provenance where the system retrieves papers, searches databases, or claims to ground an answer in identifiable evidence.
If the information sources behind a research answer cannot be established, the tool may be unsuitable for evidence-dependent work even if it remains useful for brainstorming.
Opacity About Performance Can Be More Important Than Opacity About Architecture
A researcher may know relatively little about the internal architecture of a proprietary AI system yet still have useful evidence about its performance on a defined task.
Conversely, a technically documented or open model may still lack evidence that it works reliably for the researcher's intended use.
NIST's AI Risk Management Framework treats validity and reliability as foundational characteristics of trustworthy AI and calls for systems to be evaluated under relevant conditions, with limitations on generalizability documented. Its measurement guidance likewise emphasizes that appropriate evaluation depends on context.
For many research decisions, evidence about task-specific performance is therefore more actionable than an impressive architecture diagram.
Opacity About Failure Modes Should Raise the Threshold for Use
Knowing that an AI system sometimes makes mistakes is not very informative. Researchers need to know enough to anticipate where errors are likely, how consequential they could be, and whether they can be detected.
NIST's Generative AI Profile identifies confabulation as a risk in which generative systems confidently produce erroneous or false information. It specifically notes that generated logic or citations can themselves be false and may encourage inappropriate trust.
A system becomes difficult to justify for consequential work when its important errors are both poorly characterized and difficult to detect.
That combination matters more than opacity alone. A system can be imperfect yet manageable when every output is easily checked. An opaque error that silently changes the research dataset presents a different risk.
Opacity About Data Handling Can Make a Tool Unusable Regardless of Output Quality
An AI tool might perform exceptionally well while providing insufficient information about what happens to uploaded material.
If you cannot establish how confidential transcripts, participant data, unpublished manuscripts, proprietary information, or other protected material will be processed, retained, or shared, strong technical performance does not resolve the governance problem.
UNESCO's guidance on generative AI in education and research highlights both the opacity of generative models and concerns about data privacy, while advocating human-centred validation and appropriate regulation of these systems.
Researchers should therefore examine the applicable privacy and data-handling information before treating a tool as suitable for sensitive material.
Opacity About Changes Can Undermine Reproducibility
AI systems can change while retaining the same product name. Providers may update the underlying model, retrieval process, system instructions, document handling, safety controls, or other components.
If a tool plays a substantive methodological role, researchers may need enough version or change information to understand whether the system used later in a project is materially different from the one originally evaluated.
A service that changes invisibly can make it difficult to reproduce procedures or interpret differences in outputs over time.
This does not make continuously updated systems unusable. It means reproducibility limitations should be acknowledged and, where necessary, mitigated through documentation, preserved inputs and outputs, repeated validation, or alternative methods.
A Generated Explanation Is Not Automatically a Faithful Explanation
One tempting solution to opacity is simply to ask the AI, “Explain how you reached that answer.”
That can produce useful reasoning or a helpful account of the factors relevant to the answer. It should not automatically be treated as a literal transcript of the model's internal computation.
NIST's Generative AI Profile warns that generative systems can produce confabulated logic that appears to justify an answer even when the answer itself is incorrect.
Researchers should therefore distinguish between an explanation that helps them evaluate an answer and independently established evidence about how the deployed system actually operates.
Watch Out
Do not treat a fluent “here is how I reasoned” response as an audit log unless the system actually provides a documented mechanism for that purpose. Generated explanations can themselves require verification.
Human Oversight Only Works When Humans Have Something to Inspect
“A human remains in the loop” is often presented as a safeguard. Its value depends on what the human can actually see and evaluate.
If researchers receive only a final classification with no source material, confidence information, rationale, audit trail, or practical way to test the decision, human oversight may become ceremonial rather than substantive.
Meaningful oversight requires enough information and authority for the researcher to challenge, correct, or reject the AI output.
Ask Whether You Could Defend the Tool's Role to a Skeptical Reviewer
A useful practical test is to imagine that a reviewer asks why the AI was used and how you know its contribution did not compromise the study.
Could you explain what the tool did? What information it received? What evidence supported its use? What limitations were known? How outputs were checked? What data conditions applied? Whether the system changed during the project?
You do not need an answer to every conceivable technical question. You do need answers to the questions necessary to justify the role the AI played.
Opacity Should Sometimes Reduce the Role Rather Than Eliminate the Tool
The choice is not always “use it fully” or “never use it.”
An AI system may be too opaque to make final screening decisions but still useful for suggesting records for human review. It may be unsuitable for autonomous data extraction but useful for proposing candidate values that researchers verify against the original paper. It may be inappropriate for interpreting confidential participant data but acceptable for generating generic code examples.
Reducing the tool's authority can sometimes make residual uncertainty manageable.
Some Research Roles Demand a Harder Stop
There are situations where uncertainty cannot be adequately compensated for by verification or restricted use.
If a consequential output cannot be independently checked, the system's relevant performance cannot be evaluated, the evidence base is unknown, sensitive-data conditions cannot be established, or researchers cannot reconstruct enough of the procedure to defend the resulting research decision, continuing to rely on the tool may no longer be methodologically defensible.
This is where opacity crosses from inconvenience into a reason not to use the system for that role.