03 · What You Need to Know
Where the “AI Is Just Another Research Tool” Analogy Works and Where It Breaks
Researchers Have Always Used Tools to Extend Their Capabilities
A researcher using statistical software does not personally perform millions of arithmetic operations. A reference manager can apply thousands of citation-style rules. A qualitative analysis package can retrieve coded passages almost instantly. A programming language can automate transformations across datasets containing millions of observations.
None of this makes the research inherently less scholarly.
The intellectual contribution of research has never depended on manually executing every operation. Researchers routinely rely on instruments, algorithms, software libraries, databases, and technical infrastructure created by others.
Generative AI belongs within that history of tool-assisted scholarship. This is one reason it is unhelpful to treat AI-assisted research as though any machine contribution automatically displaces human scholarship.
Calling Something a “Tool” Does Not Establish How It Should Be Used
The problem appears when “tool” becomes an argument rather than a description.
A microscope and a word processor are both research tools. That does not mean they need identical calibration, validation, training, reporting, or data-handling procedures.
The relevant questions concern what the tool does:
- What inputs does it receive?
- What operation does it perform?
- What output does it produce?
- How predictable is that output?
- How can errors be detected?
- Can the operation be independently reproduced?
- What happens to data supplied to the system?
- How consequential would an undetected error be?
Generative AI should be judged through those questions rather than either exceptionalized as mysterious intelligence or normalized through the vague statement that “software is software.”
Statistical Software Usually Executes a More Explicitly Defined Analytical Procedure
Suppose you ask statistical software to calculate the arithmetic mean of five values.
The software applies a mathematically defined operation. If the data and operation are the same, researchers generally expect the same result.
The software can still contribute to bad research. You can specify the wrong model, mishandle missing data, select inappropriate options, misinterpret the output, or feed the software incorrect data. Statistical software is not a methodological guardian.
But the computational relationship between input and output is usually more constrained and inspectable than asking a generative AI system, “Why do you think this result occurred?”
A Generative AI System Can Produce an Explanation Rather Than Calculate One
Consider a regression coefficient that changes substantially after another variable enters the model.
Statistical software calculates the new coefficient according to the specified model.
A generative AI system asked to explain the change may produce several possible mechanisms. It might discuss confounding, suppression, multicollinearity, omitted variables, measurement, or model specification.
That explanation is generated content. Unless the system has performed and validated the relevant diagnostic analysis, the explanation should not be treated as though it were another number from the statistical output.
Statistical calculation
The software executes a specified computational procedure on supplied data.
Generated interpretation
The AI produces language representing one or more plausible explanations based on its model, context, instructions, and any tools or evidence available to it.
Both involve software. They do not produce the same epistemic kind of output.
Writing Software Provides the Closest Familiar Analogy, but It Still Has Limits
Researchers have long used spelling checkers, grammar tools, autocomplete, translation software, readability tools, and style checkers.
Generative AI clearly extends that lineage. If a system changes “The results indicates” to “The results indicate,” the interaction resembles conventional grammar correction.
But generative AI can move far beyond surface correction. It can create an argument, invent an example, restructure a theoretical framework, propose a causal mechanism, summarize literature, generate references, or write an entire discussion section.
The distinction between correction and content generation matters because the latter can introduce new scholarly claims rather than merely alter their expression.
This is why AI-assisted and AI-generated contributions should not always be treated as equivalent forms of writing support.
Reference Managers Transform Stored Information More Than They Invent It
A reference manager may take bibliographic metadata and format it according to a citation style. Errors can occur because the metadata are wrong, the style is configured incorrectly, or the software has defects.
But a conventional reference manager is not expected to invent a plausible article because the bibliography seems to need one more citation.
A general-purpose language model can generate a reference-shaped sequence containing plausible authors, title, journal, year, pages, and identifier even when the corresponding publication does not exist.
NIST's Generative AI Profile explicitly identifies fabricated citations and confabulated logic among the risks of generated content.
That difference changes the appropriate verification procedure. A reference imported from a reliable bibliographic source may require metadata checking. A citation generated from an ungrounded language model first requires establishing that the source exists at all.
Traditional Software Can Be Probabilistic Too
The comparison should not be exaggerated into “traditional software is deterministic; AI is probabilistic.”
Research software can implement stochastic simulations, machine-learning algorithms, Bayesian procedures, heuristic optimization, probabilistic record linkage, random sampling, or many other methods involving randomness or uncertainty.
Likewise, some AI workflows can be configured for greater repeatability.
The more meaningful distinction concerns the combination of open-ended generation, variable outputs, broad natural-language interaction, uncertain provenance, and the ability to produce plausible content beyond what has been explicitly supplied.
Traditional Software Can Also Produce Convincing Nonsense
A statistical package will happily calculate an inappropriate model if the researcher asks for one.
A spreadsheet can propagate a bad formula across 100,000 rows with extraordinary consistency. A reference manager can reproduce incorrect metadata in every citation. A programming script can generate beautifully formatted wrong results.
Generative AI therefore did not invent the problem of researchers trusting software they do not understand.
What it changes is how easy it becomes to request sophisticated outputs through ordinary language and receive results whose evidential basis may not be immediately visible.
Natural-Language Interaction Can Hide Technical Boundaries
Traditional software often makes its boundaries obvious.
A statistical package expects variables and models. A spreadsheet expects formulas. A reference manager expects bibliographic records.
A generative AI assistant can be asked almost anything.
That flexibility can make the system appear more competent than it is. A researcher can ask about a statistical method, legal requirement, theoretical framework, programming problem, historical event, and publication policy in one conversation, receiving fluent responses to all of them.
The interface does not clearly signal when the system has moved from a task it performs reliably to one where its answer is speculative or incorrect.
This creates an evaluation burden that is less familiar with narrowly specialized research software.
Generative AI Can Blur the Difference Between Interface and Underlying Tool
Modern AI assistants increasingly use external tools.
A single conversational interface might generate text with a language model, search the web, retrieve a document, execute Python code, query a database, perform a calculation, or call another specialized service.
This means not every output shown by an AI assistant was necessarily produced by language generation.
| What you ask |
What the system might actually do |
What you need to know |
| “Calculate this statistic” |
Generate an answer or invoke a computational tool |
Which mechanism produced the result? |
| “Find recent papers” |
Generate likely references or search an external information source |
Were records actually retrieved, and from where? |
| “Summarize this PDF” |
Process supplied document content and generate a summary |
Was the complete relevant content available and accurately represented? |
| “Explain my results” |
Generate an interpretation, perhaps using supplied data or tool outputs |
Which claims are demonstrated by evidence and which are generated possibilities? |
This is why understanding the difference between AI assistants and other research information tools remains important even as the technologies converge.
Data Handling Can Differ Substantially From Desktop Research Software
A researcher may use conventional software entirely on a local computer. The data never need to leave the device.
A cloud-based generative AI service may process prompts and uploaded files through external infrastructure under terms and technical arrangements that vary among services and account types.
This difference matters when researchers handle personal information, confidential participant data, unpublished manuscripts, proprietary information, peer-review material, or other protected content.
UNESCO's guidance emphasizes data privacy and human-centred governance of generative AI, while the European Commission's current guidelines address responsible use, accountability, transparency, and risks associated with AI in research information environments.
The relevant comparison is therefore not “AI versus software” in the abstract. Researchers need to compare the actual deployment, data flow, contractual conditions, and institutional requirements of the systems they intend to use.
Generative AI Can Change More Quickly Than Conventional Research Workflows
Cloud-based generative systems may change models, interfaces, retrieval capabilities, system instructions, safety mechanisms, and other components over time.
The European Commission describes its guidance as “living” precisely because generative AI technologies and their uses continue to develop rapidly; the guidelines were updated again in May 2026 to reflect technological developments and emerging risks.
For research workflows where exact reproducibility matters, researchers may therefore need to record enough information about the system and process to understand what was used at the time.
This challenge is not unique to AI. Software versions, package dependencies, databases, operating systems, and online services also change. Generative AI simply makes versioning and behavioral change especially visible because the output itself may vary.
The Analogy Works Best When It Encourages Proportionate Evaluation
Researchers do not verify every operation of every software tool from first principles.
Few people manually recalculate every reference-manager formatting rule or independently derive every algorithm in a statistical library. Instead, researchers calibrate scrutiny according to the function, maturity of the tool, evidence of validity, consequences of error, and ability to check outputs.
Generative AI deserves the same general principle.
A suggested manuscript title does not require the same validation as AI-generated statistical code. A grammar correction does not deserve the same scrutiny as a generated interpretation of participant data. A generated reference deserves different checking from a punctuation correction.
Treating AI as software is therefore useful when it encourages rational, task-specific scrutiny rather than either unconditional trust or blanket fear.
The Analogy Fails When “Just a Tool” Becomes “No New Safeguards Needed”
Generative AI can confabulate information, generate source-like content, process externally supplied data, respond differently across interactions, and produce open-ended scholarly material through ordinary language. Those properties can create risks not encountered in the same form with more constrained software.
NIST's Generative AI Profile identifies risks including confabulation, privacy, information integrity, intellectual-property issues, and other concerns requiring generative-AI-specific risk management.
The European Commission's current research guidance likewise retains specific recommendations for generative AI rather than assuming ordinary software practices are sufficient.
This is why generative AI can require different research safeguards from traditional software even though both are tools.