Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Is Using Generative AI Like Using Statistical, Writing, or Other Research Software?

Generative AI is a software tool, but that does not make every use equivalent to using statistical, writing, reference, or other conventional research software. The analogy is useful only when researchers also account for differences in generation, predictability, provenance, verification, and data handling.

13
Generative AI vs. Research Software Guide 13 of 80
01 · The Question

Is Generative AI Really So Different From the Software Researchers Already Use?

Researchers have delegated work to software for decades.

Statistical packages perform calculations few researchers would attempt manually. Reference managers generate bibliographies. Grammar checkers alter prose. Transcription software converts speech into text. Search systems rank information. Programming libraries execute analytical procedures written by other people.

So when researchers are told to treat generative AI cautiously, a reasonable question follows: isn't this just another research tool?

In one important sense, yes. Generative AI is software, and using software to extend human capability is nothing new in research. But “it is a tool” does not tell us what safeguards are appropriate. A calculator, statistical package, spellchecker, database, and generative AI assistant are all tools precisely because that category is so broad.

02 · The Short Answer

Generative AI Is Another Research Tool, but Not Every Tool Has the Same Failure Modes

In Brief

Generative AI can reasonably be treated as research software, but using it is not always equivalent to using statistical, writing, reference, or other conventional software because generative systems can produce novel, probabilistic, source-opaque, and plausibly false outputs in response to open-ended instructions.

The comparison is most useful when it prevents unnecessary mystification of AI. It becomes misleading when “just another tool” is used to imply that all software deserves the same level of trust, verification, disclosure, data protection, or methodological scrutiny.

03 · What You Need to Know

Where the “AI Is Just Another Research Tool” Analogy Works and Where It Breaks

Researchers Have Always Used Tools to Extend Their Capabilities

A researcher using statistical software does not personally perform millions of arithmetic operations. A reference manager can apply thousands of citation-style rules. A qualitative analysis package can retrieve coded passages almost instantly. A programming language can automate transformations across datasets containing millions of observations.

None of this makes the research inherently less scholarly.

The intellectual contribution of research has never depended on manually executing every operation. Researchers routinely rely on instruments, algorithms, software libraries, databases, and technical infrastructure created by others.

Generative AI belongs within that history of tool-assisted scholarship. This is one reason it is unhelpful to treat AI-assisted research as though any machine contribution automatically displaces human scholarship.

Calling Something a “Tool” Does Not Establish How It Should Be Used

The problem appears when “tool” becomes an argument rather than a description.

A microscope and a word processor are both research tools. That does not mean they need identical calibration, validation, training, reporting, or data-handling procedures.

The relevant questions concern what the tool does:

  • What inputs does it receive?
  • What operation does it perform?
  • What output does it produce?
  • How predictable is that output?
  • How can errors be detected?
  • Can the operation be independently reproduced?
  • What happens to data supplied to the system?
  • How consequential would an undetected error be?

Generative AI should be judged through those questions rather than either exceptionalized as mysterious intelligence or normalized through the vague statement that “software is software.”

Statistical Software Usually Executes a More Explicitly Defined Analytical Procedure

Suppose you ask statistical software to calculate the arithmetic mean of five values.

A Defined Calculation
Mean = (x₁ + x₂ +... + xₙ) / n
x represents each observed value and n represents the number of observations.
For 4, 6, 8, 10, and 12: (4 + 6 + 8 + 10 + 12) / 5 = 8. The result means the arithmetic mean is 8. It does not establish whether the mean is the appropriate statistic for the research question or distribution.

The software applies a mathematically defined operation. If the data and operation are the same, researchers generally expect the same result.

The software can still contribute to bad research. You can specify the wrong model, mishandle missing data, select inappropriate options, misinterpret the output, or feed the software incorrect data. Statistical software is not a methodological guardian.

But the computational relationship between input and output is usually more constrained and inspectable than asking a generative AI system, “Why do you think this result occurred?”

A Generative AI System Can Produce an Explanation Rather Than Calculate One

Consider a regression coefficient that changes substantially after another variable enters the model.

Statistical software calculates the new coefficient according to the specified model.

A generative AI system asked to explain the change may produce several possible mechanisms. It might discuss confounding, suppression, multicollinearity, omitted variables, measurement, or model specification.

That explanation is generated content. Unless the system has performed and validated the relevant diagnostic analysis, the explanation should not be treated as though it were another number from the statistical output.

Statistical calculation The software executes a specified computational procedure on supplied data.
Generated interpretation The AI produces language representing one or more plausible explanations based on its model, context, instructions, and any tools or evidence available to it.

Both involve software. They do not produce the same epistemic kind of output.

Writing Software Provides the Closest Familiar Analogy, but It Still Has Limits

Researchers have long used spelling checkers, grammar tools, autocomplete, translation software, readability tools, and style checkers.

Generative AI clearly extends that lineage. If a system changes “The results indicates” to “The results indicate,” the interaction resembles conventional grammar correction.

But generative AI can move far beyond surface correction. It can create an argument, invent an example, restructure a theoretical framework, propose a causal mechanism, summarize literature, generate references, or write an entire discussion section.

The distinction between correction and content generation matters because the latter can introduce new scholarly claims rather than merely alter their expression.

This is why AI-assisted and AI-generated contributions should not always be treated as equivalent forms of writing support.

Reference Managers Transform Stored Information More Than They Invent It

A reference manager may take bibliographic metadata and format it according to a citation style. Errors can occur because the metadata are wrong, the style is configured incorrectly, or the software has defects.

But a conventional reference manager is not expected to invent a plausible article because the bibliography seems to need one more citation.

A general-purpose language model can generate a reference-shaped sequence containing plausible authors, title, journal, year, pages, and identifier even when the corresponding publication does not exist.

NIST's Generative AI Profile explicitly identifies fabricated citations and confabulated logic among the risks of generated content.

That difference changes the appropriate verification procedure. A reference imported from a reliable bibliographic source may require metadata checking. A citation generated from an ungrounded language model first requires establishing that the source exists at all.

Traditional Software Can Be Probabilistic Too

The comparison should not be exaggerated into “traditional software is deterministic; AI is probabilistic.”

Research software can implement stochastic simulations, machine-learning algorithms, Bayesian procedures, heuristic optimization, probabilistic record linkage, random sampling, or many other methods involving randomness or uncertainty.

Likewise, some AI workflows can be configured for greater repeatability.

The more meaningful distinction concerns the combination of open-ended generation, variable outputs, broad natural-language interaction, uncertain provenance, and the ability to produce plausible content beyond what has been explicitly supplied.

Traditional Software Can Also Produce Convincing Nonsense

A statistical package will happily calculate an inappropriate model if the researcher asks for one.

A spreadsheet can propagate a bad formula across 100,000 rows with extraordinary consistency. A reference manager can reproduce incorrect metadata in every citation. A programming script can generate beautifully formatted wrong results.

Generative AI therefore did not invent the problem of researchers trusting software they do not understand.

What it changes is how easy it becomes to request sophisticated outputs through ordinary language and receive results whose evidential basis may not be immediately visible.

Natural-Language Interaction Can Hide Technical Boundaries

Traditional software often makes its boundaries obvious.

A statistical package expects variables and models. A spreadsheet expects formulas. A reference manager expects bibliographic records.

A generative AI assistant can be asked almost anything.

That flexibility can make the system appear more competent than it is. A researcher can ask about a statistical method, legal requirement, theoretical framework, programming problem, historical event, and publication policy in one conversation, receiving fluent responses to all of them.

The interface does not clearly signal when the system has moved from a task it performs reliably to one where its answer is speculative or incorrect.

This creates an evaluation burden that is less familiar with narrowly specialized research software.

Generative AI Can Blur the Difference Between Interface and Underlying Tool

Modern AI assistants increasingly use external tools.

A single conversational interface might generate text with a language model, search the web, retrieve a document, execute Python code, query a database, perform a calculation, or call another specialized service.

This means not every output shown by an AI assistant was necessarily produced by language generation.

What you ask What the system might actually do What you need to know
“Calculate this statistic” Generate an answer or invoke a computational tool Which mechanism produced the result?
“Find recent papers” Generate likely references or search an external information source Were records actually retrieved, and from where?
“Summarize this PDF” Process supplied document content and generate a summary Was the complete relevant content available and accurately represented?
“Explain my results” Generate an interpretation, perhaps using supplied data or tool outputs Which claims are demonstrated by evidence and which are generated possibilities?

This is why understanding the difference between AI assistants and other research information tools remains important even as the technologies converge.

Data Handling Can Differ Substantially From Desktop Research Software

A researcher may use conventional software entirely on a local computer. The data never need to leave the device.

A cloud-based generative AI service may process prompts and uploaded files through external infrastructure under terms and technical arrangements that vary among services and account types.

This difference matters when researchers handle personal information, confidential participant data, unpublished manuscripts, proprietary information, peer-review material, or other protected content.

UNESCO's guidance emphasizes data privacy and human-centred governance of generative AI, while the European Commission's current guidelines address responsible use, accountability, transparency, and risks associated with AI in research information environments.

The relevant comparison is therefore not “AI versus software” in the abstract. Researchers need to compare the actual deployment, data flow, contractual conditions, and institutional requirements of the systems they intend to use.

Generative AI Can Change More Quickly Than Conventional Research Workflows

Cloud-based generative systems may change models, interfaces, retrieval capabilities, system instructions, safety mechanisms, and other components over time.

The European Commission describes its guidance as “living” precisely because generative AI technologies and their uses continue to develop rapidly; the guidelines were updated again in May 2026 to reflect technological developments and emerging risks.

For research workflows where exact reproducibility matters, researchers may therefore need to record enough information about the system and process to understand what was used at the time.

This challenge is not unique to AI. Software versions, package dependencies, databases, operating systems, and online services also change. Generative AI simply makes versioning and behavioral change especially visible because the output itself may vary.

The Analogy Works Best When It Encourages Proportionate Evaluation

Researchers do not verify every operation of every software tool from first principles.

Few people manually recalculate every reference-manager formatting rule or independently derive every algorithm in a statistical library. Instead, researchers calibrate scrutiny according to the function, maturity of the tool, evidence of validity, consequences of error, and ability to check outputs.

Generative AI deserves the same general principle.

A suggested manuscript title does not require the same validation as AI-generated statistical code. A grammar correction does not deserve the same scrutiny as a generated interpretation of participant data. A generated reference deserves different checking from a punctuation correction.

Treating AI as software is therefore useful when it encourages rational, task-specific scrutiny rather than either unconditional trust or blanket fear.

The Analogy Fails When “Just a Tool” Becomes “No New Safeguards Needed”

Generative AI can confabulate information, generate source-like content, process externally supplied data, respond differently across interactions, and produce open-ended scholarly material through ordinary language. Those properties can create risks not encountered in the same form with more constrained software.

NIST's Generative AI Profile identifies risks including confabulation, privacy, information integrity, intellectual-property issues, and other concerns requiring generative-AI-specific risk management.

The European Commission's current research guidance likewise retains specific recommendations for generative AI rather than assuming ordinary software practices are sufficient.

This is why generative AI can require different research safeguards from traditional software even though both are tools.

04 · A Practical Example

Three Software Tools, Three Different Kinds of Trust

Hypothetical Example

A Researcher Preparing the Results Section

A researcher uses three software systems while preparing a manuscript.

Statistical software The researcher specifies a regression model. The software estimates coefficients and standard errors according to the selected procedure. The researcher checks assumptions, diagnostics, data handling, model specification, and interpretation.
Reference manager The researcher inserts citations from stored bibliographic records. The software formats them according to the selected journal style. The researcher checks that the metadata and cited sources are correct.
Generative AI assistant The researcher supplies selected results and asks, “Explain why predictor X became significant after variable Z was added.” The system generates an explanation involving a suppression effect.
Different verification The researcher does not treat the generated explanation like the regression coefficient or formatted citation. Instead, the proposed suppression effect becomes a hypothesis to investigate using the actual model, variables, diagnostics, theory, and appropriate statistical literature.

All three systems are research software. The mistake would be assuming that because they are all software, every output has the same evidential status.

05 · What Researchers Often Get Wrong

Common Misunderstandings About AI as “Just Another Tool”

Misconception

If Statistical Software Is Allowed, Should Generative AI Automatically Be Allowed Too?

No. Permission depends on the particular function, data, policy, and research context. Two forms of software can raise different issues involving confidentiality, generation, disclosure, authorship, reproducibility, or validation.

Misconception

If Both Tools Can Make Mistakes, Are Their Risks Basically the Same?

No. The fact that all software can fail does not make failure modes interchangeable. Generated false citations, incorrect statistical specifications, corrupted metadata, and spreadsheet formula errors require different detection and prevention strategies.

Misconception

Is Using AI for Writing the Same as Using a Grammar Checker?

Sometimes the functions can be similar. If AI corrects grammar in researcher-authored prose, the analogy may be close. If it generates arguments, interpretations, literature summaries, or entire sections, the system is performing a substantively different task.

Misconception

Is Using AI-Generated Code the Same as Using a Statistical Package?

Not exactly. Statistical software implements defined procedures through existing code, while generative AI may create new code in response to your instructions. Generated code can be useful, but researchers need to verify what it actually implements before treating its output as methodologically valid.

Misconception

If AI Is Only a Tool, Does the Researcher Remain Responsible?

Yes. Tool use does not transfer scholarly accountability. Researchers remain responsible for the research they submit or publish, including consequential errors introduced through the tools they choose and the outputs they accept.

Misconception

Does Treating AI Differently Mean Researchers Are Applying a Double Standard?

Not if the different treatment follows different technical characteristics and risks. Research tools routinely receive different safeguards according to their function. The relevant standard is proportionality, not identical treatment of unlike systems.

06 · What This Means for You

Treat Generative AI Like Software, Then Ask What Kind of Software It Is

The most useful position lies between two extremes.

Do not treat generative AI as a mysterious artificial researcher whose outputs deserve special authority. But do not reduce it to “just software” in a way that erases meaningful differences in generation, provenance, variability, data handling, and verification.

A simple decision framework

If the tool performs a defined and inspectable operation
Validate the inputs, procedure, settings, assumptions, and resulting output in the way appropriate to that operation.
If AI transforms researcher-created material
Compare the result with the original and verify that meaning, evidence, and uncertainty have been preserved.
If AI generates new factual or scholarly content
Require independent verification of consequential claims and sources.
If the system handles sensitive research information
Evaluate the actual data flow, privacy, security, contractual, ethics, and institutional requirements rather than assuming ordinary software practices apply.
If AI output influences methods, evidence, or conclusions
Increase methodological scrutiny and document the role of the system where needed for transparency and reproducibility.

“It's a tool” is a useful starting point. It is not a risk assessment.

07 · A Quick Checklist

Before Treating Generative AI Like Your Other Research Software, Compare the Functions

For the specific AI function you intend to use, check:
Identify whether the system is calculating, retrieving, transforming, predicting, or generating the output.
Determine whether the same inputs reliably produce the same relevant output or whether variation matters to the research workflow.
Establish whether you can trace important claims or results to underlying data, sources, calculations, or retrieved evidence.
Verify generated citations and factual claims independently rather than treating them like stored bibliographic records.
Inspect and test generated code rather than assuming it is equivalent to a validated statistical procedure.
Check how prompts, uploaded files, and research data are processed and whether the proposed use is permitted.
Consider whether model or service changes could affect reproducibility and what documentation is therefore necessary.
Match the strength of verification to the consequences of an undetected error.
Verify current institutional, funder, ethics, journal, and publisher requirements for the particular use rather than assuming ordinary software rules apply automatically.
08 · Frequently Asked Questions

Frequently Asked Questions About Generative AI and Research Software

Is generative AI technically software?

Yes. Generative AI systems are implemented through software and computational infrastructure. The meaningful research question is not whether they are software but what functions they perform and what safeguards those functions require.

Is ChatGPT like SPSS, R, or Stata?

They can all support research, but their functions differ substantially. Statistical software executes analytical procedures specified through code, commands, or interfaces. A general-purpose generative AI assistant can produce open-ended language and may also invoke external computational tools. Researchers should identify which mechanism produced a particular result.

Is AI writing assistance just an advanced grammar checker?

It can function that way for limited editing tasks, but generative AI can also create entirely new substantive content. Once the system generates arguments, interpretations, summaries, or scholarly claims, the analogy with ordinary grammar correction becomes much weaker.

Can traditional research software hallucinate?

“Hallucination” or confabulation is terminology associated particularly with generative systems producing false or unsupported content. Traditional software can certainly produce incorrect results through bad data, bugs, inappropriate settings, flawed algorithms, or misuse, but the failure mechanism is not necessarily the same.

Why should AI-generated references receive more checking than references from my reference manager?

A reference manager ordinarily formats or stores an existing bibliographic record, although that record can still contain errors. An ungrounded generative model can produce an entirely nonexistent citation. You therefore need first to establish that the generated source exists and then verify its details and relevance.

Does generative AI always require disclosure when other software does not?

No universal rule applies across all research contexts. Requirements vary by institution, funder, publisher, journal, discipline, and type of use. Check the current policy governing your work rather than inferring disclosure requirements solely from the software category.

Why does generative AI need different safeguards if it is just software?

Because safeguards should follow technical characteristics and consequences rather than the generic category “software.” Open-ended generation, confabulation, source opacity, variable outputs, external data processing, and the ability to generate substantive scholarly content can create distinctive risks. These differences are examined directly in why generative AI requires different research safeguards.

09 · The Bottom Line

Generative AI Is a Research Tool, but “Tool” Is Only the Beginning of the Analysis

The Bottom Line

Generative AI is legitimately understood as research software, but using it is not always equivalent to using statistical, writing, reference, or other conventional software because its outputs and failure modes can differ substantially.

Treat AI as a tool rather than an artificial authority, then evaluate the specific function it performs. The appropriate safeguards should follow how the output is generated, how it can fail, what data the system handles, and how much the research depends on getting that output right.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes