Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Is a Large Language Model, and How Can Researchers Use One?

A large language model learns statistical patterns in token sequences and uses them to generate or process language and related content. For researchers, LLMs can be versatile assistants, but fluent output should never be confused with verified evidence.

03
Large Language Models for Researchers Guide 3 of 80
01 · The Question

What Is Actually Happening When You Ask an AI Assistant a Research Question?

You type a research question into a chatbot. Seconds later, it explains a theory, proposes variables, revises a paragraph, generates Python code, or produces what looks remarkably like a miniature literature review.

The experience can feel as though you have asked a knowledgeable person who has read an enormous library and is now retrieving the answer for you. That mental model is tempting, but it is incomplete and can become dangerous in research.

Many conversational generative AI systems are built around large language models, or LLMs. Understanding what an LLM is, even at a nontechnical level, helps explain both why these systems can be extraordinarily useful and why they can produce confident nonsense.

02 · The Short Answer

An LLM Generates Language From Learned Statistical Patterns

In Brief

A large language model is an AI model trained on very large amounts of data to learn statistical relationships among tokens and use those relationships to process and generate sequences, including human-like text.

Researchers can use LLM-based systems for tasks such as exploring terminology, explaining concepts, transforming text, assisting with code, organizing information, and generating possibilities. But an LLM's fluent response is generated output, not automatically a retrieved, verified, or evidence-based answer.

03 · What You Need to Know

How Large Language Models Work and What That Means for Research

“Large,” “Language,” and “Model” Each Tell You Something

The name is less mysterious once it is unpacked.

Large Modern LLMs are trained using enormous datasets and computational resources and may contain very large numbers of learned parameters.
Language They learn statistical structure in sequences represented as tokens and are particularly capable of processing and generating language, although modern models and systems may also work with images, audio, code, and other modalities.
Model They are computational models that learn patterns from training data and use those learned relationships to produce outputs. They are not databases containing a neat record for every sentence they can generate.

An LLM is therefore one particular technology within the much broader landscape of AI in research. Not every AI system is an LLM, and not every form of generative AI is language-based.

LLMs Work With Tokens, Not Words in Quite the Way Humans Do

Before text can be processed by an LLM, it is represented as units commonly called tokens. A token might correspond to a whole word, part of a word, punctuation, or another textual unit depending on the tokenizer used by the model.

When an LLM generates text, it operates over these token representations. Given the preceding context, the model assigns probabilities to possible subsequent tokens and generation proceeds token by token.

This simple description leaves out a great deal of technical machinery, but it captures a crucial point for researchers: the system is generating a continuation based on learned relationships and supplied context, not consulting an internal librarian who checks every sentence against the scholarly record before allowing it to appear.

Predicting Tokens Can Produce Surprisingly Sophisticated Behavior

“It predicts the next token” can sound as though an LLM is little more than predictive text on a phone. That comparison badly understates what large models can do.

To predict sequences well across diverse material, large models can acquire rich statistical representations of language, concepts, relationships, styles, programming structures, and many other patterns represented in their training and subsequent adaptation. Modern LLM-based systems can consequently perform tasks that look strikingly different from simple sentence completion.

They may summarize text, translate languages, generate computer code, classify material, transform writing, answer questions, extract structured information, compare arguments, follow multi-step instructions, or produce explanations.

Researchers should therefore avoid two opposite mistakes: pretending LLMs are trivial autocomplete, and imagining that fluent performance means they possess a human researcher's understanding, evidential judgment, or accountability.

An LLM Is Not the Same Thing as the Chatbot You Use

This distinction is easy to miss.

The LLM is the underlying model. The application through which you interact with it may add many other components: system instructions, safety mechanisms, document retrieval, web search, memory, calculators, code execution, image processing, databases, or external tools.

Layer What it may do Why researchers should care
Large language model Processes context and generates token sequences Explains many capabilities and limitations of generated language
Application Provides the interface, instructions, permissions, file handling, and workflow Two applications using related models may behave differently
Retrieval or search Supplies external documents or current information to the model An answer grounded in retrieved sources differs from one generated without them
External tools May perform calculations, execute code, query databases, or complete other operations Not every result shown by an AI assistant was generated by the LLM itself

This matters when evaluating an answer. If an AI assistant correctly calculates a complicated statistic, you may need to know whether the LLM generated the number or whether the application invoked a computational tool. If it provides current literature, you need to know whether it actually searched a scholarly source or merely generated plausible citations.

Training Is Not the Same as Search

An LLM may encode a remarkable amount of information through patterns learned during training, but asking it a question is not automatically equivalent to searching the web, querying a scholarly database, or retrieving a paper.

A research database typically returns records matching a query from a defined collection. A search engine retrieves and ranks indexed resources. A reference manager stores and organizes bibliographic information. An LLM generates an answer from its model and context unless the surrounding application gives it access to retrieval or other external tools.

That distinction is central to understanding how AI assistants differ from search engines, research databases, and reference managers.

If you ask an ungrounded LLM, “What studies show that intervention X improves outcome Y?”, its ability to produce author-like names, article-like titles, and journal-like references does not mean those publications were retrieved from a bibliographic database.

Why LLMs Can Hallucinate

Researchers commonly use the term hallucination for generated information that is false, unsupported, nonsensical, or otherwise inaccurate despite being presented as an answer.

Hallucination is not merely a matter of an LLM occasionally having a bad day. It arises from the nature of systems optimized to generate plausible continuations rather than guarantee that every generated proposition corresponds to a verified fact. Research published in Nature has examined why confident guessing can persist in language models and how common evaluation practices can reward answering rather than appropriately abstaining.

For researchers, the practical consequence matters more than the terminology: an LLM can produce a grammatically excellent, highly specific, academically styled statement that is false.

Watch Out

Specificity is not verification. An LLM can generate precise dates, quotations, statistics, DOI-like strings, article titles, methodological details, or citations without those details being correct. Verify consequential information independently.

More Capable Models Do Not Eliminate the Need for Verification

LLMs have improved substantially, and surrounding systems can reduce some failure modes through retrieval, tool use, better training, improved evaluation, and other techniques. Researchers should not assume that every limitation observed in an earlier generation applies unchanged to every newer system.

But improving average performance does not turn generated language into evidence. Even highly capable systems can make errors, misrepresent sources, misunderstand ambiguous instructions, omit important qualifications, or generate unsupported conclusions.

This is why responsible use should be calibrated to consequences. A mistaken suggestion for an alternative manuscript title is inconvenient. A fabricated reference in a systematic review, an incorrect interpretation of an inferential test, or an invented claim about participant data can affect the integrity of the research.

Context Matters Because the Model Responds to What It Can Access

An LLM does not respond to your prompt in isolation. Its effective context may include system instructions, conversation history, uploaded documents, retrieved material, tool outputs, and other information supplied by the application.

This explains why providing a model with the actual text of an article can produce a different task from asking, “What does Smith et al. say about this?” without supplying or retrieving that article. In the first case, the system may have the source in its working context. In the second, you may be asking it to generate an answer without reliable access to the publication.

Even source-provided tasks still require checking. A model can overlook qualifications, confuse sections, overgeneralize findings, or generate an interpretation not supported by the document.

Prompting Changes the Output, but It Does Not Change the Evidence

Instructions matter. A clearer prompt can specify the task, context, constraints, format, audience, or evidence the system should use. That can substantially improve usefulness.

But prompt engineering has a limit that researchers should keep firmly in view. Telling a model to “only provide real references,” “do not hallucinate,” or “make sure everything is scientifically accurate” does not itself verify the resulting answer.

A stronger instruction can influence model behavior. It cannot transform an unsupported generated statement into evidence.

Researchers Can Use LLMs Most Productively as Assistive Systems

The versatility of LLMs makes them useful across many research activities. Depending on the system, task, data restrictions, and applicable policies, they may help researchers:

  • explore terminology and alternative ways to formulate a problem;
  • explain unfamiliar concepts as a starting point for further verification;
  • generate or debug code that the researcher can inspect and test;
  • transform researcher-supplied text into another format, structure, or level of explanation;
  • compare material supplied in the context;
  • extract candidate information from documents for subsequent checking;
  • generate possible counterarguments, questions, interpretations, or analytical approaches;
  • support drafting, revision, translation, or language editing where permitted.

The broader range of research tasks generative AI can support is growing as systems gain new modalities and tools. Capability, however, should not be confused with permission or methodological suitability.

The Best Mental Model Is Not “Oracle”

For many research tasks, a productive way to think about an LLM is as a system capable of generating useful candidate outputs extremely quickly.

Candidate is the important word. A proposed search term can be tested. Generated code can be run and inspected. A suggested explanation can be investigated. A rewritten paragraph can be compared with the original. A proposed counterargument can force you to reconsider your reasoning.

Problems arise when candidate silently becomes answer, and then answer silently becomes evidence.

04 · A Practical Example

Using an LLM to Help With a Literature Search Without Letting It Become the Literature Search

Hypothetical Example

From a Broad Concept to Searchable Terminology

Suppose a doctoral researcher wants to study why university students stop using educational technologies after initially adopting them. She knows the general problem but is unsure which terminology different research traditions use.

Prompt She asks an LLM to suggest alternative concepts and terms related to discontinuance, abandonment, continued use, post-adoption behavior, and technology rejection.
Generated candidates The LLM proposes terms and combinations that could potentially broaden her search vocabulary.
Verification Instead of asking the LLM to invent a bibliography, she tests those terms in appropriate scholarly databases and examines how researchers actually use them in published literature.
Refinement She keeps useful terminology, discards irrelevant suggestions, identifies controlled vocabulary where applicable, and constructs database-specific search strategies.
Result The LLM helped with vocabulary generation. The scholarly databases supplied the literature, and the researcher evaluated the retrieved studies.

This division of labor exploits something LLMs can do well, rapidly generating plausible linguistic alternatives, without pretending that generated text is a substitute for systematic retrieval and scholarly evaluation.

05 · What Researchers Often Get Wrong

Common Misunderstandings About Large Language Models

Misconception

Is an LLM Basically a Giant Database?

No. Although training can enable a model to reproduce or reflect information encountered in data, an LLM should not be conceptualized as a conventional database containing neatly stored factual records waiting to be retrieved. It learns model parameters representing statistical patterns and generates outputs from those learned relationships and supplied context.

Misconception

If the LLM Knows About a Paper, Has It Read the Paper?

That wording can be misleading. A model may have learned patterns associated with a publication, may receive the paper through its current context, or may retrieve it through an external system. Those are different situations. Researchers should establish what information the system actually has access to rather than infer source access from a confident answer.

Misconception

If an LLM Gives a Citation, Did It Search for That Citation?

Not necessarily. Without a retrieval mechanism, an LLM may generate bibliographic-looking text. A citation should be treated as verified only after the publication and its relevant details have been checked against an authoritative source.

Misconception

Can I Eliminate Hallucinations by Writing a Better Prompt?

Better prompting can reduce some errors and make outputs more useful, but an instruction such as “do not hallucinate” is not a factual verification mechanism. For consequential claims, verification must occur outside the model's assertion that its own answer is correct.

Misconception

If the Model Explains Its Reasoning, Does That Prove How It Reached the Answer?

No. A generated explanation should not automatically be treated as a faithful audit trail of the model's internal computation. For research purposes, evaluate the answer using evidence, reproducible procedures, calculations, source material, or other appropriate external checks.

Misconception

If an LLM Is Better Than Me at a Task, Can I Delegate the Task Completely?

Performance alone does not settle whether delegation is appropriate. Consequential research tasks may involve accountability, disciplinary judgment, methodological justification, confidential information, or interpretation that researchers must retain. This becomes especially important when considering which research tasks should never be delegated completely to AI.

06 · What This Means for You

Use an LLM According to What You Can Verify

You do not need to understand every mathematical detail of transformer architectures to use an LLM responsibly. You do need a sufficiently accurate mental model of what the system is doing.

When an LLM produces something for research, classify the output before deciding what to do with it.

A simple decision framework

If the output is a suggestion
Evaluate whether it is useful. A suggestion can stimulate thinking without being treated as evidence.
If the output makes a factual claim
Check the claim against an authoritative source or underlying evidence.
If the output contains a citation
Verify that the source exists, that its bibliographic details are correct, and that it actually supports the claim for which you intend to cite it.
If the output contains code or a calculation
Run, inspect, test, and validate it using appropriate computational and methodological checks.
If the output interprets your research evidence
Return to the actual data, analysis, theory, and method. Do not substitute generated plausibility for scholarly interpretation.

The key is to exploit the model's generative capacity without quietly transferring your epistemic responsibility to it. The LLM can generate. You still have to determine what deserves to survive the generation process.

07 · A Quick Checklist

Before Using an LLM for a Research Task, Check These Points

Before relying on an LLM, check:
Identify whether the system is generating from the model alone, retrieving external information, using supplied documents, invoking tools, or combining these functions.
Decide whether you are asking for a candidate idea, a transformation, a factual claim, a calculation, an interpretation, or another type of output.
Verify factual and scholarly claims independently rather than relying on the model's confidence or writing quality.
Check every citation you intend to use against the actual publication and confirm that the source supports your claim.
Do not upload confidential, personal, proprietary, participant, peer-review, or unpublished material unless your use is permitted and the system's data handling is appropriate.
Test generated code, calculations, classifications, and transformations rather than assuming they work because they look technically sophisticated.
Be especially cautious when you lack enough subject expertise to recognize a plausible but incorrect answer.
Verify current institutional, funder, journal, and publisher rules governing the intended use.
Keep human responsibility for the research decisions and outputs that ultimately carry your name.
08 · Frequently Asked Questions

Frequently Asked Questions About LLMs in Research

Is ChatGPT an LLM?

ChatGPT is an AI application that uses underlying language models along with additional system components and capabilities. It is therefore more precise to distinguish the application you interact with from the model or models powering it.

Are all generative AI systems LLMs?

No. Generative AI includes systems designed to generate images, audio, video, and other content as well as language. LLMs are an important part of generative AI, but the categories are not synonymous.

Does an LLM understand what I write?

That depends heavily on what is meant by “understand,” a question that crosses technical, cognitive, and philosophical boundaries. For practical research use, do not infer human-like comprehension, judgment, or accountability merely from fluent behavior. Evaluate what the system can demonstrably do on the task that matters.

Does an LLM have access to the internet?

Not inherently. Internet or web access is a feature of the surrounding application or system. An LLM may be connected to search and retrieval tools, but you should not assume current web access simply because the interface can answer questions.

Can an LLM search research databases?

Only if the application has an appropriate connection or tool that actually performs that search. Generating a list of papers from model output is not equivalent to querying a scholarly database.

Why do LLMs sound confident when they are wrong?

The model generates language that fits learned patterns and context. Linguistic confidence is therefore not a calibrated guarantee of factual correctness. Research on hallucination has also shown why systems can be incentivized to guess rather than abstain when uncertain.

Can I use an LLM to write research code?

Yes, LLMs can assist with code generation, explanation, debugging, and translation between programming languages. Generated code should still be inspected, executed, tested, and validated, especially when it contributes to research results.

Should I trust an LLM more if it provides sources?

Source provision can make verification easier, particularly when the system actually retrieves relevant material, but citations do not make an answer automatically correct. Verify that each source exists, inspect the source itself, and confirm that it supports the claim attributed to it.

09 · The Bottom Line

Treat LLM Output as Generated Material Until You Establish More

The Bottom Line

A large language model learns statistical patterns in token sequences and uses those patterns to generate or process language, making it a powerful research assistant but not an automatic source of verified knowledge.

Use LLMs where generation, transformation, explanation, or exploration can help you work. When their outputs become factual, evidential, analytical, or methodologically consequential, move from generation to verification. The distinction is small in wording and enormous in research practice.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes