03 · What You Need to Know
How Large Language Models Work and What That Means for Research
“Large,” “Language,” and “Model” Each Tell You Something
The name is less mysterious once it is unpacked.
Large
Modern LLMs are trained using enormous datasets and computational resources and may contain very large numbers of learned parameters.
Language
They learn statistical structure in sequences represented as tokens and are particularly capable of processing and generating language, although modern models and systems may also work with images, audio, code, and other modalities.
Model
They are computational models that learn patterns from training data and use those learned relationships to produce outputs. They are not databases containing a neat record for every sentence they can generate.
An LLM is therefore one particular technology within the much broader landscape of AI in research. Not every AI system is an LLM, and not every form of generative AI is language-based.
LLMs Work With Tokens, Not Words in Quite the Way Humans Do
Before text can be processed by an LLM, it is represented as units commonly called tokens. A token might correspond to a whole word, part of a word, punctuation, or another textual unit depending on the tokenizer used by the model.
When an LLM generates text, it operates over these token representations. Given the preceding context, the model assigns probabilities to possible subsequent tokens and generation proceeds token by token.
This simple description leaves out a great deal of technical machinery, but it captures a crucial point for researchers: the system is generating a continuation based on learned relationships and supplied context, not consulting an internal librarian who checks every sentence against the scholarly record before allowing it to appear.
Predicting Tokens Can Produce Surprisingly Sophisticated Behavior
“It predicts the next token” can sound as though an LLM is little more than predictive text on a phone. That comparison badly understates what large models can do.
To predict sequences well across diverse material, large models can acquire rich statistical representations of language, concepts, relationships, styles, programming structures, and many other patterns represented in their training and subsequent adaptation. Modern LLM-based systems can consequently perform tasks that look strikingly different from simple sentence completion.
They may summarize text, translate languages, generate computer code, classify material, transform writing, answer questions, extract structured information, compare arguments, follow multi-step instructions, or produce explanations.
Researchers should therefore avoid two opposite mistakes: pretending LLMs are trivial autocomplete, and imagining that fluent performance means they possess a human researcher's understanding, evidential judgment, or accountability.
An LLM Is Not the Same Thing as the Chatbot You Use
This distinction is easy to miss.
The LLM is the underlying model. The application through which you interact with it may add many other components: system instructions, safety mechanisms, document retrieval, web search, memory, calculators, code execution, image processing, databases, or external tools.
| Layer |
What it may do |
Why researchers should care |
| Large language model |
Processes context and generates token sequences |
Explains many capabilities and limitations of generated language |
| Application |
Provides the interface, instructions, permissions, file handling, and workflow |
Two applications using related models may behave differently |
| Retrieval or search |
Supplies external documents or current information to the model |
An answer grounded in retrieved sources differs from one generated without them |
| External tools |
May perform calculations, execute code, query databases, or complete other operations |
Not every result shown by an AI assistant was generated by the LLM itself |
This matters when evaluating an answer. If an AI assistant correctly calculates a complicated statistic, you may need to know whether the LLM generated the number or whether the application invoked a computational tool. If it provides current literature, you need to know whether it actually searched a scholarly source or merely generated plausible citations.
Training Is Not the Same as Search
An LLM may encode a remarkable amount of information through patterns learned during training, but asking it a question is not automatically equivalent to searching the web, querying a scholarly database, or retrieving a paper.
A research database typically returns records matching a query from a defined collection. A search engine retrieves and ranks indexed resources. A reference manager stores and organizes bibliographic information. An LLM generates an answer from its model and context unless the surrounding application gives it access to retrieval or other external tools.
That distinction is central to understanding how AI assistants differ from search engines, research databases, and reference managers.
If you ask an ungrounded LLM, “What studies show that intervention X improves outcome Y?”, its ability to produce author-like names, article-like titles, and journal-like references does not mean those publications were retrieved from a bibliographic database.
Why LLMs Can Hallucinate
Researchers commonly use the term hallucination for generated information that is false, unsupported, nonsensical, or otherwise inaccurate despite being presented as an answer.
Hallucination is not merely a matter of an LLM occasionally having a bad day. It arises from the nature of systems optimized to generate plausible continuations rather than guarantee that every generated proposition corresponds to a verified fact. Research published in Nature has examined why confident guessing can persist in language models and how common evaluation practices can reward answering rather than appropriately abstaining.
For researchers, the practical consequence matters more than the terminology: an LLM can produce a grammatically excellent, highly specific, academically styled statement that is false.
Watch Out
Specificity is not verification. An LLM can generate precise dates, quotations, statistics, DOI-like strings, article titles, methodological details, or citations without those details being correct. Verify consequential information independently.
More Capable Models Do Not Eliminate the Need for Verification
LLMs have improved substantially, and surrounding systems can reduce some failure modes through retrieval, tool use, better training, improved evaluation, and other techniques. Researchers should not assume that every limitation observed in an earlier generation applies unchanged to every newer system.
But improving average performance does not turn generated language into evidence. Even highly capable systems can make errors, misrepresent sources, misunderstand ambiguous instructions, omit important qualifications, or generate unsupported conclusions.
This is why responsible use should be calibrated to consequences. A mistaken suggestion for an alternative manuscript title is inconvenient. A fabricated reference in a systematic review, an incorrect interpretation of an inferential test, or an invented claim about participant data can affect the integrity of the research.
Context Matters Because the Model Responds to What It Can Access
An LLM does not respond to your prompt in isolation. Its effective context may include system instructions, conversation history, uploaded documents, retrieved material, tool outputs, and other information supplied by the application.
This explains why providing a model with the actual text of an article can produce a different task from asking, “What does Smith et al. say about this?” without supplying or retrieving that article. In the first case, the system may have the source in its working context. In the second, you may be asking it to generate an answer without reliable access to the publication.
Even source-provided tasks still require checking. A model can overlook qualifications, confuse sections, overgeneralize findings, or generate an interpretation not supported by the document.
Prompting Changes the Output, but It Does Not Change the Evidence
Instructions matter. A clearer prompt can specify the task, context, constraints, format, audience, or evidence the system should use. That can substantially improve usefulness.
But prompt engineering has a limit that researchers should keep firmly in view. Telling a model to “only provide real references,” “do not hallucinate,” or “make sure everything is scientifically accurate” does not itself verify the resulting answer.
A stronger instruction can influence model behavior. It cannot transform an unsupported generated statement into evidence.
Researchers Can Use LLMs Most Productively as Assistive Systems
The versatility of LLMs makes them useful across many research activities. Depending on the system, task, data restrictions, and applicable policies, they may help researchers:
- explore terminology and alternative ways to formulate a problem;
- explain unfamiliar concepts as a starting point for further verification;
- generate or debug code that the researcher can inspect and test;
- transform researcher-supplied text into another format, structure, or level of explanation;
- compare material supplied in the context;
- extract candidate information from documents for subsequent checking;
- generate possible counterarguments, questions, interpretations, or analytical approaches;
- support drafting, revision, translation, or language editing where permitted.
The broader range of research tasks generative AI can support is growing as systems gain new modalities and tools. Capability, however, should not be confused with permission or methodological suitability.
The Best Mental Model Is Not “Oracle”
For many research tasks, a productive way to think about an LLM is as a system capable of generating useful candidate outputs extremely quickly.
Candidate is the important word. A proposed search term can be tested. Generated code can be run and inspected. A suggested explanation can be investigated. A rewritten paragraph can be compared with the original. A proposed counterargument can force you to reconsider your reasoning.
Problems arise when candidate silently becomes answer, and then answer silently becomes evidence.