03 · What You Need to Know
Running Without Errors Is the Beginning of Code Verification
Research code is part of the method by which computational results are produced. Nature Computational Science has argued that code review can reveal whether computational results are reproducible and whether the code is complete and reusable. Its code-review guidance recommends checking whether code and input data cover the reported experiments and whether rerunning them produces results consistent with the manuscript.
Nature Communications' computational reproducibility guidance similarly recommends validation against analytical results, standard approaches, or other appropriate tests for scientific code, alongside documentation, dependencies, tests, and examples.
These principles become especially relevant for AI-generated code because generation can produce syntactically plausible implementations whose substantive behavior the researcher has never inspected.
Start by defining what the code is supposed to do
Verification is impossible without a specification.
Before testing generated code, state the intended operation in plain language. Which data should enter the analysis? Which observations should be included or excluded? How should missing values be handled? What variables should be transformed? What statistical model should be fitted? What result should be returned?
If you cannot describe the intended analysis independently of the code, it becomes difficult to determine whether the AI implementation is correct.
Code executes
The programming language or runtime can process the instructions without terminating in an error.
Code is verified
Evidence shows that the implementation performs the intended research operations correctly under the conditions in which you plan to use it.
Read the code before trusting its output
Inspect the generated script rather than treating it as a black box.
Look for the data being loaded, filters being applied, variables being selected, transformations, loops, joins, statistical functions, parameters, randomization, output operations, and any assumptions embedded in the implementation.
You do not necessarily need to become a software engineer before using computational research methods. You do, however, need enough understanding to know what the code contributing to your research actually does. If the implementation is beyond your ability to assess, that is a reason to obtain appropriate technical review rather than a reason to trust the generator.
Check every unfamiliar function against official documentation
AI can invent functions, use deprecated interfaces, supply nonexistent arguments, or call a real function incorrectly.
For every consequential function you do not recognize, consult the official documentation for the relevant language, library, package, and version. Verify the argument meanings, defaults, return values, handling of missing data, and other behavior relevant to the analysis.
This is especially important when generated code uses a package that has changed across versions. Code that was correct for one version may behave differently or fail under another.
Inspect the data before and after each important transformation
Many research-code errors occur before the headline statistical model is fitted.
A generated pipeline might:
- remove rows containing missing values;
- convert categories into unintended numeric codes;
- reverse a filter condition;
- duplicate observations during a join;
- aggregate at the wrong level;
- standardize variables using inappropriate data;
- overwrite the original dataset;
- silently change date or text formats.
Check dimensions, variable types, category counts, missingness, ranges, identifiers, duplicates, and selected observations after important processing steps. Do not wait until the final model output to discover that the dataset entering it is not the dataset you intended.
Use small inputs whose correct answers you already know
One of the strongest practical tests of generated code is to create a small controlled example for which you can determine the expected result independently.
If AI generates a function that calculates percentages, give it simple counts. If it recodes categories, create a handful of observations covering every category. If it joins datasets, use tiny tables where the expected matches are obvious.
Known-answer tests are powerful because they verify behavior rather than appearance.
Test edge cases, not just the happy path
Code may work correctly for ordinary inputs and fail at boundaries.
Depending on the task, test missing values, zero values, empty groups, duplicated identifiers, unusual categories, extreme observations, very small samples, unexpected data types, or other plausible conditions.
The relevant edge cases depend on the research workflow. You do not need to invent every pathological input imaginable, but you should test conditions that could realistically occur in your data.
Compare important results with an independent implementation
For consequential analyses, reproduce a subset of the result using another defensible route where feasible.
You might manually calculate a simple statistic, compare the result with a trusted statistical package, implement a small example independently, or use a known benchmark dataset with expected results.
Nature Communications' computational reproducibility guidance specifically identifies comparisons with analytical results or standard approaches as useful forms of validation for scientific code.
This is the computational analogue of independent verification of AI-assisted research work: the verification route should provide evidence beyond the generated implementation itself.
Do not let AI-generated tests merely reproduce AI-generated mistakes
You can ask AI to suggest tests, and those suggestions may be useful. But there is a circularity risk if the same system generates the code, decides what the code should do, generates the tests, and then announces that the tests passed.
At least some expected results should come from an independent specification, known example, authoritative documentation, manually verifiable case, benchmark, or another defensible source.
Check the statistical method independently of the programming
Correct programming does not rescue an inappropriate analysis.
An AI-generated script might flawlessly implement an independent-samples t-test for data that require a paired analysis. Every line of code can be technically correct while the research method is wrong.
When generated code implements a statistical procedure, separately verify the statistical method and its assumptions. Programming verification and methodological verification are related but not interchangeable.
Inspect intermediate results
A final number can conceal where an error entered the pipeline.
For multi-step analyses, inspect important intermediate objects: filtered datasets, transformed variables, group counts, model matrices, train-test partitions, summary statistics, predictions, and other outputs appropriate to the workflow.
This helps localize errors and makes the analysis easier to reason about. It also reduces the temptation to judge code solely by whether the final result looks plausible.
Check randomness and reproducibility
Some research workflows contain random processes: train-test splits, bootstrapping, permutation procedures, simulations, imputation, stochastic optimization, or model initialization.
Where reproducibility is required, determine how random seeds and other sources of nondeterminism are handled. Record relevant software versions, dependencies, configurations, and computational details.
Reproducibility guidance for machine-learning research has emphasized reporting code, data, model details, dependencies, operating systems, and deterministic settings where appropriate. More recent guidance concerning LLM-based research likewise emphasizes recording model versions, prompts, configurations, scripts, and system specifications when these affect the computational workflow.
Record the software environment
Generated code does not run in a vacuum. Results can depend on language versions, package versions, operating systems, external libraries, hardware, or other environmental details.
Nature Computational Science has specifically highlighted software-library versions and hardware architectures as potential sources of inconsistency in computational research.
For consequential analyses, record enough information to recreate the environment. Depending on the project, that may involve a dependency file, environment specification, container, lock file, or equivalent record.
Verify that the complete workflow reproduces the reported result
Individual functions can pass tests while the full analysis pipeline remains wrong.
Run the workflow from its intended starting point using the appropriate data and determine whether it reproduces the tables, figures, statistics, model outputs, or other results you plan to report.
Nature Computational Science's code-review guidance recommends rerunning code associated with experiments and comparing the resulting output with what is reported in the manuscript. If the results are inconsistent, the discrepancy itself requires investigation.
Reproducibility is necessary but not identical to correctness
Code can reproduce the same wrong answer perfectly.
If a script systematically applies an incorrect formula, rerunning it tomorrow will faithfully reproduce the error. Reproducibility therefore answers an important question about whether results can be regenerated, but analytical validity requires additional checks against expected behavior, methodological requirements, and independent evidence.
| Verification layer |
Question |
Useful check |
| Syntax and execution |
Does the code run? |
Execute in the intended environment |
| Function behavior |
Do functions do what the code assumes? |
Official documentation |
| Data handling |
Are the correct observations and variables being processed? |
Inspect intermediate data and counts |
| Algorithmic correctness |
Does the implementation perform the intended operation? |
Known-answer tests and benchmarks |
| Methodological validity |
Is the implemented analysis appropriate? |
Independent methodological verification |
| Result validity |
Are outputs consistent with independent expectations? |
Alternative calculations or standard approaches |
| Reproducibility |
Can the workflow regenerate the reported results? |
Rerun from documented inputs and environment |
Security and privacy still matter in research code
Before running generated code, inspect operations involving files, credentials, network access, package installation, system commands, external services, and sensitive datasets.
Do not paste confidential research data, participant information, credentials, API keys, or restricted material into tools or services without ensuring that doing so complies with your ethical, institutional, contractual, and data-governance requirements.
Code verification is therefore partly about scientific correctness and partly about understanding what instructions you are allowing the generated program to execute.
Watch Out
A plausible final result is weak evidence that generated code is correct. Incorrect code often produces ordinary-looking numbers, tables, and figures. Verify the operations that produced the result rather than judging correctness from the appearance of the output.