03 · What You Need to Know
A Correct Number Requires More Than Correct Arithmetic
When researchers say that a calculation is "correct," they may mean several different things. The arithmetic may be correct while the formula is inappropriate. The formula may be appropriate while the inputs are wrong. The inputs may be correct while the units are inconsistent. Or the numerical result may be accurate while the interpretation attached to it is misleading.
That is why numerical verification needs several layers rather than a single calculator check.
Verify the inputs first
Before recomputing anything, establish what numbers should actually enter the calculation.
AI may copy the wrong number from a prompt, confuse two values, omit a condition, use a rounded value where an exact value was required, or silently infer a missing input. If you reproduce a calculation using the same mistaken input, you have only reproduced the AI's error more neatly.
For research data, verify the source of each consequential input. Check the original dataset, table, manuscript result, or authoritative record rather than assuming that the values stated in the AI response are reliable.
Verify that the formula matches the quantity you want
A calculation should begin with the question, not the arithmetic operation.
For example, several percentages can be calculated from the same two numbers depending on what the denominator represents. A change score, percentage change, relative risk, odds ratio, standardized effect, and simple proportion are not interchangeable merely because each involves division.
If AI supplies a formula, establish independently that it corresponds to the quantity you actually intend to calculate.
Check the formula before checking the arithmetic
Arithmetic verification answers "Did these operations produce this number?" Formula verification answers the more important question "Were these the right operations to perform?"
For statistical quantities, this distinction can be substantial. An AI may correctly execute a mathematically defined expression that is not the appropriate statistic for your research design.
When the calculation is part of a statistical analysis, verify the underlying method separately. The dedicated process for checking an AI explanation of a statistical method addresses that methodological layer.
Check every operation, not just the final result
For nontrivial calculations, reconstruct the sequence.
Check additions and subtractions, multiplication and division, exponents, logarithms, square roots, transformations, weighted averages, conversions, and intermediate rounding. A single incorrect intermediate value can contaminate the final result while leaving it superficially plausible.
Hypothetical Example
A percentage change calculation
Suppose AI says that an outcome increased from 40 to 50 and therefore increased by 20%.
The arithmetic is:
(50 − 40) ÷ 40 × 100 = 25%
The model's 20% result is incorrect. More importantly, the verification also establishes that 40 is the baseline denominator. If the research question instead concerns percentage points, a relative increase, or a different reference value, the appropriate calculation could differ.
Check units and dimensions
Unit errors can survive arithmetic checks because the operations themselves may be performed perfectly.
A calculation mixing milligrams and grams, meters and centimeters, minutes and hours, or proportions and percentages can produce a numerically neat but substantively wrong result.
When units matter, carry them through the calculation. A result should have the units expected for the quantity being calculated.
Be careful with percentages, proportions, and percentage points
These are routinely confused even without AI.
Percentage
A value expressed per hundred, such as 25%.
Percentage-point change
The arithmetic difference between two percentages, such as 60% to 75% being a 15-percentage-point increase.
AI-generated explanations can move between these concepts without making the distinction explicit. Verify what quantity is actually being reported before accepting the number.
Check rounding and precision
A calculation can be mathematically correct while its reported precision is inappropriate. AI may round intermediate values unnecessarily, retain implausible precision, or report more decimal places than the underlying measurements justify.
Where rounding affects the result, recalculate using unrounded inputs and apply rounding at the appropriate stage according to the method or reporting convention you are using.
Recalculate using an independent method
The strongest simple test is to perform the calculation outside the original generated response.
Depending on the task, that could mean using a calculator, spreadsheet, statistical software, a short independently written script, symbolic algebra system, or manual arithmetic.
The point is not that one particular tool is inherently more trustworthy than AI. The point is that the second route provides an independent computational check rather than asking the same generative process to confirm itself.
Use simple known-answer cases for formulas or functions
If AI produces a reusable calculation, test it on inputs for which the expected result is obvious.
For example, a function calculating an average should return 5 for values 3 and 7. A percentage function should handle a zero change, a negative change, and other boundary cases according to its intended specification.
Simple cases are useful because they expose logic errors without requiring you to understand every part of a complex real dataset first.
Test edge cases
Numerical procedures often fail at special values.
Depending on the calculation, consider zero denominators, negative values, missing observations, empty samples, identical values, very large or very small numbers, proportions at the boundaries, and other conditions that are plausible in the actual research.
A formula that works for ordinary inputs can still behave incorrectly at the edges.
For statistical calculations, verify the interpretation separately
A correctly computed statistic can still be explained incorrectly.
For example, a p-value can be calculated accurately while AI incorrectly describes it as the probability that the null hypothesis is true. The American Statistical Association states that p-values do not measure the probability that the studied hypothesis is true and that statistical significance should not be treated as a measure of practical significance.
Thus, verification should continue beyond "the number matches." Ask what the number means, what assumptions produced it, and what conclusions it can support.
Check calculations embedded inside tables or figures
A generated table may contain individually plausible values whose relationships are inconsistent. Percentages may not correspond to denominators, totals may not add correctly, or labels may describe a different metric from the values shown.
Recompute important totals, percentages, rates, and derived statistics from the underlying data. If AI generated the figure or table itself, verify the source data and transformation as well.
The related process for verifying AI-generated research tables and figures applies when the numerical calculation is part of a larger visual or tabular output.
Check whether the result is plausible, but do not stop there
Sanity checks are useful. If a percentage is negative when the quantity cannot logically be negative, or a calculated count exceeds the total sample, investigate.
But plausibility is a screening mechanism, not proof. Wrong values often fall comfortably inside a plausible range.
More complex calculations may need more than a second calculator
For regression estimates, confidence intervals, simulation outputs, matrix operations, bootstrapping, numerical optimization, or other computationally involved procedures, independent verification may require reproducing the calculation with a trusted implementation and checking intermediate quantities.
In these cases, manual arithmetic may not be realistic. The principle remains the same: use a verification route whose assumptions and implementation can be independently understood.
Do not assume a displayed chain of reasoning guarantees correctness
AI may provide several lines of apparently transparent calculation and then arrive at the wrong result. NIST specifically notes that generative AI can produce confabulated logic that appears to justify an incorrect answer.
Therefore, inspect the operations themselves rather than treating the presence of "working" as evidence.
Keep the underlying data available when the calculation matters
A research result should be traceable back to the values from which it was derived. When an AI-generated calculation contributes to reported findings, preserve the source inputs and the computational procedure sufficiently to allow the result to be checked later.
That makes correction much easier if a discrepancy is discovered after the manuscript has moved forward.
Watch Out
Do not verify an AI calculation by asking the same model to "double-check the math." That may help identify an obvious inconsistency, but it is not independent verification because the second answer still comes from the same generative process.