03 · What You Need to Know
An Expected Finding Is Most Informative When Alternatives Expected Something Else
Evidence Does More Work When Hypotheses Make Different Predictions
Imagine two explanations for the same phenomenon:
Explanation A: Frequent use of generative AI reduces independent problem-solving because students practice difficult cognitive tasks less often.
Explanation B: Students who already struggle with independent problem-solving subsequently use generative AI more frequently for assistance.
Both explanations predict a negative association between AI use and independent problem-solving performance in a cross-sectional dataset.
If that association appears, it is compatible with both. The result may be empirically useful, but it does relatively little to distinguish the causal stories.
Now suppose the study measures problem-solving ability before substantial AI use, tracks subsequent use, and measures later performance. The temporal patterns expected under the two explanations may differ. The evidence can potentially become more discriminating.
Competing Explanations Are More Than a Preferred Hypothesis and a Null
Researchers often frame hypothesis testing as a contest between “there is an effect” and “there is no effect.” For many substantive questions, however, the important competition is among different explanations for why an effect or pattern might exist.
A null hypothesis may tell you whether data are compatible with a particular statistical model. It does not automatically represent plausible substantive rivals such as reverse causation, confounding, a different mechanism, or an alternative theory.
This is why the older idea of strong inference remains useful. Platt argued for devising alternative hypotheses and designing experiments whose possible outcomes could exclude one or more of them, rather than repeatedly seeking observations compatible with one favored account.
First Identify the Alternatives Worth Distinguishing
Before designing a discriminating study, you need credible rivals. Start by asking what other processes could produce the finding you expect.
Some alternatives will concern substantive theory. Others may involve the causal direction running the other way, confounding, selection, or measurement.
Do not give every imaginable explanation equal status. Prioritize rivals with a credible mechanism, some grounding in prior evidence or theory, and the capacity to change your interpretation materially.
Find the Point Where the Predictions Diverge
The most useful question is not merely, “What are my competing hypotheses?” It is:
Under what observable conditions would these explanations predict different things?
Perhaps one predicts an immediate effect and another a delayed effect. One predicts change only for novices while another predicts the same pattern across expertise levels. One predicts a behavioral change while another predicts only a self-reported difference. One predicts that an association remains after a particular design control, while another predicts that it largely disappears.
The divergence may concern direction, magnitude, timing, dose-response pattern, subgroup differences, mediation, boundary conditions, or another observable implication.
A design becomes more informative when it deliberately captures that divergence.
A Study Can Fit Several Theories Without Strongly Supporting Any One of Them
This is a common interpretive problem. Researchers derive a broad prediction from a theory, observe it, and then describe the theory as supported.
Technically, the result may indeed be consistent with the theory. But if several credible theories made essentially the same prediction, the evidential gain for that particular explanation may be modest.
Watch Out
“Consistent with our theory” and “evidence that distinguishes our theory from plausible alternatives” are different claims. The first can be correct even when the study provides little discrimination among explanations.
More Data Do Not Automatically Create More Discrimination
A very large sample can estimate a shared prediction with great precision while remaining unable to distinguish the explanations producing it.
Suppose two theories both predict a correlation of roughly the same direction and magnitude between X and Y. Increasing the sample from 500 to 50,000 may estimate that correlation far more precisely. It does not necessarily tell you which theory is responsible.
Discrimination is therefore primarily a design and prediction problem, not merely a statistical power problem. Power matters once the study is aimed at a contrast capable of distinguishing the hypotheses.
Different Predictions Need to Be Specific Enough to Matter
Two explanations may technically make different predictions while remaining practically indistinguishable with the proposed measurements.
For example, one theory might predict an effect of 0.20 and another 0.21. Unless the study can estimate that difference with adequate precision and the predictions themselves are well justified, the theoretical distinction may provide little practical discrimination.
Recent methodological work on severe or informative testing similarly emphasizes the value of hypotheses whose predictions are sufficiently specific and contrastive for a design to generate discriminatory evidence.
Competing Explanations Can Change What You Measure
If every hypothesis predicts the same primary outcome, the most informative variable may be something you were not initially planning to collect.
Suppose two instructional theories both predict improved test scores. One attributes improvement to increased practice, while another attributes it to better conceptual understanding. Measuring only the final score provides limited discrimination. Measuring practice behavior, transfer to unfamiliar problems, retention, or process indicators may expose predictions on which the explanations differ.
Alternative explanations can therefore improve a design by revealing what evidence is missing.
Competing Explanations Can Change Who or What You Compare
Sometimes the crucial difference appears in a comparison condition rather than another variable.
If one explanation predicts that any additional attention improves performance while another predicts that a particular instructional mechanism matters, comparing the intervention only with no intervention may not separate them. An active comparison condition that controls for attention could be much more informative.
Likewise, subgroup comparisons, dosage contrasts, natural experiments, negative controls, or strategically selected cases may sometimes provide leverage, depending on the causal question and assumptions.
One Study Rarely Eliminates Every Plausible Explanation
The language of “ruling out” alternatives can become too strong. Empirical evidence is usually conditional on assumptions, measurement quality, design execution, statistical models, and the set of alternatives actually considered.
A study may make one explanation substantially less plausible without proving another uniquely correct. New explanations may also emerge as knowledge develops.
The practical goal is therefore often comparative rather than absolute: design evidence that changes the relative credibility of serious competing explanations.
Some Research Questions Do Not Require Explanation Discrimination
If your aim is to estimate how many students use generative AI, describe how a practice varies across institutions, develop a measurement instrument, forecast an outcome, or establish whether an intervention can be implemented, distinguishing causal explanations may not be the central purpose.
Even explanatory studies vary in ambition. Sometimes the immediate contribution is establishing that a phenomenon exists robustly before mechanisms can reasonably be separated.
This is why a study can remain worth doing even when it cannot distinguish every plausible explanation.
04 · A Practical Example
Turn One Shared Prediction Into a Test Between Explanations
Hypothetical Example
Why does retrieval practice improve later performance?
A researcher expects students who complete repeated retrieval activities to outperform students who only reread material on a later assessment. Two explanations are under consideration.
Explanation A Retrieval itself strengthens later access to the learned information.
Explanation B The advantage arises mainly because retrieval activities lead students to spend more time actively processing the material.
Shared prediction A simple retrieval group may outperform a passive rereading group. That result alone is compatible with both explanations.
Discriminating comparison The researcher adds an active comparison condition designed to more closely equate relevant time and effort without requiring retrieval.
Divergent predictions If retrieval has an effect beyond the relevant additional activity captured by the comparison, Explanation A predicts an advantage that should remain. If the difference was principally due to the controlled activity, Explanation B predicts substantial reduction of the retrieval advantage.
Interpretation The added comparison does not prove one mechanism uniquely correct, but it gives the result more power to distinguish the two explanations than the original two-group design.
The additional condition may require more participants, resources, and analytical planning. Its value comes from increasing what the study can teach, not simply making the design more elaborate.