01 · The Question
What If Researchers Measure Something Other Than What We Actually Need to Know?
A literature can contain dozens of carefully designed studies and still fail to answer its most important question.
The problem may be the outcome.
Researchers often need measurable indicators of complex phenomena. They may assess test scores instead of durable learning, intentions instead of actual behavior, satisfaction instead of effectiveness, engagement instead of achievement, biomarkers instead of outcomes experienced by patients, or immediate performance instead of long-term change.
Those measures are not automatically inappropriate. Each may answer a legitimate question. Trouble begins when the measured outcome and the conclusion researchers want to draw are not the same thing.
If the literature repeatedly measures what is convenient rather than what is consequential, the unanswered question may be hiding in plain sight.
03 · What You Need to Know
Separate What Studies Measure From What Researchers Want to Claim
An outcome is part of the research question, not an afterthought
An outcome is the result, state, event, behavior, experience, or other endpoint researchers observe to determine what happened. Choosing an outcome therefore determines what a study can actually tell you.
This sounds obvious until you examine published research. A study may ask whether an intervention improves learning but measure only students' satisfaction with it. Another may discuss behavioral change while measuring intentions. A third may claim improved well-being because participants report greater engagement.
Each measured variable may be useful. None is automatically interchangeable with the broader construct appearing in the conclusion.
When reviewing a literature, compare the substantive question with the outcomes researchers actually measured. A gap may emerge between them.
A proxy can be informative without being equivalent to the target outcome
Researchers frequently use proxy or surrogate outcomes because the outcome they ultimately care about is difficult, expensive, slow, rare, or otherwise challenging to measure.
This can be reasonable. The problem is inferential substitution: treating evidence about the proxy as though it were direct evidence about the target outcome.
Target outcome
The consequence the research question ultimately seeks to understand.
Proxy or surrogate outcome
A measurable outcome used as an indirect indicator of the target outcome.
GRADE treats outcome indirectness as one form of indirectness in a body of evidence. Concern arises when the outcome measured in available studies does not adequately correspond to the outcome specified in the question of interest. Surrogate outcomes are a particular source of this problem.
The important question is therefore not whether a proxy was used, but whether the evidence linking that proxy to the outcome of interest is strong enough for the inference being made.
Researchers may measure what is easy rather than what is meaningful
Some outcomes are attractive because they can be collected quickly. Knowledge tests, self-reported intentions, satisfaction ratings, platform activity, clicks, completion rates, and immediate post-intervention scores may be relatively inexpensive and convenient.
Convenience does not make an outcome invalid. It does, however, create a risk that the research agenda gradually becomes organized around what can be measured easily rather than what stakeholders actually need to know.
Suppose an educational technology literature repeatedly measures students' intention to use a platform. After 50 studies, researchers may understand intention extremely well while still knowing surprisingly little about whether sustained use improves learning.
The unanswered question is not hidden because nobody conducted research. It remains unanswered because the research concentrated on a different outcome.
An outcome can be relevant but still incomplete
Interventions can produce multiple consequences. Focusing on only one may create an incomplete answer.
A new teaching technology might increase short-term performance but also affect workload, student independence, equity, retention, or longer-term skill development. A health intervention may influence symptoms while also producing adverse effects. A workplace intervention may increase productivity while affecting employee well-being.
Cochrane guidance for systematic reviews reflects this multidimensionality by recommending specification of critical and important outcomes, normally including potential benefits and harms where relevant.
Thus, the question “Does it work?” is often too compressed. Work for what outcome, over what period, and at what possible cost?
Outcome heterogeneity can make a literature difficult to interpret
Another problem arises when studies ostensibly investigate the same question but measure different outcomes or define the same outcome differently.
One study might define academic success using course grades, another standardized achievement, another retention, and another self-reported learning. Pooling or comparing such findings requires careful judgment because the measures may not represent the same underlying outcome.
The COMET Initiative was established partly to address problems caused by inconsistent outcome selection and reporting. Core outcome sets specify agreed minimum outcomes that should be measured and reported in research on a particular condition or topic, while still allowing investigators to examine additional outcomes. Their purpose includes making studies easier to compare, combine, and interpret.
This does not mean every field needs identical measures. Standardization can improve comparability, but researchers should not achieve consistency by standardizing around an outcome that is poorly aligned with the substantive question.
Measurement timing is part of the outcome
“Writing performance immediately after an intervention” and “independent writing performance six months later” are not merely the same outcome measured on different calendar dates. They may represent substantively different claims.
Cochrane guidance recognizes outcome timing explicitly, recommending that reviews define relevant time categories when outcomes are synthesized.
If the important question concerns persistence, sustainability, delayed harms, or long-term development, a literature dominated by immediate outcomes may leave that question unresolved. In such cases, the problem overlaps with research in which follow-up is too short to observe the outcome that matters.
Outcome choice can determine whether effects appear meaningful
Researchers should also distinguish statistical detectability from substantive importance. A highly sensitive scale may detect a small change, but whether that change matters is a different question.
A literature may therefore provide abundant evidence that an intervention changes a score while leaving uncertain whether it produces an effect large enough to matter in practice. This becomes particularly important when research has focused on statistical significance rather than meaningful effects.
Outcome selection can create several different kinds of unanswered questions
| What studies measure |
What researchers want to know |
What may remain unanswered |
| Knowledge immediately after instruction |
Durable learning |
Whether learning is retained |
| Intention to use a technology |
Actual sustained use |
Whether intentions translate into behavior |
| Satisfaction |
Effectiveness |
Whether positive experiences produce better outcomes |
| Engagement metrics |
Learning or performance |
Whether greater activity produces meaningful improvement |
| A surrogate endpoint |
A consequential final outcome |
Whether changes in the surrogate translate into changes that matter |
| Benefits only |
Overall consequences |
Whether benefits outweigh relevant harms or burdens |
Different outcomes can also conceal apparently inconsistent evidence
Suppose one study finds a positive effect on engagement, another finds no effect on achievement, and a third finds a negative effect on workload. Calling the literature “inconsistent” may obscure the fact that the studies are measuring different consequences.
Before treating conflicting findings as evidence that nobody knows whether an intervention works, determine whether the apparent inconsistency comes from genuinely comparable evidence.
The strongest gap identifies the missing consequence, not merely the missing measure
“No study has used Scale X” is usually a weak gap statement unless Scale X captures something consequential that previous measures missed.
A more informative formulation is: “Existing studies establish that the intervention increases short-term task performance, but it remains uncertain whether students retain those gains when the assistance is removed.”
Now the missing outcome matters because it corresponds to an unresolved substantive question.
07 · A Quick Checklist
Check Whether Existing Studies Measure What Actually Matters
Before claiming an outcome-based research gap, check:
What outcome does the substantive research question actually require?
What outcomes do existing studies actually measure?
Are any commonly used measures proxies or surrogates for a more consequential outcome?
Is there adequate evidence that the proxy supports the inference researchers make from it?
Are important benefits, harms, experiences, behaviors, or longer-term consequences systematically omitted?
Are outcome definitions and measures sufficiently comparable across studies?
Is there an established core outcome set or reporting standard relevant to this research area?
Would measuring the missing outcome materially change what the literature allows researchers or stakeholders to conclude?