Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

What Questions Remain Unanswered Because Studies Measure the Wrong Outcomes?

A study can be rigorous and still leave the important question unanswered if it measures the wrong outcome. Learn how to identify gaps created by outcome choice.

648
Unanswered Questions From Wrong Outcomes Guide 648 of 899
01 · The Question

What If Researchers Measure Something Other Than What We Actually Need to Know?

A literature can contain dozens of carefully designed studies and still fail to answer its most important question.

The problem may be the outcome.

Researchers often need measurable indicators of complex phenomena. They may assess test scores instead of durable learning, intentions instead of actual behavior, satisfaction instead of effectiveness, engagement instead of achievement, biomarkers instead of outcomes experienced by patients, or immediate performance instead of long-term change.

Those measures are not automatically inappropriate. Each may answer a legitimate question. Trouble begins when the measured outcome and the conclusion researchers want to draw are not the same thing.

If the literature repeatedly measures what is convenient rather than what is consequential, the unanswered question may be hiding in plain sight.

02 · The Short Answer

Check Whether the Measured Outcome Actually Represents the Question

In Brief

Questions remain unanswered because of the wrong outcomes when existing studies measure variables that do not adequately represent the effect, experience, behavior, benefit, harm, or other consequence that the research question actually concerns.

The problem may involve surrogate outcomes, convenient proxies, incomplete outcome sets, inconsistent definitions, inappropriate measurement timing, or outcomes selected without sufficient attention to what matters to relevant stakeholders. The research gap is the missing evidence about the outcome of interest, not merely the absence of another study.

03 · What You Need to Know

Separate What Studies Measure From What Researchers Want to Claim

An outcome is part of the research question, not an afterthought

An outcome is the result, state, event, behavior, experience, or other endpoint researchers observe to determine what happened. Choosing an outcome therefore determines what a study can actually tell you.

This sounds obvious until you examine published research. A study may ask whether an intervention improves learning but measure only students' satisfaction with it. Another may discuss behavioral change while measuring intentions. A third may claim improved well-being because participants report greater engagement.

Each measured variable may be useful. None is automatically interchangeable with the broader construct appearing in the conclusion.

When reviewing a literature, compare the substantive question with the outcomes researchers actually measured. A gap may emerge between them.

A proxy can be informative without being equivalent to the target outcome

Researchers frequently use proxy or surrogate outcomes because the outcome they ultimately care about is difficult, expensive, slow, rare, or otherwise challenging to measure.

This can be reasonable. The problem is inferential substitution: treating evidence about the proxy as though it were direct evidence about the target outcome.

Target outcome The consequence the research question ultimately seeks to understand.
Proxy or surrogate outcome A measurable outcome used as an indirect indicator of the target outcome.

GRADE treats outcome indirectness as one form of indirectness in a body of evidence. Concern arises when the outcome measured in available studies does not adequately correspond to the outcome specified in the question of interest. Surrogate outcomes are a particular source of this problem.

The important question is therefore not whether a proxy was used, but whether the evidence linking that proxy to the outcome of interest is strong enough for the inference being made.

Researchers may measure what is easy rather than what is meaningful

Some outcomes are attractive because they can be collected quickly. Knowledge tests, self-reported intentions, satisfaction ratings, platform activity, clicks, completion rates, and immediate post-intervention scores may be relatively inexpensive and convenient.

Convenience does not make an outcome invalid. It does, however, create a risk that the research agenda gradually becomes organized around what can be measured easily rather than what stakeholders actually need to know.

Suppose an educational technology literature repeatedly measures students' intention to use a platform. After 50 studies, researchers may understand intention extremely well while still knowing surprisingly little about whether sustained use improves learning.

The unanswered question is not hidden because nobody conducted research. It remains unanswered because the research concentrated on a different outcome.

An outcome can be relevant but still incomplete

Interventions can produce multiple consequences. Focusing on only one may create an incomplete answer.

A new teaching technology might increase short-term performance but also affect workload, student independence, equity, retention, or longer-term skill development. A health intervention may influence symptoms while also producing adverse effects. A workplace intervention may increase productivity while affecting employee well-being.

Cochrane guidance for systematic reviews reflects this multidimensionality by recommending specification of critical and important outcomes, normally including potential benefits and harms where relevant.

Thus, the question “Does it work?” is often too compressed. Work for what outcome, over what period, and at what possible cost?

Outcome heterogeneity can make a literature difficult to interpret

Another problem arises when studies ostensibly investigate the same question but measure different outcomes or define the same outcome differently.

One study might define academic success using course grades, another standardized achievement, another retention, and another self-reported learning. Pooling or comparing such findings requires careful judgment because the measures may not represent the same underlying outcome.

The COMET Initiative was established partly to address problems caused by inconsistent outcome selection and reporting. Core outcome sets specify agreed minimum outcomes that should be measured and reported in research on a particular condition or topic, while still allowing investigators to examine additional outcomes. Their purpose includes making studies easier to compare, combine, and interpret.

This does not mean every field needs identical measures. Standardization can improve comparability, but researchers should not achieve consistency by standardizing around an outcome that is poorly aligned with the substantive question.

Measurement timing is part of the outcome

“Writing performance immediately after an intervention” and “independent writing performance six months later” are not merely the same outcome measured on different calendar dates. They may represent substantively different claims.

Cochrane guidance recognizes outcome timing explicitly, recommending that reviews define relevant time categories when outcomes are synthesized.

If the important question concerns persistence, sustainability, delayed harms, or long-term development, a literature dominated by immediate outcomes may leave that question unresolved. In such cases, the problem overlaps with research in which follow-up is too short to observe the outcome that matters.

Outcome choice can determine whether effects appear meaningful

Researchers should also distinguish statistical detectability from substantive importance. A highly sensitive scale may detect a small change, but whether that change matters is a different question.

A literature may therefore provide abundant evidence that an intervention changes a score while leaving uncertain whether it produces an effect large enough to matter in practice. This becomes particularly important when research has focused on statistical significance rather than meaningful effects.

Outcome selection can create several different kinds of unanswered questions

What studies measure What researchers want to know What may remain unanswered
Knowledge immediately after instruction Durable learning Whether learning is retained
Intention to use a technology Actual sustained use Whether intentions translate into behavior
Satisfaction Effectiveness Whether positive experiences produce better outcomes
Engagement metrics Learning or performance Whether greater activity produces meaningful improvement
A surrogate endpoint A consequential final outcome Whether changes in the surrogate translate into changes that matter
Benefits only Overall consequences Whether benefits outweigh relevant harms or burdens

Different outcomes can also conceal apparently inconsistent evidence

Suppose one study finds a positive effect on engagement, another finds no effect on achievement, and a third finds a negative effect on workload. Calling the literature “inconsistent” may obscure the fact that the studies are measuring different consequences.

Before treating conflicting findings as evidence that nobody knows whether an intervention works, determine whether the apparent inconsistency comes from genuinely comparable evidence.

The strongest gap identifies the missing consequence, not merely the missing measure

“No study has used Scale X” is usually a weak gap statement unless Scale X captures something consequential that previous measures missed.

A more informative formulation is: “Existing studies establish that the intervention increases short-term task performance, but it remains uncertain whether students retain those gains when the assistance is removed.”

Now the missing outcome matters because it corresponds to an unresolved substantive question.

04 · A Practical Example

When Positive Outcomes Still Do Not Answer Whether Students Learned

Hypothetical Example

Does generative AI improve students' writing ability?

Imagine that a hypothetical literature contains 30 studies of generative AI-assisted writing. Most report favorable results.

What studies measure Many studies assess the quality of texts students produce while they are permitted to use AI assistance. Others measure satisfaction, perceived usefulness, or intention to continue using the tool.
What the broader claim implies Claims that AI improves students' writing ability imply something about what students themselves can subsequently do, not merely the quality of a jointly produced AI-assisted text.
What remains uncertain The literature provides much less evidence about whether students become better independent writers when the tool is unavailable.
What the next study needs A study designed around independent writing performance, transfer to new tasks, or retained skill could address the unresolved learning outcome more directly.

The existing studies need not be wrong. They may provide useful evidence about AI-assisted performance and user experience. The mistake would be treating those outcomes as complete answers to a different question about skill acquisition.

05 · What Researchers Often Get Wrong

Common Mistakes When Looking for Outcome-Based Gaps

Misconception

If an Outcome Is Valid, It Must Answer the Research Question

A measure can validly assess one construct while remaining inappropriate for another. A valid satisfaction scale measures satisfaction. That does not make satisfaction a substitute for learning, effectiveness, behavior, or another target outcome.

Misconception

Surrogate Outcomes Are Always Bad

Surrogates can be useful when final outcomes are difficult or slow to observe. The problem arises when researchers infer effects on the target outcome without adequate justification that changes in the surrogate provide reliable information about that outcome.

Misconception

Researchers Should Measure Every Possible Outcome

No study can measure everything, nor should it try. Outcome selection should be driven by the research question, relevant theory, stakeholder priorities, feasibility, and the consequences necessary for interpreting the intervention or phenomenon.

Misconception

Using the Same Outcome as Previous Studies Is Automatically Best

Consistency can improve comparability, but inherited measures should still be examined critically. A field can repeatedly measure the same convenient outcome while neglecting a more consequential one.

Misconception

A New Instrument Automatically Creates a Novel Contribution

Using a different scale does not necessarily address a meaningful gap. The contribution depends on whether the new measurement captures an important aspect of the question that existing evidence cannot adequately address.

06 · What This Means for You

Audit the Outcomes Before Proposing Another Study

When reviewing a literature, create a simple outcome map. Write down the substantive claims researchers want to make, then record what their studies actually measure.

The mismatch between those two columns can be more revealing than the number of studies.

A simple decision framework

If studies rely on a proxy
Ask how strongly the proxy supports inference about the outcome you actually care about.
If studies measure only immediate outcomes
Determine whether the substantive question requires retention, persistence, delayed consequences, or longer-term effects.
If studies emphasize convenient outcomes
Ask which outcomes matter to the people, decisions, theories, or practices the research is intended to inform.
If studies use many incompatible outcomes
Determine whether a core outcome set, common outcome domain, or better harmonization could improve comparability without erasing meaningful differences.
If an important outcome is genuinely missing
Explain what conclusion cannot currently be drawn because that outcome has not been adequately measured.

This reframes your contribution. You are not adding an outcome merely because previous researchers omitted it. You are collecting the evidence necessary to answer a question their outcome choices left unresolved.

07 · A Quick Checklist

Check Whether Existing Studies Measure What Actually Matters

Before claiming an outcome-based research gap, check:
What outcome does the substantive research question actually require?
What outcomes do existing studies actually measure?
Are any commonly used measures proxies or surrogates for a more consequential outcome?
Is there adequate evidence that the proxy supports the inference researchers make from it?
Are important benefits, harms, experiences, behaviors, or longer-term consequences systematically omitted?
Are outcome definitions and measures sufficiently comparable across studies?
Is there an established core outcome set or reporting standard relevant to this research area?
Would measuring the missing outcome materially change what the literature allows researchers or stakeholders to conclude?
08 · Frequently Asked Questions

Questions About Outcome-Based Research Gaps

What makes an outcome the “wrong” outcome?

An outcome is problematic when it does not adequately address the substantive question or conclusion for which it is being used. The same outcome may be entirely appropriate for a different research question.

What is the difference between an outcome and an outcome measure?

The outcome is what researchers want to assess, such as anxiety, academic achievement, or quality of life. The outcome measure is the instrument, procedure, or operational method used to measure it. Researchers can agree on an outcome while using different measures to assess it.

Are self-reported outcomes inferior to objective outcomes?

No. Some constructs, such as perceptions, symptoms, attitudes, and experiences, may appropriately require self-report. The issue is whether the method measures the outcome of interest accurately enough for the intended inference, not whether it belongs to a supposedly superior category of measurement.

Should I choose an outcome nobody has measured before?

Novelty alone is insufficient. A new outcome is most useful when it captures a consequential aspect of the question that existing research has inadequately represented.

What is a core outcome set?

A core outcome set is an agreed minimum set of outcomes recommended for measurement and reporting in studies of a particular topic or condition. Researchers may measure additional outcomes as well. Core outcome sets are intended to improve relevance and comparability across studies.

Can measuring different outcomes explain conflicting findings?

Yes. Two studies may appear to disagree while actually estimating effects on different consequences, at different times, or using substantially different operational definitions. Compare the outcomes carefully before interpreting findings as contradictory.

09 · The Bottom Line

Ask Whether the Literature Measured the Consequence You Care About

The Bottom Line

A question can remain unanswered despite extensive research when studies repeatedly measure outcomes that do not adequately represent the consequence the question is actually asking about.

Look beyond whether an outcome was measured at all. Examine what it represents, when it was measured, whether important consequences were omitted, and whether it supports the conclusion researchers want to draw. Sometimes the most important gap is not another study, but evidence about the outcome that mattered all along.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes