Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

Will the Proposed Study Add Better Measurement?

Using a different instrument does not automatically improve a study. Better measurement matters when it captures the intended construct or outcome more appropriately and changes the credibility, precision, or interpretation of the resulting evidence.

729
Will the Study Measure It Better? Guide 729 of 899
01 · The Question

Will a new measure actually produce better evidence?

Previous studies may appear to measure the right thing while doing so imperfectly. Researchers may rely on a single self-report item for a complex construct, use an instrument developed for another population, infer behavior from a weak proxy, apply a measure without adequate evidence for its intended interpretation, or use an outcome that captures only part of what matters.

This can create a legitimate reason for another study. But replacing an old measure with a newer, longer, more technical, or more convenient instrument does not automatically improve the evidence.

The real question is whether the proposed measurement approach better represents the construct, behavior, exposure, or outcome needed for the research question and whether that improvement changes what can reasonably be inferred from the data.

02 · The Short Answer

Better measurement should improve the inference, not merely change the instrument

In Brief

A proposed study adds better measurement when its measurement strategy more appropriately captures the intended construct or outcome, reduces consequential measurement problems, or provides information that existing measures could not support adequately.

Do not assume that a validated, objective, digital, longer, or newly developed measure is automatically better. Measurement quality depends on what you intend to measure, the population and context, the interpretation you want to make from the scores or observations, and the evidence supporting that use.

03 · What You Need to Know

How to decide whether a new study genuinely measures something better

Begin with what you need to measure, not with the instrument

A surprisingly common measurement problem begins with instrument selection. Researchers find an established questionnaire, scale, platform metric, administrative variable, sensor, or test and then organize the construct around whatever that instrument happens to provide.

The stronger sequence runs in the opposite direction.

First define the construct, behavior, exposure, outcome, or attribute required by the research question. Then determine what observations would represent it adequately. Only after that should you decide which instrument or measurement procedure can generate those observations.

COSMIN guidance on outcome measurement similarly places the construct at the center of instrument selection and treats content validity as particularly important because researchers first need confidence that an instrument adequately reflects what it is intended to measure.

Measurement quality has several dimensions

Calling a measure “valid and reliable” can conceal more than it explains. Measurement quality is not one property.

COSMIN, for example, distinguishes properties including content validity, structural validity, internal consistency, cross-cultural validity or measurement invariance, reliability, measurement error, criterion validity, construct validity, and responsiveness. Which properties matter depends on the instrument and intended use.

Measurement issue Core question Why it may matter
Content validity Does the measure adequately represent the construct it is intended to measure? A precise score is of limited value if important parts of the construct are absent or irrelevant content is included.
Structural validity Does the score structure reflect the dimensionality of the construct? Combining items into a score may be difficult to justify if the assumed structure is unsupported.
Reliability Can relevant differences among people or units be distinguished consistently from measurement error? Poor reliability can obscure real differences and attenuate relationships.
Measurement error How much observed variation arises from error rather than the quantity being measured? Error can reduce precision and complicate interpretation of differences or change.
Construct validity Do scores behave as expected given the construct they are intended to represent? Evidence is needed to support interpretations based on the measured construct.
Responsiveness Can the instrument detect relevant change over time? An instrument useful for distinguishing people at one time point may not adequately detect change.

These concepts come from measurement frameworks with particular terminology and applications, and other disciplines may organize validity evidence differently. The broader lesson is portable: “better measurement” needs to specify what property is better and why that property matters to the research inference.

Validity belongs to an interpretation and use, not merely to the instrument's name

A familiar statement in research methods is that an instrument “has been validated.” That can be misleading if interpreted as permanent certification that the instrument is valid everywhere and for every purpose.

The Standards for Educational and Psychological Testing, jointly developed by the American Educational Research Association, American Psychological Association, and National Council on Measurement in Education, frame validity in terms of evidence and theory supporting interpretations of test scores for proposed uses. The Standards also emphasize that different interpretations or uses can require different validity evidence.

An instrument supported for screening one population may not automatically support a different diagnostic, evaluative, predictive, or research interpretation in another population. Translation, cultural adaptation, administration mode, scoring changes, and context can also affect what evidence is needed.

Thus, “we will use a validated questionnaire” is not the end of the measurement argument. Ask: validated for what interpretation, in whom, under what conditions, and how closely does that match the proposed use?

Reliability is not the same as validity

A measure can produce highly consistent scores while systematically failing to represent the intended construct adequately.

Imagine a bathroom scale that consistently reports the same person's weight several kilograms too high. Consistency alone does not make the measurement accurate. In psychological and educational measurement, the situation is usually more complex because many constructs do not have a simple physical reference standard, but the conceptual distinction remains useful.

COSMIN explicitly treats reliability and forms of validity as distinct measurement properties.

This matters when justifying new research. If previous evidence is limited because a construct was poorly represented, showing that the old measure had high internal consistency does not necessarily resolve the problem.

“Objective” measurement is not automatically superior to self-report

Researchers sometimes create a hierarchy in which device logs, administrative records, sensors, or behavioral traces are considered inherently superior to questionnaires or interviews.

The appropriate measurement depends on the construct.

If you want to know how many times a student opened an application, a reliable system log may be more direct than retrospective self-report. If you want to understand perceived usefulness, anxiety, confidence, motivation, or an experience that is inherently subjective, a behavioral log cannot simply replace participants' reports.

Even digital records require interpretation. Login counts may not represent meaningful use. Time with a browser tab open may not equal active engagement. Clicks can be measured with exquisite precision while remaining a questionable proxy for learning.

Objective-looking data can therefore be precisely measured proxies for something other than the construct you actually care about.

A proxy measure needs a defensible link to the target construct

Many important constructs cannot be observed directly. Researchers use indicators or proxies.

This is not inherently problematic. The problem arises when the proxy quietly becomes the construct.

Publication count is not identical to research quality. Platform activity is not identical to engagement. Course grades are not identical to learning. Citation counts are not identical to societal impact. Each may provide useful information for particular questions, but the inference depends on how well the indicator represents the target concept.

If previous studies rely heavily on a proxy whose relationship to the intended construct is weak or uncertain, a new measurement strategy may genuinely strengthen the evidence.

Better measurement can reduce one source of uncertainty without fixing the entire study

Suppose previous research uses a weak self-report measure of an exposure. Replacing it with a more defensible measure can reduce measurement-related uncertainty. But the study may still be cross-sectional, confounded, selectively sampled, or too short to answer the central question.

Measurement improvement should therefore be considered as one part of whether the proposed study overcomes consequential weaknesses in previous evidence.

Do not allow a strong instrument to lend an aura of rigor to an otherwise poorly aligned design. The instrument can only solve the problem it actually addresses.

Better measurement can mean measuring the right outcome

Sometimes the problem is not the technical quality of the instrument but the choice of outcome itself.

Imagine that studies of an educational intervention consistently measure student satisfaction but the substantive claim concerns learning. Even an exceptionally reliable satisfaction scale cannot establish whether students learned more.

Likewise, a study of a workplace program might measure intentions when the question concerns actual behavior, or an intervention study might measure a short-term surrogate when the decision depends on a later outcome.

Measurement improvement can therefore involve moving from an indirect or peripheral outcome toward one that corresponds more closely to the research question.

Better measurement can also mean measuring change appropriately

An instrument suitable for comparing people at one time point is not automatically suitable for detecting within-person change.

COSMIN defines responsiveness as the ability of an instrument to detect change over time in the construct to be measured. This becomes particularly relevant in longitudinal and intervention studies where the central question concerns improvement, deterioration, or development.

If previous studies could not distinguish meaningful change from measurement noise, using an instrument with more appropriate evidence for responsiveness may materially improve the study.

A new instrument can create a comparability problem

Replacing an established measure has costs.

If most previous studies use one instrument and your study uses a completely different measure, direct comparison with the accumulated literature may become more difficult. A new instrument may also lack evidence across populations, languages, administration modes, or time.

This creates a trade-off. The new measure may better represent the construct while reducing comparability with earlier research.

Depending on the question, researchers might address this by administering both measures in a subset, examining convergence, using validated crosswalks where defensible, or explicitly accepting reduced comparability because the previous measure is too problematic to preserve.

There is no universal solution. The important point is to recognize that measurement innovation can improve one property of the evidence while weakening another.

Watch Out

Do not develop a new instrument merely because no existing instrument is perfect. Instrument development creates an additional research problem: the new measure itself requires evidence supporting its scores, properties, and intended interpretations. Sometimes selecting and appropriately adapting an existing measure is methodologically stronger than inventing another scale.

Better measurement should change what can be concluded

The final test is inferential.

Suppose previous studies use self-reported frequency of generative AI use, while your study obtains detailed platform logs. That sounds like an improvement. But what new conclusion does it support?

If the research question concerns actual interaction frequency, the logs may reduce recall problems and provide more granular behavioral evidence. If the question concerns students' motives for using AI, the logs alone may be worse because they do not directly capture motive.

The measure is better only relative to the quantity you need to know.

This is why collecting new data should be justified by the information the new evidence adds, not by the novelty of the instrument collecting it.

04 · A Practical Example

When replacing self-report with digital traces helps, and when it does not

Hypothetical Example

Measuring students' generative AI use more directly

A literature on generative AI and student performance relies heavily on questionnaires asking students how often they use AI tools. A researcher proposes collecting platform-use records instead and describes these as “objective measures of AI engagement.”

Step 1: Define the construct The researcher initially wants to measure frequency and duration of actual interaction with the AI system, not students' attitudes toward it.
Step 2: Examine the previous measure Retrospective self-report may be affected by recall and differences in how students interpret frequency categories.
Step 3: Examine the proposed measure Platform records can provide timestamps and interaction counts, but an open session does not necessarily represent active cognitive engagement and off-platform AI use remains unobserved.
Step 4: Refine the claim The researcher describes the logs as a more direct measure of recorded platform interaction, not as a perfect measure of “engagement.”
Step 5: Match measures to the research question If motives or perceived dependence are also important, the study retains appropriate self-report measures rather than expecting behavioral logs to measure subjective constructs they cannot observe.

The measurement strategy is stronger because the researcher has specified what each source measures and what it does not. “Objective data” was not the methodological achievement. Better alignment between construct, observation, and inference was.

05 · What Researchers Often Get Wrong

Common mistakes when claiming to add better measurement

Misconception

A validated instrument is valid for every study

Validity evidence supports particular interpretations and uses of scores. Check whether the instrument has appropriate evidence for the construct, population, language, context, administration, and interpretation relevant to your study rather than treating “validated” as a permanent property.

Misconception

High reliability proves that a measure is valid

No. Reliability concerns consistency and measurement error, while validity concerns whether the intended interpretation is adequately supported. A measure can be consistent without adequately representing the construct you intend to study.

Misconception

Objective measures are always better than self-report

Measurement method should follow the construct. Behavioral records may be preferable for some observable behaviors, while perceptions, experiences, intentions, and other subjective constructs may require participants' reports. Neither category is universally superior.

Misconception

A longer questionnaire measures a construct more accurately

More items can sometimes improve measurement, but length alone is not evidence of content validity, reliability, structural validity, or usefulness. Additional items can also increase burden and redundancy.

Misconception

A new instrument automatically makes the study more novel

Developing another instrument can be unnecessary when suitable measures already exist. The contribution should come from solving a meaningful measurement problem, not from adding another scale to a literature already rich in scales and poor in comparability.

Misconception

Better measurement fixes weaknesses elsewhere in the design

Improved measurement can strengthen the evidence it directly affects, but it does not automatically solve confounding, selection, inadequate comparisons, short follow-up, or other design problems.

06 · What This Means for You

Justify the measure by the inference it improves

When measurement is part of your study's contribution, describe the problem precisely. Avoid saying only that previous studies used “subjective” or “less robust” measures. Explain what those measures captured, what they failed to capture adequately, and why that limitation matters to the conclusion.

A simple decision framework

If existing measures do not adequately represent the intended construct
Prioritize stronger content alignment and evidence supporting the intended interpretation.
If measurement error prevents sufficiently precise inference
Choose or develop a strategy that reduces or appropriately characterizes the relevant error.
If the study concerns change over time
Determine whether the measure has appropriate evidence for detecting the kind of change you intend to interpret.
If an existing instrument already measures the construct adequately
Do not replace it merely for novelty; consider the value of comparability with previous research.
If no suitable measure exists
Recognize that developing or adapting one requires its own methodological justification and evidence before substantive conclusions can safely rest on it.

The strongest measurement justification can usually complete this sentence: previous studies could not adequately support conclusion X because Y was measured in a way that introduced problem Z; the proposed measurement strategy addresses Z and therefore provides more informative evidence about X.

07 · A Quick Checklist

Before claiming that your study measures something better, check the evidence

Before changing or introducing a measure, check:
Have you clearly defined the construct, behavior, exposure, or outcome you need to measure?
What exactly is inadequate about the measurement used in previous studies?
Does that weakness materially affect the conclusion you want to draw?
Is there evidence that the proposed measure adequately represents the intended construct and use?
Have you considered relevant reliability, measurement error, validity, or responsiveness evidence rather than relying on a generic “validated” label?
Is the evidence for the measure relevant to your population, language, context, and administration method?
Are you treating a proxy as a proxy rather than quietly equating it with the underlying construct?
Will changing measures reduce comparability with previous research, and is that trade-off justified?
Can you state what stronger conclusion becomes possible because measurement has improved?
08 · Frequently Asked Questions

Questions about improving measurement in a new study

What does it mean for an instrument to be validated?

It means that evidence has been gathered to support particular interpretations or uses of its scores or observations. It should not be interpreted as a universal certificate that the instrument is valid for every population, context, purpose, or administration method.

Is Cronbach's alpha enough to show that a questionnaire is good?

No. Internal consistency addresses only one aspect of measurement and its interpretation depends on the structure of the instrument. It does not by itself establish content validity, construct validity, reliability in other senses, measurement error, responsiveness, or suitability for your intended use. COSMIN treats these as distinct measurement properties.

Are behavioral or digital measures better than questionnaires?

Only when they provide better evidence for the construct you need. Logs may measure recorded behavior more directly, while questionnaires may be necessary for perceptions, intentions, experiences, or other subjective constructs. The appropriate choice depends on the research question.

Should I create a new instrument if existing ones have limitations?

Not automatically. First determine whether an existing instrument can adequately support your intended use, perhaps with justified adaptation. A new instrument introduces the additional task of establishing evidence for its measurement properties and intended interpretations.

Can better measurement justify collecting new data?

Yes, when measurement limitations are an important reason the existing evidence cannot support the required conclusion and the proposed measurement strategy meaningfully addresses that problem. Measurement novelty alone is not enough.

Can I use a measure validated in another country or language?

Potentially, but do not assume that evidence transfers automatically. Consider translation or adaptation, cultural relevance, measurement invariance where appropriate, and whether available evidence supports the intended interpretation in the new context. COSMIN treats cross-cultural validity or measurement invariance as a distinct measurement property.

Can better measurement reveal why previous studies disagree?

It can when measurement differences are a plausible source of disagreement. If studies operationalize the same nominal construct differently, a better-aligned measurement strategy may help determine whether the conflict reflects the phenomenon itself or how it was measured. The study should still be designed so that it can clarify disagreement rather than simply add another incompatible result.

09 · The Bottom Line

Better measurement means better alignment between the construct, observation, and conclusion

The Bottom Line

A proposed study adds better measurement when it captures the quantity or construct needed for the research question more appropriately and thereby improves the credibility, precision, or interpretation of the resulting evidence.

Do not equate better measurement with a newer scale, more items, digital data, or the word “validated.” Define what must be measured, identify what existing approaches fail to capture adequately, and show how the proposed measurement strategy changes what researchers can reasonably conclude.

10 · Sources and Further Reading

Sources and further reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes