03 · What You Need to Know
How to decide whether a new study genuinely measures something better
Begin with what you need to measure, not with the instrument
A surprisingly common measurement problem begins with instrument selection. Researchers find an established questionnaire, scale, platform metric, administrative variable, sensor, or test and then organize the construct around whatever that instrument happens to provide.
The stronger sequence runs in the opposite direction.
First define the construct, behavior, exposure, outcome, or attribute required by the research question. Then determine what observations would represent it adequately. Only after that should you decide which instrument or measurement procedure can generate those observations.
COSMIN guidance on outcome measurement similarly places the construct at the center of instrument selection and treats content validity as particularly important because researchers first need confidence that an instrument adequately reflects what it is intended to measure.
Measurement quality has several dimensions
Calling a measure “valid and reliable” can conceal more than it explains. Measurement quality is not one property.
COSMIN, for example, distinguishes properties including content validity, structural validity, internal consistency, cross-cultural validity or measurement invariance, reliability, measurement error, criterion validity, construct validity, and responsiveness. Which properties matter depends on the instrument and intended use.
| Measurement issue |
Core question |
Why it may matter |
| Content validity |
Does the measure adequately represent the construct it is intended to measure? |
A precise score is of limited value if important parts of the construct are absent or irrelevant content is included. |
| Structural validity |
Does the score structure reflect the dimensionality of the construct? |
Combining items into a score may be difficult to justify if the assumed structure is unsupported. |
| Reliability |
Can relevant differences among people or units be distinguished consistently from measurement error? |
Poor reliability can obscure real differences and attenuate relationships. |
| Measurement error |
How much observed variation arises from error rather than the quantity being measured? |
Error can reduce precision and complicate interpretation of differences or change. |
| Construct validity |
Do scores behave as expected given the construct they are intended to represent? |
Evidence is needed to support interpretations based on the measured construct. |
| Responsiveness |
Can the instrument detect relevant change over time? |
An instrument useful for distinguishing people at one time point may not adequately detect change. |
These concepts come from measurement frameworks with particular terminology and applications, and other disciplines may organize validity evidence differently. The broader lesson is portable: “better measurement” needs to specify what property is better and why that property matters to the research inference.
Validity belongs to an interpretation and use, not merely to the instrument's name
A familiar statement in research methods is that an instrument “has been validated.” That can be misleading if interpreted as permanent certification that the instrument is valid everywhere and for every purpose.
The Standards for Educational and Psychological Testing, jointly developed by the American Educational Research Association, American Psychological Association, and National Council on Measurement in Education, frame validity in terms of evidence and theory supporting interpretations of test scores for proposed uses. The Standards also emphasize that different interpretations or uses can require different validity evidence.
An instrument supported for screening one population may not automatically support a different diagnostic, evaluative, predictive, or research interpretation in another population. Translation, cultural adaptation, administration mode, scoring changes, and context can also affect what evidence is needed.
Thus, “we will use a validated questionnaire” is not the end of the measurement argument. Ask: validated for what interpretation, in whom, under what conditions, and how closely does that match the proposed use?
Reliability is not the same as validity
A measure can produce highly consistent scores while systematically failing to represent the intended construct adequately.
Imagine a bathroom scale that consistently reports the same person's weight several kilograms too high. Consistency alone does not make the measurement accurate. In psychological and educational measurement, the situation is usually more complex because many constructs do not have a simple physical reference standard, but the conceptual distinction remains useful.
COSMIN explicitly treats reliability and forms of validity as distinct measurement properties.
This matters when justifying new research. If previous evidence is limited because a construct was poorly represented, showing that the old measure had high internal consistency does not necessarily resolve the problem.
“Objective” measurement is not automatically superior to self-report
Researchers sometimes create a hierarchy in which device logs, administrative records, sensors, or behavioral traces are considered inherently superior to questionnaires or interviews.
The appropriate measurement depends on the construct.
If you want to know how many times a student opened an application, a reliable system log may be more direct than retrospective self-report. If you want to understand perceived usefulness, anxiety, confidence, motivation, or an experience that is inherently subjective, a behavioral log cannot simply replace participants' reports.
Even digital records require interpretation. Login counts may not represent meaningful use. Time with a browser tab open may not equal active engagement. Clicks can be measured with exquisite precision while remaining a questionable proxy for learning.
Objective-looking data can therefore be precisely measured proxies for something other than the construct you actually care about.
A proxy measure needs a defensible link to the target construct
Many important constructs cannot be observed directly. Researchers use indicators or proxies.
This is not inherently problematic. The problem arises when the proxy quietly becomes the construct.
Publication count is not identical to research quality. Platform activity is not identical to engagement. Course grades are not identical to learning. Citation counts are not identical to societal impact. Each may provide useful information for particular questions, but the inference depends on how well the indicator represents the target concept.
If previous studies rely heavily on a proxy whose relationship to the intended construct is weak or uncertain, a new measurement strategy may genuinely strengthen the evidence.
Better measurement can reduce one source of uncertainty without fixing the entire study
Suppose previous research uses a weak self-report measure of an exposure. Replacing it with a more defensible measure can reduce measurement-related uncertainty. But the study may still be cross-sectional, confounded, selectively sampled, or too short to answer the central question.
Measurement improvement should therefore be considered as one part of whether the proposed study overcomes consequential weaknesses in previous evidence.
Do not allow a strong instrument to lend an aura of rigor to an otherwise poorly aligned design. The instrument can only solve the problem it actually addresses.
Better measurement can mean measuring the right outcome
Sometimes the problem is not the technical quality of the instrument but the choice of outcome itself.
Imagine that studies of an educational intervention consistently measure student satisfaction but the substantive claim concerns learning. Even an exceptionally reliable satisfaction scale cannot establish whether students learned more.
Likewise, a study of a workplace program might measure intentions when the question concerns actual behavior, or an intervention study might measure a short-term surrogate when the decision depends on a later outcome.
Measurement improvement can therefore involve moving from an indirect or peripheral outcome toward one that corresponds more closely to the research question.
Better measurement can also mean measuring change appropriately
An instrument suitable for comparing people at one time point is not automatically suitable for detecting within-person change.
COSMIN defines responsiveness as the ability of an instrument to detect change over time in the construct to be measured. This becomes particularly relevant in longitudinal and intervention studies where the central question concerns improvement, deterioration, or development.
If previous studies could not distinguish meaningful change from measurement noise, using an instrument with more appropriate evidence for responsiveness may materially improve the study.
A new instrument can create a comparability problem
Replacing an established measure has costs.
If most previous studies use one instrument and your study uses a completely different measure, direct comparison with the accumulated literature may become more difficult. A new instrument may also lack evidence across populations, languages, administration modes, or time.
This creates a trade-off. The new measure may better represent the construct while reducing comparability with earlier research.
Depending on the question, researchers might address this by administering both measures in a subset, examining convergence, using validated crosswalks where defensible, or explicitly accepting reduced comparability because the previous measure is too problematic to preserve.
There is no universal solution. The important point is to recognize that measurement innovation can improve one property of the evidence while weakening another.
Watch Out
Do not develop a new instrument merely because no existing instrument is perfect. Instrument development creates an additional research problem: the new measure itself requires evidence supporting its scores, properties, and intended interpretations. Sometimes selecting and appropriately adapting an existing measure is methodologically stronger than inventing another scale.
Better measurement should change what can be concluded
The final test is inferential.
Suppose previous studies use self-reported frequency of generative AI use, while your study obtains detailed platform logs. That sounds like an improvement. But what new conclusion does it support?
If the research question concerns actual interaction frequency, the logs may reduce recall problems and provide more granular behavioral evidence. If the question concerns students' motives for using AI, the logs alone may be worse because they do not directly capture motive.
The measure is better only relative to the quantity you need to know.
This is why collecting new data should be justified by the information the new evidence adds, not by the novelty of the instrument collecting it.