01 · The Question
What happens when an entire literature depends on one questionable measure?
You may encounter a surprisingly consistent pattern while reviewing a field: study after study measures the central construct with the same questionnaire, index, coding scheme, test, or proxy. Researchers cite earlier papers that used it, later researchers follow suit, and eventually the measure can begin to look established simply because it has become familiar.
Then you look for evidence supporting the interpretation of its scores and find very little.
This creates a deeper problem than one study choosing an imperfect instrument. If the measure does not adequately represent the construct researchers claim it represents, repeated use can propagate the same measurement uncertainty throughout the literature. Researchers may have replicated a numerical pattern without adequately establishing what those numbers mean.
03 · What You Need to Know
The problem is not that the measure is popular but that its interpretation is insufficiently supported
“Validated” is not a permanent label attached to an instrument
Researchers often speak of a “validated questionnaire” as though validity were a permanent property of the instrument. Contemporary measurement standards take a more careful position. Validity concerns the degree to which evidence and theory support particular interpretations of test scores for their intended uses. The relevant question is therefore not simply whether someone once validated the instrument, but whether sufficient evidence supports the way you intend to interpret its scores in your context.
That distinction matters when instruments migrate across populations, languages, disciplines, delivery modes, or purposes. Evidence supporting an interpretation among one population does not automatically establish that the same interpretation is appropriate everywhere else.
Likewise, reporting a reliability coefficient does not establish validity. Reliability concerns the consistency or precision of scores under specified conditions. A measure can produce highly consistent scores while systematically measuring something different from what researchers believe it measures.
Reliability evidence
Addresses aspects of score consistency or precision.
Validity evidence
Supports the interpretation and use researchers intend to make from the resulting scores.
Repeated use can create an illusion of validation
A measure can become conventional through citation rather than through accumulating appropriate validity evidence. One paper develops an instrument. Another adopts it. Subsequent studies cite those papers as precedent. Eventually authors may describe the measure as “widely used,” which is true, and then quietly treat widespread use as evidence that the measurement problem has been settled, which does not necessarily follow.
Usage establishes usage. It does not, by itself, establish content relevance, internal structure, appropriate response processes, relationships with other variables, measurement invariance, or the defensibility of a particular score interpretation.
This distinction is especially important when a measure was created informally for an early exploratory study and then inherited by an expanding literature without substantial psychometric investigation.
Measurement problems can propagate into substantive conclusions
Suppose researchers repeatedly study whether Construct A predicts Construct B. If Construct A is poorly operationalized, the problem is not confined to the methods section. Estimates involving A may be attenuated, distorted, unstable, or difficult to interpret. Researchers can then debate the relationship between A and B for years while the more basic question remains unresolved: are they measuring A adequately?
The consequences depend on the type of measurement problem. Poor content coverage may omit important parts of the construct. Items may inadvertently capture another construct. Translation may alter meaning. Respondents may interpret questions differently from what researchers intend. A scale assumed to be unidimensional may contain distinguishable dimensions. Scores may also function differently across groups.
These possibilities should not be assumed merely because an instrument lacks extensive validation. They are hypotheses about measurement that require investigation. The methodological gap is the absence of adequate evidence, not proof that the instrument is defective.
Validity evidence is broader than one statistical test
The Standards for Educational and Psychological Testing , developed jointly by the American Educational Research Association, American Psychological Association, and National Council on Measurement in Education, treats validation as an accumulation of evidence supporting intended score interpretations. Relevant evidence can concern test content, response processes, internal structure, relationships with other variables, and the consequences associated with testing.
Which evidence matters most depends on what the instrument is intended to measure and how its scores will be used. Factor analysis alone, for example, cannot answer every validity question. Nor does a high Cronbach's alpha transform an inadequately conceptualized collection of items into a valid measure.
Using the same operationalization can make a literature methodologically dependent
If dozens of studies use the same instrument, their results may be independent in terms of samples or research teams while remaining dependent on one operational definition of the construct.
Imagine that every study of “digital research competence” uses the same six-item questionnaire. Thirty studies finding similar relationships with academic performance would show that scores from that questionnaire behave somewhat consistently across those studies. But if the questionnaire primarily captures confidence rather than demonstrated competence, the literature may have accumulated evidence about perceived competence while discussing it as actual competence.
That is why a methodological weakness shared across many studies can matter more than the limitation of any single paper. The same assumption travels from study to study.
An alternative measure can test whether the finding depends on operationalization
One response is to use another measure with stronger evidence for the intended interpretation. An even more informative design may measure the same construct in more than one defensible way.
If a substantive relationship appears across meaningfully different operationalizations, confidence that it is not merely an artifact of one instrument may increase. If the relationship appears with one measure but disappears with another, the disagreement becomes scientifically useful. Researchers then have reason to investigate what the measures actually capture.
Watch Out
Do not label an instrument “unvalidated” merely because you did not immediately find a validation paper. Search the original development study, later psychometric studies, adaptations, translations, manuals, and evidence for the specific population and interpretation before making the claim.
Sometimes measurement validation should come before another substantive study
If the central variable in a literature cannot be measured with sufficient confidence, adding another correlational, experimental, or predictive study may be premature. The more consequential contribution may be to clarify the construct and establish stronger evidence for its measurement.
This can involve developing a new instrument, evaluating an existing one, comparing competing measures, testing dimensionality, examining measurement invariance, studying response processes, or combining self-report with behavioral or observational indicators where appropriate.
The objective is not psychometric sophistication for its own sake. It is to make subsequent substantive conclusions interpretable.
04 · A Practical Example
When the measurement problem becomes more important than another correlation
Hypothetical Example
Twenty studies, one questionnaire, little validity evidence
Suppose you review research on university students' “AI literacy.” You find 20 studies, and 17 use essentially the same eight-item self-report questionnaire. The instrument asks students to rate statements such as whether they understand AI and can use AI tools effectively. The studies repeatedly report relationships between questionnaire scores and other educational variables.
Observation
The instrument dominates the literature and produces apparently coherent findings.
Verification
You trace the measure to its original paper and search subsequent psychometric work rather than assuming that frequent citation establishes validity.
Problem
You find limited evidence that questionnaire scores correspond to demonstrated AI knowledge or performance, and little investigation of whether respondents interpret the items as intended.
Uncertainty
Existing studies may support conclusions about perceived AI literacy more strongly than demonstrated AI literacy.
New contribution
A study compares the questionnaire with a carefully developed performance-based assessment and investigates the relationships among both measures and relevant outcomes.
If self-reported and performance-based scores converge appropriately, the result adds useful validity evidence. If they diverge substantially, that result may be even more consequential because it changes how earlier findings should be interpreted. Either outcome addresses an uncertainty that another study using the same questionnaire would have left untouched.
06 · What This Means for You
Decide whether your field needs another application or better measurement evidence
A simple decision framework
If strong validity evidence exists for your intended interpretation and population
Do not manufacture a measurement gap simply because many studies use the same instrument.
If validity evidence exists but not for your population, language, or intended use
Determine which aspects require additional evidence before treating scores as equivalent.
If the field repeatedly uses the measure with little relevant validity evidence
Consider whether evaluating the instrument or using an independently supported operationalization would reduce a more fundamental uncertainty.
If competing measures operationalize the construct differently
Consider comparing them rather than assuming that one operationalization defines the construct.
Your literature review should therefore distinguish between evidence about the phenomenon and evidence about the instrument used to observe it. A large body of substantive findings cannot fully compensate for uncertainty in the measurement on which those findings depend.
If nearly every paper uses the same measure, ask a simple but uncomfortable question: what do we know about the construct that does not depend on this instrument? If the answer is “very little,” you may have identified a consequential methodological gap.
07 · A Quick Checklist
Before claiming that a literature relies on an unvalidated measure, verify the evidence
Before treating measurement as the research gap, check:
Locate the original instrument-development publication or technical documentation.
Search for later validation, psychometric, adaptation, translation, and measurement-invariance studies rather than relying only on the original paper.
Identify exactly what construct the instrument was designed to measure and compare that with how later researchers describe it.
Check whether available validity evidence applies to the population, language, setting, and score interpretation relevant to your study.
Do not treat reliability coefficients alone as evidence that the intended interpretation is valid.
Look for studies using alternative operationalizations of the same construct and compare whether their findings converge.
Determine whether measurement uncertainty could materially change the substantive conclusions of the literature.
If proposing a new measure, plan to generate appropriate validity evidence rather than merely replacing the old instrument.
11 · Cite this Guide
How to Cite This Guide
This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.
Recommended (Field Guide)
APA
MLA
Chicago
Copy Citation