Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

How Do You Evaluate a Measure You Have Never Seen Before?

An unfamiliar measure should not be accepted or rejected merely because you do not recognize its name. Evaluate what it measures, how it was developed, what evidence supports it, and whether that evidence fits its current use.

372
How to Evaluate an Unfamiliar Measure Guide 372 of 899
01 · The Question

What Should You Do When You Do Not Recognize the Measure?

You are reading a paper and encounter an instrument you have never seen before. Perhaps it is a questionnaire with an impressive acronym, a researcher-developed scale, a performance task, an observational coding system, a device-derived measure, or an index assembled from several variables.

Not recognizing it tells you almost nothing about its quality. An unfamiliar instrument may be carefully developed and well supported. A familiar one may be poorly suited to the study in front of you.

The useful question is therefore not, “Have I heard of this measure?” It is: “What evidence would justify interpreting the measurements in the way these researchers do?” Answering that question requires a structured examination of the instrument rather than a reputation check.

02 · The Short Answer

Evaluate the Evidence, Not Your Familiarity With the Name

In Brief

When you encounter an unfamiliar measure, identify exactly what it is intended to measure, inspect how it was developed and scored, examine the relevant evidence for its measurement properties, and determine whether that evidence applies to the population, language, context, and use in the current study.

You do not need to know the instrument beforehand. What matters is whether you can trace a defensible path from the construct to the measurement procedure and from the resulting observations or scores to the interpretation the researchers make.

03 · What You Need to Know

Work From the Construct Outward

When an instrument is unfamiliar, it is tempting to start with statistics. You find a Cronbach's alpha, see that it is high, and feel somewhat reassured. That is often too early.

Before examining coefficients, determine what the instrument is supposed to measure and what its scores mean. Measurement evidence only makes sense in relation to a specified construct, population, context, and intended interpretation.

First, Identify Exactly What the Measure Is Supposed to Represent

Find the construct definition given by the study authors or the instrument developers. A measure labeled “digital readiness,” for example, might assess technical skills, attitudes toward technology, access to devices, confidence, willingness to adopt technology, or some combination of these.

Then inspect the actual content or measurement procedure whenever possible. What questions are asked? What behaviors are observed? What tasks must participants perform? What does the device record? How are observations converted into scores?

This lets you address the more fundamental question of whether the study appears to measure the construct it claims to measure.

Find the Original Instrument Source

If the paper cites the instrument's development or validation study, follow that citation rather than relying entirely on the current paper's summary. The original source may explain why the instrument was created, how items or tasks were generated, who participated in its development, what scores or subscales it produces, and what evidence was gathered to support their interpretation.

If the current paper gives an instrument name but no citation, insufficient description, or no route to determining what the measure contains, that is itself relevant to your appraisal. Readers need enough information to understand and evaluate how important variables were measured.

Examine Whether the Content Matches the Construct

Content validity concerns whether the content of a measurement instrument adequately represents the construct or outcome being measured. COSMIN's framework emphasizes three related considerations: relevance, comprehensiveness, and comprehensibility.

Relevance concerns whether the content belongs to the construct and is appropriate for the target population and context. Comprehensiveness concerns whether important aspects of the construct are missing. Comprehensibility concerns whether the people involved understand the questions, tasks, instructions, or other content as intended.

These questions are useful far beyond questionnaires. A performance task can omit important capabilities. An observational protocol can focus on behaviors that only partially represent a construct. A device can precisely record a physical signal that is only an imperfect proxy for the phenomenon researchers want to discuss.

Understand How the Score Is Produced

Do not treat the final number as self-explanatory. Determine how responses or observations become the score used in the analysis.

For a questionnaire, this may involve summing or averaging items, reverse-scoring particular responses, creating subscales, applying cutoffs, or transforming raw scores. For an observational measure, researchers may code behaviors according to specified rules. For a performance measure, the score may depend on accuracy, speed, task difficulty, or a scoring rubric.

Pay particular attention to what a higher or lower score means. If the researchers divide participants into categories, investigate how those thresholds were established and whether they are appropriate for the intended use.

Look for Several Kinds of Measurement Evidence

There is no single statistic that establishes the quality of an unfamiliar instrument. Which evidence matters depends on the type of measure and the interpretation being made.

Question to Ask Evidence That May Help What It Can Tell You
Does the content represent the intended construct? Instrument-development work, expert evaluation, target-population input, content-validity studies Whether important content is relevant, sufficiently represented, and understood as intended
Does the proposed structure make sense? Factor analysis or other evidence concerning internal structure Whether relationships among items or components are consistent with the proposed scoring structure
Are measurements sufficiently consistent? Internal consistency, test-retest reliability, inter-rater reliability, or other appropriate reliability analyses Whether particular sources of measurement variation are acceptably controlled for the intended use
Does the measure relate to other variables as theory predicts? Convergent, discriminant, known-groups, criterion-related, or hypothesis-testing evidence Whether observed relationships support the proposed interpretation
Does measurement behave appropriately across groups? Measurement invariance, differential item functioning, cross-cultural validity, or related analyses Whether scores may be comparable across relevant populations or groups

Not every instrument requires every analysis in this table. The relevant evidence depends on what is being measured and how the scores are being used. The goal is not to collect statistical badges. It is to determine whether the available evidence addresses the interpretation that matters in the current study.

Do Not Let Reliability Substitute for Validity

A reliability coefficient can answer an important question, but not the whole question. High internal consistency, for example, does not establish that the items cover the intended construct comprehensively or that the score means what researchers say it means.

COSMIN likewise treats content validity, structural validity, internal consistency, reliability, measurement error, criterion validity, hypothesis testing, and cross-cultural validity as distinct measurement properties or domains rather than interchangeable evidence.

Watch Out

If the only justification for an unfamiliar scale is a high Cronbach's alpha reported in the current sample, you still know relatively little about whether its scores support the intended construct interpretation. Consistent items can consistently measure something narrower or different from what researchers intended.

Check Who the Evidence Came From

Suppose an instrument has substantial supporting evidence, but all of it comes from middle-aged clinical patients in one country. The current study uses the instrument with secondary-school students elsewhere. That does not automatically invalidate the measure, but it creates a transfer question that deserves attention.

Ask whether the evidence is relevant to the participants in the current study. Age, culture, education, clinical characteristics, and other population differences can matter, particularly when they affect how people understand or respond to instrument content.

This is why the question of whether an instrument was evaluated in the same or a sufficiently relevant population deserves separate consideration.

Check the Language and Version

Instrument names can create a false impression of continuity. The scale used in the current study may not actually be identical to the version examined in earlier research.

Determine whether it was administered in the same language for which the supporting evidence was established. If translated, look for information about the adaptation process and evidence supporting use of the translated version.

Likewise, determine whether the researchers shortened the measure, removed questions, rewrote items, changed response options, altered instructions, combined subscales, or changed the recall period. A modified instrument may require additional evaluation rather than inheriting every property attributed to the original version.

Separate Evidence About the Instrument From Evidence About This Use

This distinction is especially important. You may find extensive literature supporting an instrument, but the relevant question is whether that evidence supports the particular interpretation being made here.

A measure can perform adequately for one purpose without being established for another. Screening, diagnosis, group comparison, individual decision-making, longitudinal change, and prediction can impose different measurement demands.

Do not ask simply, “Is this a good instrument?” Ask, “Is there adequate evidence for using these measurements this way, with these participants, for this conclusion?”

04 · A Practical Example

Evaluating an Unfamiliar “Online Learning Readiness” Scale

Hypothetical Example

You Encounter a Scale You Have Never Heard Of

Suppose a paper reports that students with greater “online learning readiness” achieve higher grades. Readiness is measured using a 20-item questionnaire unfamiliar to you. The authors report a high internal consistency coefficient and cite an earlier instrument-development paper.

1. Define the construct You check what “online learning readiness” means in this instrument. The developers define several dimensions, including self-directed learning, technology confidence, and online communication.
2. Inspect the content and scoring You determine which items represent each dimension and whether the study analyzes separate subscales or a total score.
3. Examine the supporting evidence You look beyond internal consistency and examine how the instrument was developed, whether its proposed structure was evaluated, and whether relationships with other variables behaved as expected.
4. Compare populations and versions You discover that the original evidence came from university students and that the current study also involves university students, but in another language. You therefore examine evidence for the translated version rather than assuming the original-language evidence transfers automatically.
5. Judge the current claim You assess whether the available evidence is sufficient to interpret the scores as online learning readiness in this population, rather than treating the reported reliability coefficient as the end of the appraisal.

Notice what you did not need: prior familiarity with the scale. The instrument's name initially told you very little. Your appraisal came from examining the construct, content, scoring, supporting evidence, and current application.

05 · What Researchers Often Get Wrong

Common Mistakes When Encountering an Unfamiliar Measure

Misconception

“I Have Never Heard of It, So It Must Be Questionable”

Familiarity is not a measurement property. Specialized disciplines contain many legitimate instruments that researchers outside those areas will not recognize. Evaluate the supporting evidence rather than the instrument's fame.

Misconception

“It Appeared in a Published Paper, So Someone Must Have Validated It”

Publication does not substitute for examining measurement evidence. A paper may use a researcher-developed measure, provide limited validation evidence, or cite earlier research that does not support the particular use being made.

Misconception

“The Alpha Is Above the Usual Threshold, So I Can Move On”

An internal consistency estimate cannot establish content validity, structural validity, population appropriateness, or the meaning of a score by itself. It answers a narrower measurement question and should be interpreted within the instrument's measurement model and intended use.

Misconception

“The Authors Cited a Validation Paper, So I Do Not Need to Read It”

The cited study may have examined another version, population, language, scoring method, or purpose. Following the citation is often the only way to determine what evidence actually exists.

Misconception

“A Long Validation Section Means the Measure Is Strong”

Quantity of analysis is not the same as relevant evidence. Several coefficients may still leave the central construct interpretation poorly supported. Ask what each analysis contributes to the argument for interpreting the measurements as intended.

06 · What This Means for You

You Can Appraise an Instrument Without Being an Expert on It

You do not need to reproduce an entire psychometric evaluation every time you read a paper. The depth of your appraisal should depend on how important the measure is to the study and how consequential the conclusions are.

A simple decision framework

If the unfamiliar measure represents the study's primary exposure or outcome
Examine its construct definition, development, scoring, relevant measurement properties, population, language, and version carefully.
If the measure is adequately described and supported by directly relevant evidence
Treat unfamiliarity itself as irrelevant. You have a substantive basis for judging the measurement.
If the evidence comes from a substantially different use or population
Treat transferability of the measurement interpretation as an open question rather than assuming the evidence applies unchanged.
If you cannot determine what the instrument contains, how it is scored, or what evidence supports it
Recognize that your confidence in conclusions depending on that measure is limited by insufficient measurement information.

The practical objective is not to classify every instrument as simply “valid” or “invalid.” It is to determine how strongly the available evidence supports the interpretation required by the study. That judgment can then inform how much weight you give the study's findings.

07 · A Quick Checklist

A Fast Appraisal of an Unfamiliar Measure

When you encounter an unfamiliar measure, check:
Identify the construct the instrument is intended to measure and how that construct is defined.
Inspect the instrument's items, tasks, observations, indicators, or measurement procedure when available.
Find the original development or validation source rather than relying only on the current paper's summary.
Determine how responses or observations are converted into the score used in the analysis.
Look for evidence relevant to content, structure, reliability, relationships with other variables, or other measurement properties required by the intended interpretation.
Compare the population and context supporting the measure with those in the current study.
Verify the language and version actually administered.
Check whether the researchers changed items, scoring, instructions, response options, or other important features.
Judge whether the evidence supports the specific interpretation and use required by the study.
08 · Frequently Asked Questions

Questions About Evaluating Unfamiliar Research Measures

Do I need to read the original validation paper for every instrument?

Not necessarily. The depth of checking should be proportional to the measure's importance to the study and the adequacy of the information already available. For a central exposure or outcome, consulting the original development and key validation evidence is often worthwhile.

Is a researcher-developed questionnaire automatically weaker than an established scale?

No. Its quality depends on how well the construct was defined, how the content was developed, how the instrument was evaluated, and whether the evidence supports its intended interpretation. An established instrument has the advantage of an existing evidence base, but its reputation alone does not make it appropriate for every study.

What if I cannot find the instrument's actual questions?

You may still be able to evaluate some aspects from the development paper, supplementary materials, instrument documentation, or cited sources. If the content and measurement procedure remain inaccessible, however, your ability to judge whether the measure adequately represents the construct is correspondingly limited.

What is the most important validity evidence to check first?

There is no universal first statistic. Begin by asking whether the instrument's content adequately represents the intended construct for the relevant population and context. COSMIN emphasizes relevance, comprehensiveness, and comprehensibility when evaluating content validity. Other evidence then addresses additional aspects of the intended interpretation.

Can I judge a measure from Cronbach's alpha alone?

No. Internal consistency addresses a narrower question about relationships among items and depends on assumptions about the scale. It does not establish by itself that the instrument covers the intended construct or that its scores support the study's interpretation.

Should an unfamiliar measure make me distrust the study?

Not by itself. Lack of familiarity is a property of the reader, not the instrument. Concern is more justified when the study provides insufficient information, relevant supporting evidence is weak, or the measure is used in ways that the available evidence does not support.

09 · The Bottom Line

An Unfamiliar Name Is a Starting Point, Not a Verdict

The Bottom Line

Evaluate an unfamiliar measure by reconstructing the argument behind it: what construct it represents, how it produces measurements, what evidence supports their interpretation, and whether that evidence applies to the population, language, version, context, and use in the study you are reading.

You do not need prior familiarity to make a useful appraisal. Once you replace “Have I heard of this?” with “What supports this interpretation?”, an unfamiliar instrument becomes a set of answerable methodological questions.

10 · Sources and Further Reading

Sources and Further Reading

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes