01 · The Question
When does adapting to available data become shopping for a question?
Secondary analysis rarely begins with unlimited freedom. The observations have already been collected. The variables, measurements, population, time period, and sample size are largely fixed. If your original question requires information the dataset does not contain, some adaptation may be unavoidable.
That is not automatically a methodological problem. A researcher might narrow a population, replace an impossible question with a related one the available measurements genuinely support, or recognize that only two of five planned survey waves contain the required variable.
But there is another kind of adaptation. You examine dozens of variables, outcomes, subgroups, transformations, and model specifications until something produces an appealing pattern. Then you write the paper as though that was the question you intended to investigate all along.
The distinction is not whether the data influenced the question. In secondary research, they often will. The important issue is whether the question is being adapted to the capabilities of the evidence or selected because of the result the evidence happens to produce.
03 · What You Need to Know
Secondary analysis requires negotiation between scientific interest and data constraints
Researchers collecting new data can often design measurements around a question. Secondary analysts work in the opposite direction as well: they must determine whether an already existing set of measurements can address the question they care about.
This creates legitimate feedback between question development and data feasibility. Methodological guidance on secondary analysis acknowledges that many major features of an existing study are fixed. Researchers cannot retroactively alter the sample size, measurement instruments, collection schedule, or original design simply because a new question would benefit from different data.
The appropriate response is sometimes to refine the question. The difficulty is doing so without allowing observed results to become an undisclosed selection mechanism.
There is a difference between data-constrained and result-driven adaptation
Suppose you want to investigate whether generative AI use predicts students' writing performance. You discover from the codebook that the dataset contains no measure of writing performance but does contain a standardized measure of academic engagement.
You might decide that the original question is impossible with this dataset. If there is a genuine theoretical reason to investigate AI use and engagement, you could formulate that as a different question.
Now consider a different process. You calculate the associations between AI use and 25 outcomes, discover that engagement produces the strongest association, and then present an AI-engagement hypothesis as though engagement had been selected beforehand.
Both processes allow available data to influence the eventual question. Only the second uses the observed substantive results to select which question appears to have been intended.
Data-constrained question refinement
The question changes because the dataset's variables, measurements, population, period, sample, or structure determine what can actually be studied.
Result-driven question shopping
The question or hypothesis is selected substantially because an inspected association, effect, subgroup, outcome, or model produced an attractive result.
Let feasibility information narrow the question
Several kinds of information can reasonably influence the scope of a secondary-data question before the focal result is examined.
You may discover that the dataset:
- contains a narrower measure than the construct you initially wanted;
- covers only one part of the intended population;
- contains the relevant variable in only certain waves;
- has too few observations for a proposed subgroup comparison;
- does not permit two required files to be linked;
- contains substantial missingness in an essential variable; or
- cannot support the temporal sequence implied by the original question.
Responding to these constraints can improve methodological alignment. For example, narrowing “university students nationally” to the population actually represented by the dataset can be more defensible than preserving a broader claim the data cannot support.
This is why limited preliminary inspection of existing data can sometimes be useful. The aim is to determine what is feasible without unnecessarily examining the focal relationships that could drive question selection.
Do not change the construct merely because a convenient variable exists
Adaptation has limits. Suppose your theoretical interest concerns academic self-efficacy, but the dataset measures general satisfaction with university life. Both are student attitudes, but they are not interchangeable.
Changing the label in the research question does not change what was measured. If the available variable addresses a different construct, ask whether that different construct supports a meaningful question in its own right.
This connects directly with evaluating whether available variables were measured in the way the question requires. A secondary-data study becomes more defensible when the question follows the measurement rather than asking the measurement to impersonate the preferred construct.
The number of analytical possibilities matters
Large secondary datasets can contain hundreds or thousands of variables, numerous subgroups, repeated waves, alternative operationalizations, and many defensible analytical specifications. This flexibility is scientifically useful, but it also creates many opportunities to choose analyses after seeing their results.
Methodological work on preexisting datasets identifies this flexibility as a transparency problem because researchers may make choices about outcomes, predictors, covariates, exclusions, transformations, or models after learning how those choices affect the result.
If only the successful analytical path appears in the final paper, readers cannot tell how many alternative questions or specifications were considered.
Question shopping is not defined by statistical significance alone
A researcher does not need to search explicitly for p <.05 for result-driven selection to occur.
You might choose the outcome with the largest effect, the subgroup with the clearest pattern, the model producing the most theoretically convenient coefficient, or the operationalization that makes the narrative most compelling. The general problem is that the observed result influences which analytical question receives privileged status.
Statistical significance is only one possible selection criterion.
Exploration is legitimate when it is presented as exploration
Exploring a rich dataset can be an excellent way to generate research questions. The difficulty begins when the history of that discovery is rewritten.
Research on preregistration of preexisting data emphasizes the distinction between exploratory, hypothesis-generating analysis and confirmatory hypothesis testing. If a pattern discovered during exploration motivates a new hypothesis, that hypothesis can be scientifically useful. It simply carries a different evidential status from a prediction specified without knowledge of that pattern.
One reasonable workflow is:
Explore Identify an interesting and theoretically meaningful pattern in existing data.
Explain the origin Report that the question or hypothesis emerged from exploration rather than presenting it as prespecified.
Develop the rationale Connect the observed pattern to relevant theory and prior evidence without pretending that literature search preceded the discovery if it did not.
Seek stronger confirmation when needed Test the hypothesis in independent data, a later wave, a held-out sample, or another suitable source when the scientific claim requires confirmation.
Exploration generates knowledge. Confirmation asks whether that knowledge survives a test not constructed around the same observed pattern.
HARKing is one form of the broader problem
The term HARKing refers to hypothesizing after the results are known while presenting the hypothesis as though it had been specified beforehand. In secondary-data research, the risk can be particularly difficult to recognize because the data already exist and researchers may have prior knowledge from earlier publications, previous analyses, collaborators, or preliminary inspection.
Methodological recommendations for secondary analysis therefore encourage researchers to document what was already known about the dataset and when key decisions were made.
Not every post hoc hypothesis is illegitimate. The misleading part is concealing its post hoc origin when that origin matters for how the evidence should be interpreted.
Preregistration can establish a decision boundary
Once feasibility is understood and the question is sufficiently developed, preregistration can record the hypotheses, operationalizations, exclusions, preprocessing decisions, and planned analyses before the focal tests are conducted.
Secondary-data preregistration templates specifically ask researchers to report prior data access and relevant knowledge of the dataset. This is important because preregistration does not require pretending that the data did not exist or that nothing was known about them.
Its value is partly in making the sequence visible: what was known, what was decided, and what remained to be tested.
Watch Out
A preregistration written after you have tried many outcomes or model specifications cannot retroactively turn the selected result into an independently generated prediction. It can still constrain future decisions, but relevant prior exploration should be disclosed.
Use theory and substantive importance to choose among feasible questions
Suppose the dataset supports five plausible questions. How should you choose?
Prefer criteria that do not depend on which result looks most attractive. Consider theoretical importance, unresolved evidence in the literature, practical relevance, quality of the available measurements, population fit, temporal appropriateness, statistical information, and whether the design can support the intended inference.
The result itself should not secretly become the principal selection criterion if you intend to present the eventual analysis as a test of a prespecified question.
Sometimes the right answer is that the dataset cannot answer the question
Researchers can become attached to a dataset because it is convenient, prestigious, expensive to obtain, or already cleaned. That creates pressure to find a question it can answer.
There is nothing wrong with discovering a different worthwhile question for an existing resource. But you should not progressively weaken the conceptual requirements of the original study merely to preserve the dataset.
If essential information is absent or the design cannot support the intended inference, the more defensible decision may be to reconsider the research idea when the available evidence is too weak rather than forcing an attractive dataset into service.