Manuel B. Garcia

Manuel B. Garcia serves as the Senior Director for Educational Technology and Digital Learning at FEU Institute of Technology, Manila, Philippines. Read More

Contact Info

1607, FEU Tech Building,
P. Paredes St, Sampaloc,
Manila, Philippines
mbgarcia@feutech.edu.ph

Follow Me

How Much Can You Let Available Data Shape the Research Question Without Becoming Data-Driven Question Shopping?

Secondary research inevitably requires some negotiation between the question you want to ask and the data that actually exist. Learn how to adapt responsibly without searching the dataset for whichever question produces the most attractive result.

465
How Much Should Available Data Shape the Question? Guide 465 of 603
01 · The Question

When does adapting to available data become shopping for a question?

Secondary analysis rarely begins with unlimited freedom. The observations have already been collected. The variables, measurements, population, time period, and sample size are largely fixed. If your original question requires information the dataset does not contain, some adaptation may be unavoidable.

That is not automatically a methodological problem. A researcher might narrow a population, replace an impossible question with a related one the available measurements genuinely support, or recognize that only two of five planned survey waves contain the required variable.

But there is another kind of adaptation. You examine dozens of variables, outcomes, subgroups, transformations, and model specifications until something produces an appealing pattern. Then you write the paper as though that was the question you intended to investigate all along.

The distinction is not whether the data influenced the question. In secondary research, they often will. The important issue is whether the question is being adapted to the capabilities of the evidence or selected because of the result the evidence happens to produce.

02 · The Short Answer

Let data constrain feasibility, not secretly determine the preferred result

In Brief

Available data can legitimately shape the scope, operationalization, population, and feasibility of a research question, but the process becomes data-driven question shopping when observed results substantially determine which question, hypothesis, outcome, subgroup, or analysis is ultimately presented without that exploratory process being disclosed.

A useful safeguard is to separate information about what the dataset can support from information about what result it produces. When a question emerges from observed patterns, treat it transparently as exploratory or data-informed rather than reconstructing it as an independently specified hypothesis.

03 · What You Need to Know

Secondary analysis requires negotiation between scientific interest and data constraints

Researchers collecting new data can often design measurements around a question. Secondary analysts work in the opposite direction as well: they must determine whether an already existing set of measurements can address the question they care about.

This creates legitimate feedback between question development and data feasibility. Methodological guidance on secondary analysis acknowledges that many major features of an existing study are fixed. Researchers cannot retroactively alter the sample size, measurement instruments, collection schedule, or original design simply because a new question would benefit from different data.

The appropriate response is sometimes to refine the question. The difficulty is doing so without allowing observed results to become an undisclosed selection mechanism.

There is a difference between data-constrained and result-driven adaptation

Suppose you want to investigate whether generative AI use predicts students' writing performance. You discover from the codebook that the dataset contains no measure of writing performance but does contain a standardized measure of academic engagement.

You might decide that the original question is impossible with this dataset. If there is a genuine theoretical reason to investigate AI use and engagement, you could formulate that as a different question.

Now consider a different process. You calculate the associations between AI use and 25 outcomes, discover that engagement produces the strongest association, and then present an AI-engagement hypothesis as though engagement had been selected beforehand.

Both processes allow available data to influence the eventual question. Only the second uses the observed substantive results to select which question appears to have been intended.

Data-constrained question refinement The question changes because the dataset's variables, measurements, population, period, sample, or structure determine what can actually be studied.
Result-driven question shopping The question or hypothesis is selected substantially because an inspected association, effect, subgroup, outcome, or model produced an attractive result.

Let feasibility information narrow the question

Several kinds of information can reasonably influence the scope of a secondary-data question before the focal result is examined.

You may discover that the dataset:

  • contains a narrower measure than the construct you initially wanted;
  • covers only one part of the intended population;
  • contains the relevant variable in only certain waves;
  • has too few observations for a proposed subgroup comparison;
  • does not permit two required files to be linked;
  • contains substantial missingness in an essential variable; or
  • cannot support the temporal sequence implied by the original question.

Responding to these constraints can improve methodological alignment. For example, narrowing “university students nationally” to the population actually represented by the dataset can be more defensible than preserving a broader claim the data cannot support.

This is why limited preliminary inspection of existing data can sometimes be useful. The aim is to determine what is feasible without unnecessarily examining the focal relationships that could drive question selection.

Do not change the construct merely because a convenient variable exists

Adaptation has limits. Suppose your theoretical interest concerns academic self-efficacy, but the dataset measures general satisfaction with university life. Both are student attitudes, but they are not interchangeable.

Changing the label in the research question does not change what was measured. If the available variable addresses a different construct, ask whether that different construct supports a meaningful question in its own right.

This connects directly with evaluating whether available variables were measured in the way the question requires. A secondary-data study becomes more defensible when the question follows the measurement rather than asking the measurement to impersonate the preferred construct.

The number of analytical possibilities matters

Large secondary datasets can contain hundreds or thousands of variables, numerous subgroups, repeated waves, alternative operationalizations, and many defensible analytical specifications. This flexibility is scientifically useful, but it also creates many opportunities to choose analyses after seeing their results.

Methodological work on preexisting datasets identifies this flexibility as a transparency problem because researchers may make choices about outcomes, predictors, covariates, exclusions, transformations, or models after learning how those choices affect the result.

If only the successful analytical path appears in the final paper, readers cannot tell how many alternative questions or specifications were considered.

Question shopping is not defined by statistical significance alone

A researcher does not need to search explicitly for p <.05 for result-driven selection to occur.

You might choose the outcome with the largest effect, the subgroup with the clearest pattern, the model producing the most theoretically convenient coefficient, or the operationalization that makes the narrative most compelling. The general problem is that the observed result influences which analytical question receives privileged status.

Statistical significance is only one possible selection criterion.

Exploration is legitimate when it is presented as exploration

Exploring a rich dataset can be an excellent way to generate research questions. The difficulty begins when the history of that discovery is rewritten.

Research on preregistration of preexisting data emphasizes the distinction between exploratory, hypothesis-generating analysis and confirmatory hypothesis testing. If a pattern discovered during exploration motivates a new hypothesis, that hypothesis can be scientifically useful. It simply carries a different evidential status from a prediction specified without knowledge of that pattern.

One reasonable workflow is:

Explore Identify an interesting and theoretically meaningful pattern in existing data.
Explain the origin Report that the question or hypothesis emerged from exploration rather than presenting it as prespecified.
Develop the rationale Connect the observed pattern to relevant theory and prior evidence without pretending that literature search preceded the discovery if it did not.
Seek stronger confirmation when needed Test the hypothesis in independent data, a later wave, a held-out sample, or another suitable source when the scientific claim requires confirmation.

Exploration generates knowledge. Confirmation asks whether that knowledge survives a test not constructed around the same observed pattern.

HARKing is one form of the broader problem

The term HARKing refers to hypothesizing after the results are known while presenting the hypothesis as though it had been specified beforehand. In secondary-data research, the risk can be particularly difficult to recognize because the data already exist and researchers may have prior knowledge from earlier publications, previous analyses, collaborators, or preliminary inspection.

Methodological recommendations for secondary analysis therefore encourage researchers to document what was already known about the dataset and when key decisions were made.

Not every post hoc hypothesis is illegitimate. The misleading part is concealing its post hoc origin when that origin matters for how the evidence should be interpreted.

Preregistration can establish a decision boundary

Once feasibility is understood and the question is sufficiently developed, preregistration can record the hypotheses, operationalizations, exclusions, preprocessing decisions, and planned analyses before the focal tests are conducted.

Secondary-data preregistration templates specifically ask researchers to report prior data access and relevant knowledge of the dataset. This is important because preregistration does not require pretending that the data did not exist or that nothing was known about them.

Its value is partly in making the sequence visible: what was known, what was decided, and what remained to be tested.

Watch Out

A preregistration written after you have tried many outcomes or model specifications cannot retroactively turn the selected result into an independently generated prediction. It can still constrain future decisions, but relevant prior exploration should be disclosed.

Use theory and substantive importance to choose among feasible questions

Suppose the dataset supports five plausible questions. How should you choose?

Prefer criteria that do not depend on which result looks most attractive. Consider theoretical importance, unresolved evidence in the literature, practical relevance, quality of the available measurements, population fit, temporal appropriateness, statistical information, and whether the design can support the intended inference.

The result itself should not secretly become the principal selection criterion if you intend to present the eventual analysis as a test of a prespecified question.

Sometimes the right answer is that the dataset cannot answer the question

Researchers can become attached to a dataset because it is convenient, prestigious, expensive to obtain, or already cleaned. That creates pressure to find a question it can answer.

There is nothing wrong with discovering a different worthwhile question for an existing resource. But you should not progressively weaken the conceptual requirements of the original study merely to preserve the dataset.

If essential information is absent or the design cannot support the intended inference, the more defensible decision may be to reconsider the research idea when the available evidence is too weak rather than forcing an attractive dataset into service.

04 · A Practical Example

Two ways the same dataset can shape a research question

Hypothetical Example

Adapting to measurement versus selecting the best-looking outcome

A researcher is interested broadly in how generative AI use relates to university students' learning experiences. An existing dataset contains AI-use measures and several student outcomes.

Scenario A: Data-constrained refinement Before examining associations, the researcher reviews the questionnaire and discovers that the dataset does not measure learning performance but does contain a well-documented academic-engagement measure. Prior literature provides a substantive reason to investigate AI use and engagement, so the researcher reformulates the question accordingly.
Decision The question has been shaped by what the dataset can measure, but not by knowledge that engagement produces a desirable result.
Scenario B: Result-driven selection The researcher calculates associations between AI use and 18 available outcomes. Most are weak, but academic engagement produces the largest association.
Decision The researcher selects engagement as the outcome, develops a hypothesis around it, and reports the analysis as though that hypothesis preceded examination of the 18 outcomes.
Interpretation The second process conceals the role of the observed results in selecting the question. The analysis could still be reported as exploratory, but presenting it as an independently specified test would misrepresent how the hypothesis arose.

The distinction is methodological rather than moral. Researchers naturally notice patterns. The task is to preserve the history of how a question emerged so that readers can interpret the resulting evidence appropriately.

05 · What Researchers Often Get Wrong

Common mistakes when adapting questions to existing data

Misconception

The data should never influence the research question

That is often unrealistic in secondary analysis. Existing measurements, populations, periods, and sample structures legitimately constrain what can be studied. The important distinction is whether feasibility information or observed substantive results drive the adaptation.

Misconception

If the new question has a theoretical explanation, it was not data-driven

A plausible theory can often be constructed after a pattern is observed. The existence of a theoretical rationale does not change the chronology of discovery. Report whether the hypothesis preceded or followed examination of the result.

Misconception

Only searching for significant p-values counts as question shopping

Result-driven selection can involve effect size, direction, subgroup patterns, model fit, narrative appeal, or other desirable features. Statistical significance is only one way results can influence question selection.

Misconception

An exploratory question is weaker and should be made to look confirmatory

Exploratory research can be scientifically valuable. Concealing its exploratory origin does not strengthen the evidence; it makes the evidential history harder to evaluate.

Misconception

Preregistration removes any influence of earlier exploration

Preregistration can constrain decisions made after registration, but it cannot erase information already seen. Prior access and knowledge relevant to the hypothesis should be reported.

06 · What This Means for You

Use the dataset to define what is possible, then use scientific reasons to choose what is worth asking

A useful discipline is to separate two decisions. First ask what questions the dataset can credibly support. Then ask which of those questions is scientifically worth pursuing. The first decision depends heavily on the data's design and contents. The second should depend on substantive considerations rather than which answer looks best.

A simple decision framework

If the question changes because a required variable, population, period, or sample is unavailable
Refine the question to what the data genuinely support, provided the revised question remains substantively meaningful.
If several feasible questions remain before you have inspected their results
Choose among them using theory, prior evidence, measurement quality, relevance, and the inferential strengths of the design.
If the question emerged because you observed an interesting association or subgroup result
Report the analysis as exploratory or data-informed and consider testing the hypothesis with independent evidence.
If repeated adaptations are needed to make the dataset support some version of the original idea
Ask whether the research question and dataset have become too poorly aligned to justify continuing with that combination.

Before settling on the revised question, revisit the basic feasibility checks: whether the required variables actually exist, whether the population and period fit, whether enough relevant observations remain, and whether the measurements correspond to the constructs you intend to name.

The goal is not to prevent the data from informing your thinking. Good researchers learn from data. The goal is to prevent the final paper from presenting a cleaner chronology than the research process actually had.

07 · A Quick Checklist

Adapt the question without hiding how it was chosen

When available data begin shaping the research question, check:
Identify whether the proposed change is responding to a feasibility constraint or to an observed substantive result.
Confirm that any replacement variable genuinely represents the construct named in the revised question.
Choose among feasible questions using substantive and methodological criteria rather than whichever produces the most attractive association.
Keep a record of outcomes, predictors, subgroups, operationalizations, and analyses examined before the final question was selected.
Distinguish hypotheses specified before focal results were inspected from hypotheses generated after observing those results.
Use preregistration when useful to document unresolved analytical decisions before conducting the focal tests.
Present exploratory analyses as exploratory rather than reconstructing them as prespecified tests.
If the dataset requires repeated conceptual compromises, reconsider whether another dataset or research question would provide a better fit.
08 · Frequently Asked Questions

Questions about letting existing data shape the research question

Is it wrong to develop a research question after finding a dataset?

No. Existing datasets can legitimately motivate new research questions. What matters is how the question was generated, whether the dataset can validly address it, and whether exploratory discovery is represented transparently rather than described as an independently prespecified hypothesis.

Can I change my research question if an important variable is unavailable?

Yes, when the revised question is substantively meaningful and accurately reflects what the available variables measure. Changing the question is preferable to pretending an unsuitable variable measures something it does not.

What is question shopping?

In this context, question shopping refers to examining many possible outcomes, predictors, subgroups, or specifications and allowing the observed results to determine which research question receives emphasis, particularly when that selection process is not disclosed.

Can an exploratory finding become a research hypothesis?

Yes. Exploratory findings are a legitimate source of hypotheses. For stronger confirmation, the resulting hypothesis can subsequently be evaluated using independent data, a held-out sample, another wave, or another design appropriate to the claim.

Does finding supporting literature afterward make a data-generated hypothesis confirmatory?

No. Relevant literature can strengthen the theoretical interpretation of a finding, but it does not change when the hypothesis was generated. Preserve the chronology and report that the hypothesis emerged after examining the data when that is what occurred.

Can I preregister a question that was developed using an existing dataset?

Yes. Preregistration can still document the hypothesis and constrain subsequent analytical choices. Disclose relevant prior access, analyses, and knowledge of the data so readers can understand what the preregistration does and does not establish.

How do I know when I have adapted the question too far?

A warning sign is that the revised question requires you to relabel constructs, ignore important population or temporal mismatches, accept weak proxies, or make claims the design cannot support merely to keep using the dataset. At that point, another dataset or a different research idea may be more defensible.

09 · The Bottom Line

Let available data constrain the question without concealing how the question emerged

The Bottom Line

It is legitimate for existing data to shape a research question when their variables, measurements, population, time coverage, and sample determine what can realistically be studied; it becomes question shopping when observed results drive the selection of the preferred question and that selection process is hidden.

Use feasibility information to define what the dataset can support, substantive reasoning to decide what is worth asking, and transparent reporting to preserve the distinction between questions specified before focal results were known and questions generated through exploration.

10 · Sources and Further Reading

Sources and further reading on secondary-data question development

11 · Cite this Guide

How to Cite This Guide

This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.

Has the Field Guide helped your research?

If a guide helped clarify a question, inform a research decision, or move your work forward, I would love to hear about your experience. Your story may also help other researchers discover the Field Guide.

Share Your Experience
Takes only a few minutes