03 · What You Need to Know
How Do You Assess Whether Research Data Are Truly Available?
Data feasibility begins with the research question. You first determine what evidence the question requires, then assess whether candidate data sources can provide it. Reversing that order can encourage researchers to ask whatever question an available spreadsheet happens to permit.
For studies using existing data, fit-for-purpose frameworks formalize this process by connecting a specific research question and study-design requirements to assessment of candidate data sources. One such framework, the Structured Process to Identify Fit-For-Purpose Data, explicitly considers both data relevance and operational access issues.
Write the Minimum Data Requirements Before Looking at the Dataset
List what the study must contain for the question to remain answerable. Depending on the design, this might include an exposure, outcome, comparison group, relevant covariates, timestamps, geographic identifiers, repeated observations, textual records, linkage variables, or sufficient cases meeting particular criteria.
Distinguish essential requirements from desirable ones. If the absence of a variable would make the central analysis uninterpretable, label it accordingly before you discover whether it happens to be available.
This reduces the temptation to redefine essential evidence as optional after seeing the dataset.
Verify the Data Dictionary, Not Just the Variable Names
A column labeled “engagement,” “performance,” “AI_use,” or “completion” tells you very little by itself.
How was the variable generated? What values are possible? Did its definition change? Does “completion” mean opening a module, viewing every page, passing an assessment, or receiving course credit? Does “AI use” indicate self-reported frequency, access to a tool, logged interactions, or merely enrollment in a course where AI was available?
Variable names can create an illusion of conceptual alignment. Examine documentation, coding rules, instruments, metadata, collection procedures, and operational definitions before concluding that the dataset measures what your question requires.
Check Whether the Necessary Population and Time Period Are Covered
A dataset may contain the right variables but the wrong people or years.
Perhaps electronic records became reliable only recently. A policy change altered how an outcome was recorded. The dataset includes undergraduate students but not postgraduate students. One campus contributes data while another does not. Earlier years contain only aggregated records.
Coverage determines what population and period your evidence can represent. Verify it rather than inferring it from the database title.
Granularity Can Determine Whether the Question Is Answerable
Aggregated data may answer questions about institutions or groups while being incapable of answering questions about individuals. Annual totals cannot necessarily establish event sequences. Course-level averages cannot reconstruct student-level relationships.
Similarly, data collected monthly may be too coarse for a process that unfolds over hours or days.
Ask whether the unit and temporal resolution of the available data match the unit and temporal logic of your question.
Data Linkage Should Never Be Assumed
Researchers may discover that the variables they need exist, but in separate systems.
One database contains demographic characteristics. Another contains academic performance. A third records platform use. The proposed study requires them together.
Can records actually be linked? Is there a stable identifier? Are identifiers available to researchers? Is linkage permitted under applicable governance and consent conditions? How accurate is the linkage? If identifiers were removed before data release, the theoretical possibility of linkage may no longer help you.
Confirm the complete data pathway, not merely the existence of each component.
Missing Data Can Change More Than the Sample Size
A variable being present in the schema does not mean it is populated adequately.
Suppose socioeconomic information appears in the dataset but is missing for half the participants. More importantly, suppose missingness is concentrated among particular groups. The resulting problem may affect both statistical precision and the credibility of the intended inference.
Before committing to the study, inspect expected completeness for essential variables where possible. If you cannot inspect the actual data yet, obtain documentation or credible estimates rather than assuming that listed fields are complete.
Data Quality Must Be Evaluated for the Intended Use
There is no single threshold at which a dataset becomes universally “high quality.” A source can be excellent for one research question and unsuitable for another.
Administrative data, for example, are collected primarily for administrative purposes. Platform logs record what the system was designed to log. Clinical records support care. None automatically captures every construct a researcher later wishes to investigate.
The relevant question is whether the data are fit for your intended purpose. The SPIFD framework similarly evaluates candidate sources in relation to specific study-design elements and the decision the resulting evidence is intended to inform.
Data Access Is Not the Same as Data Existence
A colleague telling you that a database exists is not permission to use it.
Access may require ethics approval, data-use agreements, institutional authorization, information-security review, consent restrictions, contractual negotiations, fees, or approval from a data custodian. Some sources can be analyzed only within secure environments. Others may prohibit exporting record-level data.
Operational access deserves explicit assessment because delays in contracting and data access can determine whether evidence can be produced within the required timeline.
Resolve critical permissions early, especially when the study has a fixed dissertation, grant, reporting, or policy deadline.
Publicly Available Does Not Mean Immediately Analysis-Ready
Open datasets can still require substantial preparation. Files may use unfamiliar formats, contain inconsistent coding across waves, require merging, lack documentation, or demand considerable computational resources.
There may also be different public and restricted versions. The public file may omit geographic detail, sensitive variables, exact dates, or linkage identifiers required by your question.
Download and inspect the actual accessible version whenever possible before finalizing the design.
Primary Data Collection Has Data Feasibility Problems Too
Researchers sometimes treat primary data collection as the solution to unavailable secondary data: if the variable does not exist, simply collect it.
That creates another set of feasibility questions. Can the phenomenon be measured adequately? Will participants provide the information accurately? Can observations occur at the necessary frequency? Is the instrument valid for the intended use? Can researchers collect the data consistently? Will enough complete observations be obtained?
Data feasibility applies whether the evidence already exists or must be generated.
Do Not Let the Dataset Rewrite the Question Without Admitting It
Suppose your question concerns students' critical evaluation of AI-generated claims. The available dataset contains only frequency of AI use and final grades. You might be tempted to substitute grades for critical evaluation because they are conveniently available.
That does not solve the data problem. It changes the question.
Changing the question can be entirely defensible. What matters is recognizing the change and evaluating whether the revised question still addresses a problem important enough to investigate.
Consider the Time Needed to Obtain and Prepare the Data
“Available next semester” may be functionally unavailable for a project due in three months. Likewise, six months of cleaning, linkage, transcription, coding, or digitization can make an otherwise excellent source impractical.
Assess not only acquisition time but the complete path from raw material to analyzable evidence. This should eventually be considered alongside whether the study is feasible within the time you actually have.
Watch Out
Do not write a proposal around a dataset you have never inspected when the study depends critically on its contents. A data dictionary, administrator's assurance, or database description can help, but none substitutes for verifying the actual variables, coverage, coding, completeness, access conditions, and usable form of the evidence whenever verification is possible.