01 · The Question
When researchers study online data, whom are they actually studying?
A researcher downloads thousands of public posts and analyzes patterns in the text. Another researcher joins an online community, interviews members, and asks them to complete a survey. A third constructs detailed profiles by linking pseudonymous posts across several platforms.
All three projects involve information created by internet users. Are those users research participants?
The answer matters because formal classification can affect ethics review, informed consent, data handling, and other regulatory requirements. Yet “participant” and “data source” are sometimes used too casually, as though researchers can choose whichever label makes a project administratively easier.
The more defensible approach is to examine what the researcher actually does with and to the people behind the data.
03 · What You Need to Know
“Participant” is not simply another word for “person whose data appear in my dataset”
Formal definitions depend on the framework governing the research
Different jurisdictions and institutions define human research participation through their own regulations and policies. Researchers should therefore avoid assuming that one definition applies universally.
The U.S. Common Rule provides a useful example. Under the current HHS definition, a human subject is a living individual about whom a researcher obtains information or biospecimens through intervention or interaction and uses, studies, or analyzes them, or about whom the researcher obtains, uses, studies, analyzes, or generates identifiable private information or identifiable biospecimens.
This definition immediately shows why internet research cannot be classified merely by saying, “The data came from people.” The relevant questions include how the information was obtained, whether the researcher interacted with anyone, whether the information is private, and whether it is identifiable.
Direct interaction makes the relationship easier to recognize
If researchers recruit people through social media and then interview them, send them questionnaires, communicate with them for research purposes, or manipulate their online environment as part of an experiment, the research relationship is relatively clear.
Under the U.S. Common Rule, interaction includes communication or interpersonal contact between investigator and subject, while intervention includes physical procedures and manipulations of the subject or their environment performed for research purposes.
An online survey does not stop involving people because the questionnaire appears in a browser. A virtual interview remains an interaction. Internet mediation changes the channel, not necessarily the regulatory relationship.
Existing online material creates a more difficult classification problem
Now consider a researcher who never contacts users and instead analyzes posts that already exist.
Under the Common Rule definition, the absence of interaction does not automatically mean the project falls outside human-subjects research. Research can also involve human subjects when researchers use identifiable private information about living individuals. Conversely, research using information that does not meet the relevant criteria may fall outside that regulatory definition.
That distinction helps explain why the question of whether researchers may use websites, forums, and social media posts without consent cannot be answered merely by calling the material “secondary data.”
Public information can be treated differently from private information
Publicity can materially affect regulatory classification.
U.S. advisory guidance on internet research has treated information legally available to internet users without specific authorization as public rather than private for purposes of the Common Rule analysis, while recognizing that law, privacy policies, access mechanisms, and terms governing the information can affect that determination. The same guidance notes that internet data can create distinctive identification problems because information can be mined, matched, and combined across datasets.
UKRI guidance similarly states that information intentionally made public in internet spaces may be considered in the public domain, while cautioning researchers to examine critically the public nature of online communications, privacy, consent, and identifiability.
Researchers therefore need to distinguish the formal status of public social media material as research data from the broader ethical choices involved in using it.
A data source can still represent a real person
There is a danger in allowing regulatory terminology to become ethical shorthand.
Suppose an institutional determination concludes that a study of unrestricted public posts does not meet the applicable definition of human-subjects research. That conclusion may settle an important regulatory question. It does not make the people behind the posts imaginary.
Researchers may still expose identities, amplify sensitive disclosures, misrepresent communities, reproduce searchable quotations, or create profiles that affect people who never expected their scattered online activity to be aggregated.
Regulatory classification
Determines whether a project falls within a particular formal definition of research involving human subjects or participants.
Ethical responsibility
Concerns the foreseeable effects of collection, analysis, linkage, interpretation, storage, and dissemination on the people and communities represented in the data.
The first can inform the second. It should not replace it.
The person providing information is not always the person the research is about
Another complication arises when one internet user provides information about someone else.
HHS guidance emphasizes the phrase “about whom” in the Common Rule definition. A person supplying information is not necessarily the human subject if the information being collected is actually about another individual.
Imagine a public post in which a parent describes a child's medical experience, an employee describes a colleague, or a forum member discusses a spouse. Researchers may be collecting content authored by one person while extracting information about another.
This distinction becomes especially important when researchers build person-level datasets from posts that mention third parties.
Pseudonymous users can still be identifiable
An account called “ResearchFan247” is not necessarily anonymous in an ethical or regulatory sense.
Identity may be recoverable from profile information, photographs, locations, links, repeated usernames, distinctive quotations, or combinations of attributes. HHS guidance treats identifiability contextually rather than relying on a simple checklist of direct identifiers.
This matters because an online quotation can lead readers back to its author even when the research report omits the account name.
Public data can become person-level profiling
A project may begin by collecting public posts but then link them by username, infer demographics, map social networks, estimate location, classify psychological attributes, or combine information from several platforms.
At that point, describing the material simply as “posts” can obscure what the study actually does. The unit of analysis may effectively have shifted from documents to people.
This does not automatically determine regulatory status, because formal definitions still apply. It does mean that privacy, identification, proportionality, and potential harm deserve closer examination.
Large-scale data complicate the participant-data distinction
At very large scales, researchers may never know the names of the people represented in a dataset. Yet automated analysis can still generate detailed inferences about groups and individuals.
The Association of Internet Researchers recommends examining ethical issues across the lifecycle of internet research, including collection, analysis, storage, existing datasets, identifiability, inference, and derivative uses. Its guidance encourages researchers to ask whether data collection is necessary and proportionate and whether apparently non-identifying information can be combined to reconstruct sensitive characteristics.
The ethical implications of large-scale data collection therefore cannot be settled by announcing that the dataset contains “data points rather than participants.”
Do not self-classify strategically
Researchers should not choose terminology based on which category appears to require less paperwork.
Watch Out
Calling people “data sources,” “users,” “accounts,” or “records” does not determine whether a study falls within a formal definition of human-subjects research. Describe what the study actually does and allow the applicable regulatory and institutional criteria to determine its status.
Where classification is uncertain, researchers should seek the appropriate institutional determination rather than assuming that secondary or publicly accessible online data automatically fall outside ethics oversight.
04 · A Practical Example
The same platform can support very different research relationships
Hypothetical Example
Three studies using the same social media platform
Three research teams investigate public discussion of remote work using the same platform.
Study A: Existing public posts
Researchers analyze aggregate linguistic patterns in unrestricted posts without interacting with users or constructing person-level profiles.
Study B: Online recruitment
Researchers message users, recruit them into interviews, obtain consent, and ask about their work experiences.
Study C: Individual profiling
Researchers link users' posts across several online services and infer occupations, locations, employers, relationships, and behavioral patterns.
Classification
The teams apply the definitions and requirements of the ethics framework governing each study rather than assuming that all three projects have the same status because they involve the same platform.
Ethical analysis
Regardless of formal classification, each team examines identification, privacy, data minimization, security, reporting, and potential harm in light of what its particular method does with information about people.
The website is not what determines the research relationship. The protocol does.
07 · A Quick Checklist
Before classifying internet users as participants or data sources, check the protocol
Before making the classification, check:
Identify the exact regulatory or institutional definition of a human subject or research participant that governs the study.
Determine whether researchers communicate or interact with users specifically for research purposes.
Determine whether researchers manipulate users or their online environment as part of the study.
Establish whether the information is public, private, restricted, or subject to meaningful access conditions.
Assess whether the people represented in the information can be identified directly or through linkage, inference, quotation, or contextual details.
Check whether information authored by one user actually concerns another identifiable person.
Reassess identifiability if datasets will be linked or person-level characteristics will be inferred.
Separate the formal ethics-review classification from broader responsibilities concerning privacy, harm, security, and reporting.
Obtain an institutional determination when the study's human-subjects status is uncertain.