01 · The Question
If nobody says the participant's name, is the recording anonymous?
You record a research interview. The participant never says their name. You label the file with a participant code rather than a name, and perhaps even remove identifying details from the transcript.
Is the original audio now anonymous?
Not necessarily. Identifiability is not determined solely by whether a name appears in the file. A participant may be recognizable from their voice, what they say, or information that can be connected with other available knowledge. That makes audio different from a dataset in which a direct identifier can simply be deleted from a column.
03 · What You Need to Know
Why an unnamed voice recording may still reveal identity
Identifiability does not require knowing someone's name
A common mistake is to equate identification with naming. In research ethics and data protection, the relevant question is generally broader: can a particular person be distinguished, recognized, or connected to the information?
Under the U.S. Common Rule, for example, identifiable private information is information for which a participant's identity is or may readily be ascertained by the investigator or associated with the information. The U.S. Office for Human Research Protections also emphasizes that determining whether information is identifiable is contextual rather than dependent on a fixed list of identifiers.
Similarly, UK data protection guidance explains that a person can be identifiable directly or indirectly and that you do not necessarily need to know someone's name to distinguish them from other people.
Unnamed
The person's name does not appear in the recording.
Unidentified
The researcher or listener does not currently know who the person is.
Anonymous
The information does not relate to a person who remains identifiable under the applicable standard and context.
Those conditions are not equivalent. A participant can be unnamed and still readily recognizable.
The participant's voice may itself make them recognizable
Think about hearing a close colleague, friend, supervisor, or family member speaking from another room. You often do not need an introduction to know who is talking.
The same problem can arise in research recordings. A listener who already knows the participant may recognize their voice. Depending on the circumstances, speech may also contain characteristics such as accent, dialect, habitual expressions, speech patterns, or other features that contribute to recognition.
Recognition risk is therefore audience-dependent. A recording might reveal very little to a member of the general public while being immediately recognizable to someone from the participant's workplace, family, community, or professional network.
An ordinary voice recording is not automatically biometric data
There is an important technical distinction here. The fact that a voice can contribute to identification does not mean that every audio recording should automatically be described as biometric data.
Under UK GDPR terminology, biometric data result from specific technical processing of physical, physiological, or behavioural characteristics that allows or confirms unique identification. The Information Commissioner's Office gives voice as a characteristic that can become biometric data when it is processed using specific technologies for this purpose.
An ordinary research interview recording may therefore be identifiable personal information without necessarily being biometric data under that definition. If researchers use voice-recognition or speaker-identification technology, however, additional biometric-data considerations may arise under the applicable law.
What the participant says may identify them even if their voice does not
Suppose the voice has been altered so that recognition by sound is unlikely. The recording can still reveal identity through its content.
A participant might mention their employer, job title, neighborhood, educational institution, research specialty, unusual career history, family relationship, distinctive event, or another person. Some clues may appear harmless individually but become identifying when combined.
This is sometimes called the jigsaw or mosaic effect. Current Information Commissioner's Office guidance describes linkability as the possibility of combining information within or across sources in ways that allow a person to be singled out or identified. Qualitative data are particularly susceptible because contextual richness is often precisely what makes the material analytically valuable.
The same issue can survive transcription. Removing a speaker's name does not necessarily make an interview transcript anonymous if its narrative details still reveal who the speaker is. The question of whether an apparently anonymous quotation can identify a participant therefore involves much the same contextual reasoning.
Who has access changes the identification risk
Identifiability should not be assessed as though every possible listener knows the same things.
Imagine an interview with the only pediatric neurosurgeon in a small region. A stranger may have no idea who is speaking. A colleague who hears the recording may need only the participant's role and a few contextual details to identify them. In another study, a distinctive voice alone may be enough for coworkers to recognize the speaker.
Current anonymisation guidance from the Information Commissioner's Office similarly treats identifiability as contextual and asks whether means reasonably likely to be used by the relevant parties could identify someone. This includes considering linkability with other information.
Researchers should therefore ask who will have access to the recording and what those people could realistically know, not merely whether an unknown member of the public could identify the speaker.
Replacing the filename with a participant code does not anonymise the audio
Changing Maria_Santos_Interview.mp3 to P017.mp3 is sensible data management, but it addresses only the filename. It does not alter the contents of the recording.
If a separate key connects P017 with the participant, the coded system may also preserve a route back to identity. OHRP guidance specifically recognizes that coded information may remain linkable to individuals through coding systems, depending on who can access the key and the circumstances of the research.
This distinction is why identifiable recordings should be managed according to their actual disclosure risk , rather than assuming that a pseudonym or participant number has made the original material anonymous.
A transcript and the original audio may have different identification risks
Audio and transcripts should not automatically be assigned the same confidentiality status. A carefully de-identified transcript can remove or generalize names, organizations, locations, and other contextual identifiers. The original audio preserves the participant's voice and everything that was said.
This does not mean a transcript is automatically anonymous either. UK Data Service guidance on qualitative text warns that distinctive life events, occupations, geographic references, and combinations of contextual details can permit identification after names have been removed.
The distinction matters when deciding who needs access to the original audio, what version can be shared with collaborators, and whether the identifiable recording still needs to be retained after transcription .
04 · A Practical Example
An interview with no name can still point to one person
Hypothetical Example
An unnamed teacher describes a distinctive experience
A researcher interviews teachers about adopting educational technology. Participant P12 never states her name. During the recording, however, she mentions that she teaches chemistry at the only senior high school in a small municipality, became department chair last year, and led a highly publicized robotics competition. Her natural voice remains in the audio.
Name
No name appears in the recording or filename.
Voice
Colleagues who know the participant may recognize her when they hear the recording.
Context
Her school, subject, position, and distinctive activity substantially narrow the possible speakers.
Result
The researcher should not classify the original recording as anonymous merely because P12 never says her name.
Action
The researcher manages the original audio as identifiable material and separately assesses what contextual information should remain in any transcript or quotation intended for wider disclosure.
The example also shows why identifiability is not a property of one word. Several details that seem innocuous separately may become revealing when they appear together.
07 · A Quick Checklist
Before treating an audio recording as non-identifiable
Check whether a listener could identify the participant through:
Recognition of the participant's natural voice.
Names or explicit identifiers spoken anywhere in the recording.
Occupation, employer, institution, location, relationships, or other contextual details.
Distinctive experiences or combinations of details that substantially narrow who the participant could be.
A participant code, linkage key, recruitment record, or other information available to the research team.
Information reasonably available to the people who will receive or access the recording.
The applicable ethics, institutional, legal, and data-management requirements before changing how the recording is classified or shared.
11 · Cite this Guide
How to Cite This Guide
This guide is intended to be read, shared, and used in research, teaching, and academic work. If you draw on its ideas, explanations, or other content, please acknowledge the source by citing the guide. Doing so gives appropriate credit and helps your readers locate the original resource.
Recommended (Field Guide)
APA
MLA
Chicago
Copy Citation