01 · The Question
How Can Combining Two Apparently Safe Datasets Reveal Someone's Identity?
A researcher removes names, email addresses, and participant numbers before sharing a dataset. The remaining file contains age, occupation, municipality, and several research variables. Nothing seems to identify anyone directly.
Then someone obtains another dataset containing names alongside age, occupation, and municipality. Suddenly, the variables that looked harmless in the first file can function as a bridge between the two.
This is one form of re-identification risk. Removing obvious identifiers can break the direct connection between a record and a person, but other information may allow that connection to be reconstructed.
03 · What You Need to Know
Re-Identification Risk Comes From What the Data Can Be Connected To
Re-Identification Re-Establishes an Identity Connection
NIST defines re-identification as a process in which information is attributed to de-identified data in order to identify the individual to whom those data relate. It also describes the concept more generally as re-establishing the relationship between identifying data and a data subject.
This can happen in several ways. A code key might be recovered. An unusual combination of variables might point to one person. Records may be matched with another dataset containing names. Contextual knowledge may reveal who a record describes.
The central idea is the same: an identity connection that was removed, hidden, or unavailable becomes available again.
Removing Direct Identifiers Does Not Remove Every Matching Variable
Suppose Dataset A contains no names but includes:
- age;
- sex;
- occupation;
- municipality; and
- year of appointment.
None may identify a person directly. They are potential indirect or quasi-identifying variables.
Now suppose Dataset B contains names together with occupation, municipality, and year of appointment. Those overlapping variables can provide a linkage route. If only one person has the same combination in both datasets, the previously nameless record may be connected to a named individual.
The identifying information was not necessarily hidden somewhere inside Dataset A. It emerged from the relationship between A and B.
Combining Data Adds Constraints to the Identity Puzzle
Imagine trying to identify one participant from a population of 50,000 people.
Knowing only that the participant is 45 years old may leave hundreds of candidates. Knowing that the participant is 45, works as a dentist, lives in a particular municipality, and received a professional award in a particular year may leave far fewer.
Each additional matching characteristic constrains the set of possible people. If another source associates those characteristics with a name, the identification problem may become much easier.
This is why a participant can be identified without their name appearing in the original research dataset.
Auxiliary Information Can Come From Many Places
The second source does not have to be another formal research dataset. Relevant additional or auxiliary information might come from:
- administrative records;
- professional directories;
- institutional websites;
- public registers;
- social-media profiles;
- news reports;
- commercial datasets;
- previously released research data;
- data breaches; or
- personal knowledge of the research population.
The relevant question is therefore broader than "What other dataset are we sharing?" Researchers should consider what information a realistic recipient already possesses or could reasonably obtain.
Current ICO anonymisation guidance explicitly requires consideration of whether another person could identify individuals from the information itself or by combining it with other information they possess or may obtain.
More Datasets Do Not Automatically Mean More Risk, but They Can Create More Linkage Opportunities
It would be too simplistic to say that combining any two datasets always increases identification risk. If the datasets contain no useful overlapping information, linkage may add little. Strong transformations, aggregation, controlled access, or other safeguards may also constrain identification.
The concern arises when combined information makes people more distinguishable or supplies a bridge to identity.
A dataset containing broad age groups and general research outcomes may have limited identification value. Add precise geography, employment records, event dates, or another source containing corresponding identifiers, and the situation can change materially.
Re-identification risk therefore depends on the informational relationship among datasets, not simply their number.
Rare Combinations Are Particularly Revealing
A common combination such as "female, age 30–39, teacher" may correspond to many people. A combination such as "female, age 67, university president, municipality X" may correspond to one.
This is one reason rare values and unusual combinations deserve attention during anonymization. The more distinctive a record becomes relative to the underlying population, the easier it may be to single out.
Small research populations can amplify this effect. When only a handful of people satisfy the inclusion criteria, even broad demographic information may substantially narrow the possibilities.
Linkage Can Reveal More Than Identity
Re-identification is concerning not only because someone may learn a participant's name. Once a research record is linked to an identified person, the recipient may also learn the sensitive attributes associated with that record.
Imagine Dataset A contains demographic information and a sensitive research outcome but no names. Dataset B contains names and overlapping demographic information but no sensitive outcome. Matching the records may reveal both who the participant is and the sensitive information recorded about them.
Identity disclosure
A record is connected to a particular person.
Attribute disclosure
Previously unknown information about a person can be inferred or learned, potentially even when their exact record is not fully reconstructed.
Privacy-risk assessments therefore need to consider what a successful or partial linkage would reveal, not merely whether a name can be attached to a row.
Pseudonymized Data Have an Intentional Re-Identification Route
Not every ability to restore identity is an attack or failure. In longitudinal research, researchers may deliberately retain a protected code key so participants can be linked across study waves.
Those data are generally better described as pseudonymized or coded rather than anonymous. The additional information is intentionally retained so authorised people can restore attribution when necessary.
The distinction between pseudonymization and anonymization matters because re-identification through an authorised code key is conceptually different from reconstructing identity from data intended to be anonymous.
Re-Identification Risk Depends on the Release Model
The same dataset can create different risks depending on where and how it is made available.
A restricted environment may limit users, prohibit external linkage, monitor outputs, and control which additional datasets are available. Public release makes the information available to an unknown audience with potentially diverse external data and computational resources.
NIST's guidance on de-identifying datasets consequently recommends considering the data-sharing model itself, including public release, protected enclaves, query interfaces, and other arrangements, rather than treating de-identification as a transformation detached from the release environment.
Watch Out
A dataset that presents an acceptable identification risk inside a controlled research environment should not automatically be assumed suitable for unrestricted public release. The recipients, available auxiliary information, linkage opportunities, and controls may be completely different.
Re-Identification Risk Changes Over Time
External information does not remain static. New datasets are released, online profiles expand, computational tools improve, and records that were once difficult to obtain may become searchable.
This means an anonymization or disclosure-risk assessment may need reconsideration when the information environment changes materially. The relevant question is not whether a dataset was judged safe once, but whether the assumptions underlying that judgment still hold for its current use and release environment.
NIST notes that research has demonstrated that some de-identified data can sometimes be re-identified, while its later guidance recommends evaluating disclosure risks and, where appropriate, conducting re-identification studies to assess those risks.
Re-Identification Is a Risk, Not an Automatic Fate of Every Dataset
The existence of published re-identification examples does not mean every de-identified dataset can inevitably be reconstructed. Risk varies with data granularity, population characteristics, available auxiliary information, access conditions, anonymization techniques, and realistic capabilities of potential recipients.
Researchers should therefore avoid both extremes: assuming that deleting names makes re-identification impossible, or assuming that anonymization is futile because some datasets have been re-identified.
The appropriate task is risk assessment.
| Situation |
What Changes? |
Possible Effect on Re-Identification Risk |
| Names are removed |
A direct linkage route disappears |
Risk may decrease, but indirect linkage can remain |
| Detailed geography is added |
Records become more geographically distinctive |
Risk may increase |
| A second dataset shares several variables |
New matching opportunities appear |
Risk may increase substantially for distinctive matches |
| Data are aggregated |
Individual-level detail is reduced |
Risk may decrease, depending on group sizes and outputs |
| Access moves from controlled to public |
Recipients and auxiliary information become harder to constrain |
Risk may increase |
| A new public database becomes available |
Additional linkage information enters the environment |
Previous assumptions may need reassessment |
07 · A Quick Checklist
Check the Linkage Environment Before Sharing Research Data
Before releasing de-identified or anonymized research data, check:
Identify variables that could support matching, including age, dates, geography, occupation, institutional role, and unusual characteristics.
Look for rare combinations that may single out individual participants.
Identify realistic public, administrative, commercial, research, or institutional sources containing overlapping information.
Consider what information intended recipients already possess or could reasonably obtain.
Assess what sensitive attributes could be inferred if records were successfully linked.
Evaluate the actual release model rather than assuming a dataset suitable for controlled access is suitable for public release.
Consider whether generalisation, suppression, aggregation, or another disclosure-control technique can reduce unnecessary linkage opportunities.
Reassess risk when new datasets, recipients, technologies, or uses materially change the information environment.