Metroline Step: Pseudonymisation, anonymisation and key-file management
‘Pseudonymisation is a foundational technique to mitigate data protection risks.’ (European Data Protection Supervisor)
This step explains how to protect participants’ identities while keeping research data useful. It helps researchers decide whether data should be pseudonymised or anonymised and how to separate identifying information from the research dataset. It also explains how to create, store, use and eventually archive or destroy a key file securely. For more complex studies, it points to approved systems, trusted third parties and emerging EHDS practices.
Short description
Research involving people often requires the collection and processing of personal data. Under the GDPR, personal data is any information relating to an identified or identifiable person. A person may be identified directly, for example by their name, photograph or identification number, or indirectly through a combination of variables such as age, location, diagnosis, profession, dates or rare characteristics, particularly when these can be combined with other available information. Removing names and other obvious identifiers does not therefore automatically make research data anonymous.
Pseudonymisation means replacing identifying information with a code while retaining the information needed to link that code back to the participant. This additional information may be stored in a key file, also called a subject identification log, managed through a dedicated linkage system or held by a trusted third party. Because the link can still be restored, pseudonymised data remains personal data under the GDPR.
Key files and other linkage information often contain direct identifiers, such as names and email addresses, and therefore require a high level of protection. Appropriate safeguards may include strict role-based access controls, storage in an approved environment, encryption, secure backups and auditable procedures. The measures used should reflect the sensitivity of the data, the risks associated with unauthorised access and applicable institutional requirements.
See the How to section for practical guidance on creating, storing, using and retaining key files.Anonymisation means processing data in such a way that individuals can no longer reasonably be identified, either from the data itself or by combining it with other information that is reasonably available. Properly anonymised data is no longer personal data. In practice, however, effective anonymisation can be difficult, particularly for detailed research datasets.
In many studies, retaining a controlled link to participants is necessary for follow-up, data linkage, record correction, withdrawal management or study monitoring. In such cases, pseudonymisation is often more appropriate than anonymisation.
This Metroline step explains how to choose and document an appropriate pseudonymisation approach, separate identifying information from research data and securely manage a key file throughout and after the study.
Why is this step important?
Managing pseudonymisation and linkage information properly is essential for several reasons.
Protect participants and meet data-protection obligations. Under Article 32 GDPR, controllers and processors must implement technical and organisational measures that ensure a level of security appropriate to the risk. These measures may include pseudonymisation and encryption, taking into account:
the nature of the data, such as health data and other special categories;
the purposes of the processing;
the likelihood and severity of risks to participants.
Key files and linkage systems that connect pseudonymised research data to identifiable personal data typically require a high level of protection. This usually includes a combination of controlled access, encryption, secure backups and auditable procedures.
Enable necessary study activities. Preserves a controlled link for purposes such as follow-up, data linkage, record correction, withdrawal management and study monitoring.
Maintain data integrity. Helps ensure that records remain linked to the correct participant and that authorised reidentification can be performed reliably when required.
Establish clear governance. Defines where linkage information is stored, who may access and use it, for which purposes and how long it should be retained.
Appropriate key management ensures that pseudonymisation provides meaningful protection throughout the research data lifecycle.
How to
The steps below are primarily based on practical LCRDM guidance for pseudonymisation and key-file management in research. They also incorporate relevant principles from emerging TEHDAS2 guidance for more complex secondary-use and data-linkage settings under the European Health Data Space. Steps 1–3 focus on choosing and documenting an appropriate approach. Step 4 covers separating and coding the data, while Steps 5–7 address the storage, use and retention of key files.
Running example: Alice’s study
Alice is studying recovery after a medical treatment. During recruitment, she collects participants’ names, email addresses, dates of birth, postcodes and diagnoses.
For the analysis, however, she does not need names, email addresses, exact dates of birth or full postcodes. She needs only a participant code, age group, broad geographical area and diagnosis category.
Alice must still be able to email participants for follow-up and identify their records if they withdraw. She therefore decides to pseudonymise the data using a separate key file.
The purple boxes in the seven steps below show how Alice moves from the identifiable information collected during the study to a pseudonymised research dataset containing only the information needed for analysis.
Step 1 – Determine whether a link is needed
Decide whether the research data must be linkable back to participants during or after the study, for example for follow-up, data linkage, record correction, withdrawal management, source-data verification or study monitoring.
Also determine how long the link must remain available. When no reversible link is needed, assess whether anonymisation is appropriate and achievable. Effective anonymisation may require more than removing direct identifiers, particularly when participants may still be identifiable from combinations of variables or other available information.
Example: Alice’s study
Alice needs to email participants for follow-up and must be able to identify their records if they withdraw. She therefore needs to retain a controlled link and chooses pseudonymisation rather than anonymisation.
Step 2 – Plan how pseudonymisation will be organised
Check which institutional policies, approved systems and procedures apply and determine how the link will be managed. Depending on the study, this may involve:
a key file;
a dedicated linkage system;
a trusted third party;
separate local keys in a multicentre study.
Choose an approach that is proportionate to the scale, complexity and risks of the study.
An approved system may combine several of these functions. For example, Ldot can be used to securely store personal data and generate pseudonyms, as illustrated by the AOCR biobank in the community example below.
For complex data-sharing or linkage workflows, including future EHDS use, the chosen arrangement may need to support project-specific pseudonyms, controlled linkage across organisations, clear separation of responsibilities and oversight by an authorised linkage party.
Example: Alice’s study
Alice checks her institution’s policies and approved storage options. Because this is a small, single-site study, she decides to manage the link through a separate key file.
Step 3 – Define and document the pseudonymisation approach
Apply data minimisation
Determine which personal data is necessary, which identifying information must be kept separately and which data should not be collected or retained.
Data minimisation is not limited to reducing the number of variables. Also consider whether the volume or level of detail can be reduced without compromising the research purpose. For example, exact ages may be replaced with age groups, full postcodes with broader geographical areas or exact dates with relative time intervals. This lifecycle-based approach to minimisation is also emphasised in TEHDAS2 guidance.
Identify quasi-identifiers
Identify indirect identifiers, sometimes referred to as quasi-identifiers. Variables such as detailed dates, age, location, diagnosis, profession or rare characteristics may create a risk of identification when combined. Where these variables are necessary, decide whether their level of detail can be reduced without compromising the research purpose.
Define the coding approach
Define how participant codes will be created. Codes should be unique and should not reveal or be predictably derived from identifying information, such as names, initials, dates of birth or patient numbers.
Also decide how duplicate codes will be prevented and whether different datasets, projects or recipients require different codes to prevent unintended linkage.
When datasets will be linked, assess the identification risk of the combined dataset. Variables that are not identifying in isolation may become identifying when datasets are combined.
Document the approach
Document these decisions in the data management plan or other study documentation. Include:
the purpose and expected duration of the link;
the information to be kept separately;
the coding approach;
roles and responsibilities;
storage and access arrangements;
when pseudonymisation will take place;
what will happen to the linkage information after the study.
Planning these measures before data collection helps embed privacy protection in the design of the study.
Example: Alice’s study
Alice reviews the information she collects. She decides that names and email addresses will be kept only in the key file, dates of birth will be converted into age groups and full postcodes will be converted into broader geographical areas. Each participant will receive a unique code in the format A001, A002, A003 and so on. The codes will be assigned sequentially and will not be derived from names, dates of birth or other identifying information. Alice documents these decisions, who may access the key file and how long the link will be retained.
Step 4 – Separate and code the data
Implement the pseudonymisation approach defined in Step 3.
Organise the data into three categories:
identifying and contact information, which should be stored separately;
research data, which should contain only the information needed for the study and should not include direct identifiers;
unnecessary data, which should not be collected or retained.
Where pseudonymisation is used, assign the agreed participant codes and replace direct identifiers as early as possible, preferably during data collection or promptly after receiving identifiable data.
Remove or delete data that is no longer needed where this is permitted and consistent with applicable research, legal and institutional requirements.
Check that direct identifiers have not remained in free-text fields, file names, metadata or other parts of the dataset.
Example: Alice’s study
Alice assigns one participant the code A001. She records the participant’s name and email address in the key file under A001. In the research dataset, she replaces the date of birth with the age group 50–59 and the full postcode with the area North.
Key file
Participant code Name Email address
A001 Jane Doe jane_doe@example.org
Research dataset
Participant code Age group Area Diagnosis category
A001 50–59 North Diagnosis BStep 5 – Store and protect the key file
Store the key file separately from the research data. Merely using a different file, spreadsheet tab or subfolder does not provide meaningful separation when both locations have the same access permissions.
Use approved storage and apply appropriate safeguards, including:
separate access permissions, so that only authorised roles can access the key files;
(mandatory) encryption where required or appropriate to the risk;
secure authentication (e.g. institutional accounts, strong passwords or multifactor authentication);
secure and recoverable backups;
protection against unauthorised copying, export or transfer.
Encryption is an additional safeguard, not a replacement for separate storage and restricted access. However, in health research involving special categories of personal data, key files holding the link between codes and identities should normally always be encrypted. Relying only on folder separation or subdirectories (for example within a Trial Master File/Investigator Site File/study folder) is not sufficient if both areas share similar access permissions or if access rights are not rigorously maintained.
Example: Alice’s study
Alice stores the key file in an approved restricted location with separate access permissions. She keeps the pseudonymised research dataset in a different approved location.
Step 6 – Control access to and use of the key file
Implement the access arrangements defined in Step 3.
Grant access only to people who need the key file for their role, and update permissions when responsibilities change. Access management is error-prone in practice. Misconfigured permissions on shared folders, cloud storage or collaborative platforms are a common source of incidents. To reduce risk:
give access only to authorised roles;
use role‑based access control instead of ad hoc individual permissions;
centralise the management of permissions where possible;
remove access when someone leaves the study or changes role;
review access permissions periodically;
ensure that linkage or reidentification is carried out only for an agreed purpose;
control and, where possible, record changes to the key file;
prevent uncontrolled copies, unprotected email and unnecessary merging with the research dataset;
maintain one authoritative version of the key file.
Where feasible and proportionate, separate responsibility for managing the key file from responsibility for analysing the pseudonymised research data. At the same time, ensure that authorised access does not depend on one individual, so that the key remains available during absence or staff changes.
Example: Alice’s study
Only Alice and the authorised data manager may access the key file when follow-up or withdrawal management requires it. The researcher conducting the analysis receives only the pseudonymised research dataset.
Step 7 – Retain, archive or destroy the key file
Carry out the retention or disposal plan defined in Step 3, and review it when the study ends or when the link is no longer needed.
Before acting, confirm that the decision is still consistent with:
the research purpose;
participant information and, where applicable, consent arrangements;
follow-up and withdrawal requirements;
institutional policy;
applicable legal and retention requirements.
If the key file is retained or archived, keep it accurate, protected and accessible only to authorised people.
If the key file is destroyed, document when and how this was done. Destroying the key file does not automatically make the remaining dataset anonymous. Individuals may still be identifiable from the data itself or through other reasonably available information, especially in rich clinical datasets or in combination with external registries.
Identifiability should be assessed in the context in which the data will be retained, used or shared, including any planned or foreseeable data-linkage under the European Health Data Space. Emerging TEHDAS2 guidance notes that anonymisation is not necessarily permanent: technological developments or additional available datasets may change whether identification is reasonably possible.
Example: Alice’s study
At the end of the study, Alice checks whether the link is still needed and follows the approved retention plan. She securely archives or destroys the key file as required and documents what was done.
Practical examples from the community
AOCR. The Archipelago of Ovarian Cancer Research biobank uses Ldot to securely store patient personal data and generate study-specific pseudonyms. These pseudonyms support linkage with the Netherlands Cancer Registry, managed by IKNL, and with Palga. ZorgTTP facilitates the linkage between Ldot and Palga.
In this setup:directly identifying data is stored in an approved, access‑controlled environment;
pseudonyms are generated in a way that avoids predictability or derivation from identifiers;
linkage across organisations is coordinated by an authorised trusted third party;
the governance framework specifies who may initiate linkages, for which purposes, and under which legal basis.
This illustrates how pseudonymisation, key management and trusted third parties can be combined in practice for complex, multi-centre research and secondary use of health data.
Training
Todo: ask training team.
Suggestions
This page is under construction. Learn more about the contributors here and explore the development process here. If you have any suggestions, visit our How to contribute page to get in touch or fill in our Quick feedback form.