Skip to main content
The state of ai impact assessment
Category: Privacy Principles

Pseudonymization

Also known as: Pseudonymisation
Simply put

Pseudonymization is a data protection technique that replaces information identifying a person with a substitute value, such as a token or pseudonym, so the data can no longer be linked to a specific individual without additional, separately held information. It is intended to reduce the risk of directly identifying people while still allowing the data to be used. Unlike full anonymization, pseudonymized data can generally be re-linked to an individual if the additional information is available.

Formal definition

Pseudonymization is a de-identification technique in which one or more identifiers for a data subject (or data principal) are replaced, removed, or transformed—commonly substituted with a pseudonym or cryptographically generated token—with the identifying information stored separately and subject to controls. The goal is to prevent attribution of data to a specific individual absent the additional, separately kept information required for re-identification. It should be distinguished from anonymization: pseudonymized data generally remains capable of re-linkage and is therefore typically treated as personal data under applicable data protection regimes such as the UK GDPR, whereas the treatment of specific implementations depends on the technique used, the strength of separation between the data and the re-identification key, and the governing jurisdiction and facts. Note that terminology and legal characterization vary across frameworks and standards (for example, ICO guidance and NIST usage), and readers should verify against the current authoritative source relevant to their context.

Why it matters

Pseudonymization matters because it offers a middle path between using personal data in its raw, directly identifying form and stripping it of all utility through full anonymization. By replacing identifying fields with pseudonyms or tokens and keeping the re-identification information separately, organizations can reduce the risk of directly attributing data to an individual while retaining the ability to analyze, process, or share it for legitimate purposes. This makes it a practical technique for reducing exposure in data processing, analytics, and data sharing arrangements.

A critical point for compliance professionals is that pseudonymization is not the same as anonymization. Because pseudonymized data can generally be re-linked to a specific person if the separately held information is available, it is typically treated as personal data under data protection regimes such as the UK GDPR. That means the underlying legal obligations—lawful basis, data subject rights, security measures, and accountability—generally continue to apply. Treating pseudonymized data as if it were outside the scope of data protection law is a common and consequential misunderstanding.

The legal characterization and terminology surrounding pseudonymization vary across frameworks and jurisdictions—for example, ICO guidance and NIST usage frame the technique in related but distinct ways. Whether a particular implementation achieves its intended risk-reduction effect depends on the technique used, the strength of the separation between the data and the re-identification key, and the governing facts and jurisdiction. Because these frameworks and interpretations are periodically revised, readers should verify against the current authoritative source relevant to their context and treat application to any specific situation as requiring professional judgment.

Who it's relevant to

Data protection officers and privacy specialists
Those responsible for privacy compliance need to understand that pseudonymized data is generally still personal data under regimes such as the UK GDPR, meaning data protection obligations typically continue to apply. Pseudonymization can serve as a risk-reduction and security measure, but it should not be assumed to remove data from regulatory scope. Its legal characterization depends on the technique and the facts, so verification against current authoritative guidance is advisable.
Information security and engineering teams
Teams implementing de-identification must design the separation between pseudonymized data and the re-identification information carefully, since the technique's risk-reduction depends on that separation and the controls around the additional information or key. Choices such as tokenization versus other substitution methods affect the strength of the outcome and should be evaluated against the intended use case.
Data sharing and analytics functions
Where data is shared or analyzed, pseudonymization can allow continued use of records while reducing direct identifiability. However, because pseudonymized data can generally be re-linked when the separately held information is available, these functions should treat it as personal data and apply appropriate safeguards rather than assuming it is anonymized.
Legal counsel and auditors
Counsel and auditors assessing data handling should distinguish pseudonymization from anonymization and recognize that terminology and legal treatment vary across frameworks and jurisdictions. Whether a given implementation meets a particular standard or regulatory expectation is fact-specific and depends on the current authoritative source; conclusions require professional judgment applied to the specific circumstances.

Inside Pseudonymization

Separation of identifying information
Pseudonymization involves processing personal data so that it can no longer be attributed to a specific individual without the use of additional information, such as a key or mapping table.
Additional information kept separately
The 'additional information' needed to re-identify individuals must be kept separately and be subject to technical and organizational measures that prevent attribution to an identified or identifiable person.
Reversibility
Unlike anonymization, pseudonymization is reversible by design; the original identities can be restored where the additional information is available, which is why pseudonymized data generally remains personal data.
Technical and organizational safeguards
The measures protecting the separation and the additional information are central to the concept, encompassing controls such as access restrictions, key management, and governance processes.
Status as a data protection measure
Under the EU GDPR, pseudonymization is treated as a risk-reduction and safeguarding technique rather than a means of removing data from the scope of data protection law.

Common questions

Answers to the questions practitioners most commonly ask about Pseudonymization.

Does pseudonymization make personal data anonymous and therefore exempt from the GDPR?
No. Under the GDPR, pseudonymized data remains personal data because the information can still be attributed to an individual through the use of additional information kept separately. This distinguishes it from anonymization, which — where it genuinely and irreversibly prevents re-identification — generally takes data outside the scope of the GDPR. Because pseudonymized data stays in scope, controllers and processors remain subject to their obligations, though pseudonymization may support compliance and reduce risk. Interpretations of what constitutes effective anonymization continue to evolve, so verify against current guidance and the applicable regulatory text.
Is pseudonymization the same thing as encryption?
Not exactly. Both are referenced in the GDPR as measures that can help protect personal data, and they are related but distinct techniques. Encryption transforms data into an unreadable form that can be reversed with a key, and is typically applied to protect confidentiality of data at rest or in transit. Pseudonymization replaces identifying elements with pseudonyms so that data can no longer be attributed to a specific individual without additional information held separately. Encryption may be used as one method to achieve pseudonymization, but the concepts serve different purposes and should not be treated as interchangeable. Application to a particular setup requires professional judgment.
How should the additional information needed to re-identify individuals be handled?
The GDPR framing generally requires that such additional information be kept separately and be subject to technical and organizational measures that prevent attribution to an identified or identifiable person. In practice this often means storing the mapping or key apart from the pseudonymized dataset, with access controls governing who can bring the two together. The specific measures appropriate to a given case depend on the risk level, data category, and organizational context, and this entry does not prescribe a particular technical arrangement. Verify requirements against the current authoritative text.
Does pseudonymization reduce a controller's compliance obligations in any way?
It may support compliance rather than remove obligations. Because pseudonymized data remains personal data, the full set of applicable obligations generally continues to apply. That said, pseudonymization is recognized as a measure that can help demonstrate data protection by design and by default, support the security of processing, and factor into risk-based assessments. Whether and how it affects specific obligations in a given scenario is fact-specific and depends on the processing context, so professional judgment is needed.
Can pseudonymization on its own satisfy security requirements?
Generally no. Pseudonymization is one measure among several and is typically expected to be combined with other technical and organizational safeguards appropriate to the risk. It addresses attribution of data to individuals but does not by itself guarantee confidentiality, integrity, or availability. Note that privacy and security, while related, are distinct concerns, and a measure that supports one does not automatically address the other. The adequacy of any given combination of measures depends on the specific circumstances.
How does pseudonymization relate to data minimization and retention practices?
Pseudonymization can complement data minimization by limiting the extent to which datasets directly identify individuals, but it does not substitute for reducing the amount of personal data collected or for applying appropriate retention limits. The two operate together rather than one replacing the other, and their application should be assessed against the purpose of processing and the applicable requirements. This entry does not cover specific retention periods, which depend on the relevant legal basis and jurisdiction and should be verified against current sources.

Common misconceptions

Pseudonymized data is anonymous and therefore falls outside data protection law.
Because pseudonymization is reversible using separately held additional information, pseudonymized data generally remains personal data and, under the EU GDPR, stays within the scope of the regulation. Anonymization, by contrast, aims to make re-identification no longer reasonably possible.
Pseudonymization is a mandatory legal requirement in all cases.
Under the EU GDPR it is generally presented as a recommended safeguard and one way to help demonstrate appropriate technical and organizational measures, rather than an absolute obligation in every situation. Whether it is expected in a given case tends to depend on risk, data category, and context, and requirements differ across jurisdictions.
Simply replacing names with codes automatically achieves compliant pseudonymization.
Effective pseudonymization also depends on keeping the additional information separate and protecting it with technical and organizational measures. Without those safeguards, replacing direct identifiers may not deliver the intended risk reduction.

Best practices

Store the additional information (such as keys or mapping tables) separately from the pseudonymized data set, and restrict access to it through documented technical and organizational controls.
Treat pseudonymized data as personal data for compliance purposes unless you can demonstrate, against the current authoritative text and guidance, that a higher standard such as anonymization has been met.
Document the pseudonymization methods and safeguards applied so they can support demonstration of appropriate technical and organizational measures where required.
Apply robust key and access management governance to control who can reverse the process and under what conditions.
Reassess the adequacy of pseudonymization measures in light of the risk level, data category, and evolving regulatory interpretation, and verify obligations against the applicable jurisdiction's current requirements.
Involve appropriate legal and data protection expertise when determining whether pseudonymization is sufficient for a specific processing activity, since application to particular circumstances requires professional judgment.
a promotional banner asking how ready are you for PCI DSS 4.0? With a call-to-action to get the checklist now.