Skip to main content
The state of ai impact assessment
Category: Technical Controls

Hashing

Also known as: Hash function, Hashing algorithm
Simply put

Hashing is the process of running data through a mathematical algorithm to produce a fixed-length value that represents the original data. This value, often an alphanumeric string of predetermined length, is generally designed to be irreversible, meaning the original data cannot be practically reconstructed from it. In cybersecurity, hashing is commonly used to help protect sensitive information such as passwords, messages, and documents, while in computing more broadly it is also used to store and retrieve data efficiently.

Formal definition

Hashing applies a deterministic mathematical algorithm (a hash function) to input data of arbitrary size to produce a numeric or alphanumeric output of fixed, predetermined length that is representative of that input. Cryptographic hash functions are generally designed to be one-way (irreversible), such that deriving the original input from the output is computationally infeasible; this property underpins their use in protecting sensitive data such as passwords and in verifying data integrity. Note that hashing serves distinct purposes across contexts: in data structures it enables near-constant-time access and reduced storage overhead, whereas in security contexts the emphasis is on irreversibility and collision resistance. Not all hash functions are suitable for security use; non-cryptographic hashes used for indexing or lookup do not provide the same guarantees, and specific algorithm choices and their continued suitability change over time and should be verified against current authoritative guidance.

Why it matters

Hashing is foundational to how organizations protect sensitive information and verify that data has not been altered. In security contexts, storing a hash of a password rather than the password itself means that even if a data store is compromised, an attacker does not directly obtain the original credentials. Because cryptographic hash functions are generally designed to be irreversible, they provide a mechanism for confirming that a value matches a known input without retaining the input in a recoverable form. This distinction matters for compliance programs concerned with data minimization and reducing the impact of a breach.

Hashing also supports data integrity verification: comparing the hash of a received message or document against an expected value can reveal whether the underlying data was changed in transit or at rest. This makes hashing relevant to a range of obligations that touch on the confidentiality and integrity of information. However, it is important not to overstate its protective effect. Hashing is not encryption, and not every hash function is appropriate for security use — non-cryptographic hashes used for indexing or lookup do not provide irreversibility or collision-resistance guarantees. The suitability of specific algorithms also changes over time as cryptanalysis advances.

Because of this, hashing should be treated as one control among several rather than a complete safeguard. Whether a particular hashing approach adequately protects sensitive data is a fact-specific question that depends on the algorithm chosen, how it is implemented, and current authoritative guidance. Organizations should verify their choices against up-to-date standards and recognize that a technique considered sound today may be deprecated in the future.

Who it's relevant to

Information Security Professionals
Those responsible for protecting credentials and sensitive data rely on hashing to store passwords and other information in a form that is generally designed to be irreversible. They must distinguish cryptographic hash functions suitable for security from non-cryptographic hashes used only for indexing, and verify that chosen algorithms remain appropriate against current authoritative guidance, since suitability changes over time.
Auditors and Assessors
Practitioners evaluating an organization's data protection controls need to understand where hashing is used to protect sensitive information and to verify data integrity. Assessing whether a given implementation is fit for purpose is fact-specific and depends on the algorithm and its continued suitability, so evidence should be checked against up-to-date standards rather than assumed to be adequate.
Software and Data Engineers
Those building systems encounter hashing in two distinct roles: as a security control emphasizing irreversibility and collision resistance, and as a data-structure technique enabling near-constant-time access and reduced storage overhead. Recognizing that these purposes require different hash functions is important to avoid using a lookup-oriented hash where a security guarantee is needed.
Data Protection and Compliance Officers
Personnel overseeing how sensitive data is handled may treat hashing as one measure supporting confidentiality and integrity. It is not a complete safeguard and is distinct from encryption; whether it satisfies a particular obligation depends on the specific circumstances and requires professional judgment against the latest authoritative sources.

Inside Hashing

Hash Function
A deterministic algorithm that maps input data of arbitrary size to a fixed-length output (the hash value or digest). The same input generally produces the same output, while even a minor change to the input typically yields a substantially different digest.
One-Way Property
A cryptographic hash function is designed to be computationally infeasible to reverse, meaning the original input generally cannot be reconstructed from the digest alone. This distinguishes hashing from encryption, which is reversible with the appropriate key.
Collision Resistance
The property that it should be computationally impractical to find two distinct inputs producing the same digest. Some older or weaker algorithms have known weaknesses in this area, which is why algorithm selection matters.
Salt
A unique, typically random value added to an input (commonly a password) before hashing. Salting helps defend against precomputed lookup attacks by ensuring identical inputs do not produce identical stored digests.
Distinction from Encryption and Pseudonymisation
Hashing is not encryption (it is not intended to be reversed) and is not automatically equivalent to anonymisation. Under some data protection regimes, a hash that can be linked back to an individual may still be treated as personal data or as pseudonymised data rather than anonymous data; this depends on the facts and the applicable framework.
Common Use Cases
Typical applications include verifying data integrity, storing password representations rather than plaintext credentials, and supporting digital signatures and file fingerprinting. The suitability of a given approach depends on the use case and threat model.

Common questions

Answers to the questions practitioners most commonly ask about Hashing.

Does hashing personal data mean it is no longer personal data under data protection law?
Not necessarily. Hashing is often mistaken for a technique that removes data from the scope of regulations such as the GDPR, but a hash of an identifier can frequently be linked back to an individual, particularly where the input space is small or predictable and the hash can be recomputed or matched. In most cases, hashed identifiers are treated as pseudonymised rather than anonymised data, meaning they generally remain personal data and within regulatory scope. Whether a given hashing approach achieves genuine anonymisation is fact-specific and depends on the wider context, so this should be verified against current guidance and professional judgment rather than assumed.
Is hashing the same as encryption?
No. Hashing and encryption are distinct concepts that are commonly conflated. Encryption is designed to be reversible by a party holding the appropriate key, so the original data can be recovered. Hashing is a one-way transformation intended not to be reversed; there is no key that converts a hash back to its input. This difference matters when selecting controls for a given purpose: encryption is generally used to protect confidentiality where the data must later be recovered, whereas hashing is typically used for integrity verification or storing values that never need to be reconstructed. The two serve different roles and are not interchangeable.
How should hashing be applied when storing passwords?
For password storage, hashing is generally applied together with additional measures rather than on its own. Common practice is to use a per-value salt to defend against precomputed lookup attacks and a function designed to be computationally costly so that guessing attempts are slowed. The appropriate configuration depends on the threat model, performance constraints, and the sensitivity of the accounts involved. Because recommended approaches evolve as computing power increases and as prior methods are deprecated, implementers should confirm current recognised practice against authoritative technical sources before finalising a design.
What role does a salt play, and when should one be used?
A salt is an additional value combined with the input before hashing so that identical inputs produce different hash outputs. Its main purpose is to frustrate attacks that rely on precomputed tables of hashes and to prevent two identical stored values from being visibly the same. Salts are generally advisable when hashing values drawn from a limited or guessable range, such as credentials. Whether and how to apply a salt is an implementation decision that should be made in light of the specific use case and current technical guidance.
How does hashing support data integrity verification?
Because a hash function produces a consistent output for a given input and a change to the input generally produces a different output, comparing a freshly computed hash against a previously recorded one can indicate whether data has been altered. This makes hashing useful for detecting unintended or unauthorised modification of files or messages. It is worth distinguishing this integrity use from confidentiality protection, which hashing does not provide. The strength of the assurance depends on the properties of the specific hash function used, and functions once considered adequate may later be deprecated.
Does using hashing satisfy a specific regulatory or standards obligation on its own?
Hashing may form part of the technical measures an organisation adopts, but its use alone should not be assumed to satisfy any particular obligation. Regulatory requirements and voluntary standards typically call for measures appropriate to the risk rather than mandating a named technique, and how hashing contributes depends on the data category, purpose, and surrounding controls. Whether a given implementation meets a specific requirement is fact-specific and should be assessed against the current authoritative text and with appropriate professional judgment.

Common misconceptions

Hashing is a form of encryption.
Hashing and encryption are distinct. Encryption is designed to be reversible with a key, whereas a cryptographic hash function is intended to be one-way and not reversible to recover the original input. Conflating the two can lead to incorrect security and compliance decisions.
Hashing personal data automatically makes it anonymous and therefore outside the scope of data protection law.
Whether hashed data falls outside a regime such as the GDPR is fact-specific. Where a hash can still be linked to an individual, it may generally be treated as pseudonymised or as personal data rather than truly anonymous. Readers should assess this case by case and verify against the applicable authoritative text and guidance.
All hash algorithms provide equivalent security.
Hash functions vary in their properties, and some older algorithms have known weaknesses affecting collision resistance. Algorithm choice should reflect the intended use, current recommendations, and the relevant threat model rather than an assumption of interchangeability.

Best practices

Select a hash function appropriate to the use case, avoiding algorithms with known weaknesses and preferring those aligned with current authoritative recommendations, which change over time and should be reverified periodically.
For storing password representations, use a unique salt per credential to reduce the effectiveness of precomputed lookup attacks.
Do not treat hashing as encryption; where confidentiality that must be reversible is required, use encryption instead and keep the two mechanisms conceptually separate.
Do not assume that hashing personal data removes it from the scope of data protection obligations; assess on a case-by-case basis whether the result is anonymous, pseudonymised, or still personal data under the applicable regime.
Document the algorithm, parameters, and rationale for hashing decisions so they can be reviewed and updated as recommendations evolve.
Verify algorithm choices and any compliance-related assumptions against the latest authoritative guidance and, where the application to specific circumstances is unclear, seek appropriate professional judgment.
Promotional banner graphic asking if you are ready for PCI DSS 4.0 with a call-to-action to get the guide