Reference
What is data anonymisation?
Data anonymisation is the process of altering personal data so that no individual can be identified from it, by anyone, using any means reasonably likely to be used. Done properly it takes the data outside the scope of GDPR entirely. That last part is why the bar is so much higher than most people expect.
The definition, and the word that carries it
Anonymisation removes or alters the information that links a record to a person, so that the person can no longer be identified either directly or indirectly. The load-bearing word is indirectly. Deleting a name does not anonymise a record that still contains a postcode, a date of birth and a job title, because those three together identify most people in a population.
The test is not whether you removed the obvious identifiers. It is whether identification remains possible for anyone holding the result, including by combining it with other data they can reach.
Anonymisation, pseudonymisation and masking are not synonyms
These three get used interchangeably and mean quite different things in law.
Anonymisation is irreversible for every party. There is no key. GDPR Recital 26 states the regulation does not apply to anonymous information, so genuinely anonymised data carries no obligations.
Pseudonymisation replaces identifiers but keeps a separate piece of information that can reverse it. Under Article 4(5) the reversing information must be kept separately and protected. Pseudonymised data is still personal data and stays in scope. Articles 25 and 32 treat it as a recognised safeguard, not an exemption.
Masking describes the mechanical act of obscuring characters, such as showing only the last four digits of a card number. It may achieve either of the above or neither, depending on what is left.
Most tools sold as anonymisers perform pseudonymisation: they hold a durable key — a vault, a lookup table, a mapping column — that turns the placeholders back into people. That key is the additional information Article 4(5) is about, and while it exists the data stays in scope.
DataAnonymiser does not keep one. The mapping lives in memory for the session you created it in and is destroyed when that session ends, so for every party you hand the output to there is no key in existence to reverse it with. What that does not settle is the second half of the test: whether the remaining text still identifies someone indirectly. Detection is best-effort, and quasi-identifiers survive it. Judge the output, not the label on the tool.
The main techniques
Suppression deletes the identifying field outright. Simple, and costly in analytical value.
Generalisation reduces precision until a value stops being distinctive: an exact age becomes a band, a full postcode becomes its first half, a timestamp becomes a month.
Perturbation adds calibrated noise, so individual records are wrong but aggregate statistics stay close to true. Differential privacy is the rigorous form, giving a mathematical bound on what any single record contributes to the output.
Aggregation reports only group totals, never rows, with a minimum group size so small groups cannot be isolated.
Synthetic data generates an artificial dataset with similar statistical structure. Useful, though a model trained on real records can leak them if it overfits.
k-anonymity and its refinements l-diversity and t-closeness are the measures used to check the result: each record must be indistinguishable from at least k-1 others on the quasi-identifiers.
Why anonymisation keeps failing
The recurring failure is the same one every time. Data that looks anonymous in isolation becomes identifying when joined against another dataset. Sparse, high-dimensional data is especially fragile, because the more attributes a record has, the more likely its combination is unique to one person.
Anonymisation is also not permanent. A dataset released today may be re-identifiable in five years against auxiliary data that does not exist yet. This is why regulators treat anonymisation as an ongoing assessment rather than a one-time transformation, and why conservative practice keeps controls on the output rather than declaring the problem solved.
Free text and images are a different problem
Most anonymisation literature assumes structured data, where the columns holding identifiers are known in advance. Free text has no columns. A name, a diagnosis or an account number can appear anywhere, in any phrasing, and the identifying detail is often circumstantial rather than a recognisable format.
That is why free-text redaction is best-effort by nature: what is not detected is not removed. Anything derived from unstructured text should be reviewed before it is treated as safe.
Where a local tool fits
If the goal is to stop identities reaching a third party such as an AI assistant, a support desk or an analytics vendor, full anonymisation is usually more than you need and less than you can verify. Pseudonymising locally, before the text leaves the machine, addresses the actual exposure: the recipient never receives the identities, and the information needed to reverse it never leaves your device.
DataAnonymiser works this way, and holds the reversing mapping in memory for the current session only, so once you close the app no key to that output exists anywhere. It is a practical aid, not a compliance determination.