PSEUDONYMISATION
A local pseudonymisation tool to replace and restore identifiers
Replace detected names, emails and other identifiers with consistent placeholders, then restore recognised placeholders in a reviewed AI response while your session mapping is available. DataAnonymiser is a local pseudonymisation tool for contracts, correspondence and other sensitive text. In the Individual desktop app, processing and restoration stay on your device.
What is pseudonymisation?
Pseudonymisation is the processing of personal data so that attributing it to a specific person requires additional information. That information must be kept separately and protected by technical and organisational measures. Replacing a name with a code is one possible step; how the surrounding information and the lookup are handled matters too. [1]
Consider a review document in which Elena becomes [first_name_1]. A reader can follow that person's actions without seeing the name in each paragraph. A separate association between Elena and the placeholder allows an authorised person to put the name back when needed.
The practical benefit is selective access to detail. Someone editing a letter can work with its wording while the person responsible for the letter keeps the identifying values. Start by deciding what the reviewer needs to know and whether a later return to those values serves a legitimate purpose.
Pseudonymisation vs anonymisation: the distinction that matters
The EDPB distinguishes pseudonymisation, which reduces a person's linkability, from anonymisation, which removes identifiability. Truly anonymous information falls outside EU data protection law. The presence of placeholders alone does not establish that outcome. [2]
For an organisation that can reconnect the records to people using separately held information, the records remain personal data. Pseudonymisation is a safeguard, not an automatic exemption from obligations governing the processing. Assess the intended recipient and context before making claims about a particular disclosure. [1]
Deleting one mapping does not prove that every route to identification has disappeared. The original document may still exist, or the text may describe a recognisable incident. For example, an unnamed employee who received a publicly announced award last week may remain easy to identify from that description.
In practice, decide whether you need temporary substitution for editing, controlled linkage for analysis or an output assessed as anonymous. Those are different objectives. A short document-review session and a research database with years of follow-up require different arrangements.
Pseudonymisation methods: placeholders, tokens, hashing and encryption
Methods include tokenisation, hashing and encryption. Tokens substitute an identifier with another value and can use a separately controlled lookup. Hashing derives a value from an input; predictable inputs and unprotected hashing can permit guessing attacks. Encryption can transform identifiers using a key. The suitability of each method depends on its design and protections. [3]
For prose, readable placeholders offer a useful representation. A sentence containing [first_name_1] and [first_name_2] is easier to edit than one containing long opaque identifiers. The labels also reveal categories, however, so decide whether revealing that a value is a name or an account number is appropriate.
DataAnonymiser substitutes detected values with labelled placeholders and supports local restoration from the session mapping. These placeholders are not encrypted versions of the entire document. The surrounding text remains readable and needs its own review.
Masking only part of a value, such as leaving several account-number digits visible, serves a different practical purpose from replacing the whole value. For a proofreading task, even those last digits may be unnecessary. Choose what the task needs rather than retaining detail simply because a particular mask allows it.
The ICO source describes these methods in the UK context and is currently marked as under review following legislative changes. It provides technical background here; it is not presented as a statement of EU law or a product certification. [3]
Example 1: review a document without losing track of the parties
The examples below use fictional details and illustrative substitutions. They show the decisions involved in a workflow, not measured product outputs or proof of legal anonymity.
Original: “Elena will send the draft to Jonas. Jonas will return his comments to Elena before the review meeting.”
Prepared text: “[first_name_1] will send the draft to [first_name_2]. [first_name_2] will return his comments to [first_name_1] before the review meeting.”
A reviewer can check the sequence of responsibilities without seeing the names. Replacing both people with an identical blank would lose that distinction. A request to shorten the paragraph can explicitly tell the reviewer to preserve the placeholders and who performs each action.
Check aliases and spelling variations separately. A name, surname and email address may refer to the same person without being the same extracted value. Consistent substitution is not a guarantee of identity resolution. Read the result to confirm that it still expresses the intended relationships.
Example 2: draft a response with a local return to the real values
Original: “Lena at lena@example.com reported a duplicate invoice. Amir will check the billing records and contact Lena with an update.”
Prepared text: “[first_name_1] at [email] reported a duplicate invoice. [first_name_2] will check the billing records and contact [first_name_1] with an update.”
A drafting request could say: “Write a short acknowledgement to [first_name_1]. Say that [first_name_2] will investigate. Preserve the placeholders exactly and do not invent a refund decision or deadline.”
An illustrative reply is: “Hello [first_name_1], thank you for flagging the duplicate invoice. [first_name_2] will check the billing records and follow up with an update.” Local restoration can reinsert Lena and Amir while the relevant session mapping remains available.
The mapping should stay out of the material supplied for drafting. After restoration, inspect the message for correct names, responsibilities and claims. The reverse step replaces recognised tokens; it cannot tell you whether the draft made the right person responsible for a task.
Example 3: recognise when a research workflow needs a lasting identifier
Imagine a study comparing questionnaire answers across three visits. A participant code lets authorised researchers connect the visits without displaying the participant's name in every analysis file. The lookup, access arrangements and retention rules are part of that study's design.
That requirement differs from editing a document during one working session. Researchers may need to reconnect a code to a participant months later, and to distinguish that participant from everyone else throughout the study. They need an approved process designed to maintain that continuity.
DataAnonymiser's session mapping is not a persistent participant registry. Do not rely on a placeholder such as [first_name_1] as a durable identity across independent runs. The same label in another document does not establish that the two documents concern the same person.
A document tool can still help prepare a particular excerpt for review, but it does not replace the study's identity-management process. Establish which system is responsible for long-term linkage before choosing how to transform individual documents.
What happens to the mapping when the session ends?
In the standalone desktop workflow, DataAnonymiser holds the placeholder-to-original associations in memory for the working session. It does not write that mapping to disk. When the app closes and the mapping is lost, reopening the output does not recreate those associations.
Finish any restoration that you need while the relevant mapping is available. Keep the response's placeholder spelling and numbering intact. If a reviewer or model changes a token, the reverse step may leave it unmatched. If it swaps two valid tokens, both can restore successfully while assigning the wrong values to the sentences.
The practical boundary is the mapping the app creates. Closing DataAnonymiser does not delete the original source document, a copy you saved elsewhere or information another person already knows. Session-only storage therefore limits persistence of the lookup; it is not a guarantee that nobody can identify a person from the remaining material.
Once values have been restored, the result contains those details again. Decide where that final document belongs before copying it into another application. A workflow can minimise the drafting input and still require careful handling of the finished output.
A five-step pseudonymisation workflow for document review
1. Define the review task. Identify the recipient, required context and whether restoration is needed. A grammar check, an action-item summary and a long-term analysis have different information requirements. Select the smallest excerpt that can support the task.
2. Prepare the text. Run the relevant material through DataAnonymiser and inspect the substitutions. Where your plan includes custom terms, add sensitive client references or internal names that are specific to the work. Check that different people remain distinguishable.
3. Review what remains. Look beyond names and contact details to job titles, locations, dates and distinctive events. Remove unnecessary passages manually when they reveal too much. Resolve residual-risk warnings and incomplete detection before making a sharing decision.
4. Share the reviewed version through the approved workflow. Keep the original and the mapping separate from what the recipient receives. Check accompanying attachments, quoted messages and filenames. If a return step is planned, instruct the recipient to leave every placeholder unchanged.
5. Restore and verify. Bring the edited response back while the relevant mapping remains available. Check each person's role and the factual meaning, then use the result only for its intended purpose. Record the review decision in your organisation's process without copying unnecessary identifying details into that record.
What pseudonymisation does not remove
A document can disclose sensitive facts without containing a name. “The only person in the branch who started parental leave this week” may identify a colleague to that branch. A detailed chronology can connect a case to public reporting. Ask what a recipient could infer from the whole text.
It can also reveal confidential business information unrelated to personal identity. Replacing participant names does not conceal an unannounced acquisition, pricing strategy or proprietary design. Decide whether those facts belong in the recipient's hands before treating a clean identifier scan as sufficient.
Readable output needs appropriate handling after substitution. Avoid bundling the original document with the reviewed copy or pasting the mapping underneath it for convenience. If the task genuinely requires the original identities, use the approved process for sharing that information.
Pseudonymisation can reduce exposure while supporting useful work. It does not supply a lawful basis, permission to reuse a record or a universal retention period. Those decisions depend on the processing purpose and applicable obligations. [1]
How to choose a pseudonymisation tool for your workflow
Evaluate a tool with representative fictional documents before using real material. Include repeated names, two people with similar names, an identifier split across lines and a distinctive description without an obvious name. Check what the output preserves as well as what it removes.
Examine the return step. Edit a response while keeping its placeholders intact, then check restoration. Try a changed or missing token so you understand how the workflow behaves when a reviewer rewrites it. Decide whether a person can reliably spot and resolve that situation.
Ask where the original is processed, where the lookup lives, who can access it and how long it remains available. A session-only mapping may fit a short drafting task; a workflow requiring recovery next month needs a different retention design. Include the original records and external copies in that decision.
Also check coverage and review controls. Rules-only detection, model-assisted extraction and explicit custom terms have different roles. No detection mode can know all the context a recipient has. Make the human review responsibility clear before introducing the tool into a recurring process.
Use DataAnonymiser on infrastructure you control
The Individual edition processes content on your device. Enterprise can process it on a server your organisation installs and operates within its own network, activated through explicit administrator configuration. Neither edition routes document content to DataAnonymiser-operated infrastructure for processing.
For standalone document work, consistent placeholders and an in-memory mapping support a local editing round trip. DataAnonymiser presents review information alongside the output so you can inspect detected substitutions and residual-risk warnings before copying the text elsewhere.
Start with one recurring task, such as reviewing a draft or preparing a customer reply. Define what the next reader needs, try the workflow with fictional text and check the result through restoration.
Frequently asked
Questions about this workflow.
What is the difference between pseudonymisation and anonymisation?
Pseudonymisation reduces the link to a person while separately held information can allow attribution. Anonymisation reaches the different outcome of removing identifiability. A document containing placeholders needs assessment in its actual context; the visual appearance of redaction does not decide its legal status.
Where does DataAnonymiser keep the placeholder mapping?
In the standalone desktop workflow, the mapping is held in memory for the working session and is not written to disk. It enables local restoration while available. This statement concerns the app's mapping; it does not mean that original documents or copies elsewhere have been deleted.
Can I reverse a document after restarting the app?
No. Restarting the app does not recover the previous session's mapping. Finish the restoration you need while that mapping is available. Reprocessing a source document is a new run and should not be assumed to recreate identical placeholder assignments.
Does pseudonymisation make us GDPR compliant?
No. It can support safeguards, but the organisation must still assess the purpose, applicable legal basis, recipient, security and retention arrangements. A tool cannot determine all those conditions from the text. Review both the transformed document and the wider process.
Will the same person get the same placeholder in every document?
Do not assume identity continuity across independent runs. Consistency within a working document does not make its placeholders permanent person identifiers. A workflow that links people across documents or time needs an explicit linkage design and appropriate controls.
Can an AI response always be restored correctly?
Restoration can replace recognised placeholders while the corresponding mapping is available. It cannot repair a changed token, recover omitted wording or detect that the AI assigned an action to the wrong person. Ask for unchanged placeholders and inspect the result after restoration.
Related
Where else this comes up.
ANONYMIZE DATA FOR CHATGPT
Anonymize data for ChatGPT with local PII redaction
Redact personal details before pasting into ChatGPT, then restore recognised placeholders in the response locally during the same session.
LEGAL DOCUMENTS AND CLIENT CONFIDENTIALITY
Anonymize legal documents locally before AI review
Prepare contracts and client correspondence with local redaction software, consistent placeholders and session-based restoration after approved AI review.
COMPARE
How this compares to the alternatives
Cloud redaction APIs, browser-only extensions and self-hosted libraries — what each one trades away.
REFERENCE
What is data anonymisation?
How anonymisation, pseudonymisation and masking differ in law, and which one actually takes data outside GDPR.
DataAnonymiser supports redaction and review. It does not guarantee complete detection, legal anonymity or GDPR compliance. This guide provides general information, not legal advice; assess outputs and their intended use.