How to remove personal data from a document: 4 methods compared

The short answer. Really removing personal data means: the values are gone from the text and the metadata, and nobody can bring them back. Black bars in a PDF usually fail that test; the text is still underneath. Manual rewriting works but is slow and misses things. The fastest and most verifiable route: automatic detection with a review step. ShareSafe.ai finds names, addresses, numbers and amounts, shows you what it found, and delivers a safe file in under a minute.

Remove personal data from your document

Four methods, compared honestly

MethodDoes it work?Risk
Black bars (PDF markup)Often notThe text sits under the bar and can be copied or searched
Manual rewritingYes, if completeSlow, inconsistent, indirect identifiability gets missed
Word document inspectorMetadata onlyAuthor and revisions disappear, the content stays
Automatic detection + reviewYesDetection can miss something, which is why review is a mandatory step

The most notorious failure is method 1. There are plenty of public incidents where "redacted" documents turned out to be fully searchable. A bar over text is formatting, not removal.

What counts as personal data in a document?

More than names. GDPR counts everything that identifies a person directly or indirectly:

  • names, email addresses, phone numbers, addresses;
  • national ID numbers, employee numbers, case numbers, licence plates, IBANs;
  • job titles in small organisations ("the controller of...");
  • combinations of date, place and event;
  • signatures and handwritten notes;
  • metadata: author, company name, revision history.

Indirect identifiability is what manual work misses most. One stray detail is rarely a problem; three details together point at one person.

Remove or replace?

Two flavours, depending on your goal:

  • Anonymise: the data is gone or generalised for good. No key, no way back. Choose this for publication, archiving or sharing outside your organisation.
  • Pseudonymise: the data becomes consistent placeholders ([NAME_01], [AMOUNT_01]) and you hold an Identity Key to translate them back. Choose this when you still want to use the document with AI or need to map results back. The difference explained: [Pseudonymisation vs anonymisation].

Done in one minute

  1. Drop your PDF, Word, Excel or text file into ShareSafe.ai. Up to 25 MB, no account.
  2. Check the before and after view. Everything detected is marked per category. Missing something? Add it.
  3. Choose anonymise or pseudonymise and download the clean file, plus the processing receipt.

Processing happens in the EU and your original is not kept after the run.

FAQ

Does this remove metadata too?
The safe file is rebuilt from the document's content; detected personal data leaves the text. Document properties of your original (author, company name, revision history) fall outside content detection. For high-sensitivity material, check those yourself or remove them in your word processor first. The before and after view shows what was found in the content.
Is an anonymised document still usable?
Yes, that is the point of consistent labels. "[NAME_01] invoices [ORG_02] for [AMOUNT_01]" remains a readable, analysable document, just without the real values.
Can I do an entire case file at once?
The free tool works per document. For recurring files and team work there is VaultLM, the workspace where safe files become searchable and shareable.
What if detection misses something?
That is why the tool shows everything it found and lets you add items before downloading. No automatic detection is infallible; detection plus human review is the safe method.

ShareSafe.ai is part of VaultLM. Raw files stay in the EU. Minimal retention. You hold the key. Try it with your own document →