Skip to content
Delete Image Metadata
How-to

What to Check Before Sending a PDF

Four things routinely survive deletion in a PDF — and every one of them looks removed in the viewer you checked it in.

10 min read

The pattern behind all four

PDF leaks share a design. In each case the thing you removed is still in the file, and the viewer you checked it in is showing you the state you intended rather than the state on disk. The check that would catch the problem is never the check you naturally perform.

That is why these failures keep happening to careful people. Nobody skips verification out of laziness — they verify in the one place that cannot show the leak.

1. Document properties and XMP

The first is the easiest to fix and the easiest to forget. Every PDF carries an information dictionary — Title, Author, Subject, Keywords, Creator, Producer and two timestamps — and most also carry an XMP packet duplicating much of it, plus document and instance identifiers that persist across exports.

Author is usually your operating system account name, and Creator names the application, its version, and by implication your organisation's software stack. On a document going to an opposing party, a journalist or a job applicant pool, that is more than most people intend to send.

Check it in any PDF reader's properties dialog — or load the file into the tool on this site, which lists both the dictionary and the XMP fields it can decode.

2. Old revisions

This is the one that defeats the obvious fix. PDF allows incremental saving: an editor appends changed objects and a new cross-reference table rather than rewriting the file. Clear the Author field, save incrementally, and the file now contains an empty information dictionary *and* the original one, a few kilobytes earlier.

The properties dialog reads the current dictionary and shows nothing. The document looks clean in every check you would think to run, while your name sits in the bytes.

Count the occurrences of startxref in a text editor: more than one means multiple saved states. Only a full rewrite of the document removes them.

3. Content you hid rather than removed

Two versions of this, and both are worse than the metadata problems because they leak the actual content of the document.

A black rectangle drawn over a paragraph adds a rectangle. The text underneath stays in the content stream, selectable and copyable by anyone, and extractable in bulk by any PDF library. The result looks identical to a real redaction, which is exactly why it keeps being published.

Cropping has the same shape. In PDF, cropping a page usually sets a CropBox — a rectangle telling viewers which part to display. The content outside it is not deleted. Change the box back, or extract the page's content stream, and the cropped-away material is still there.

Proper redaction removes the underlying content rather than covering it. Acrobat's Redact tool does this; drawing a shape does not, and neither does a highlighter set to black. This site's tool does not redact and does not claim to — it removes metadata, which is a different job.

4. Things attached to the document

The last category is content that was never part of the page at all, and therefore never looked like something to remove.

  • Embedded files — a PDF can carry whole attachments, including the spreadsheet a chart was built from. They do not appear on any page.
  • Comments and annotations — review notes, sticky notes and mark-ups, often with author names and timestamps attached to each one.
  • Form field values — filled-in data can persist in the form's data structure even when a field looks blank on screen.
  • Hidden layers — optional content groups can hold entire drafts or alternate versions that simply are not displayed.
  • Bookmarks and named destinations, which sometimes preserve headings from a structure that has since been edited away.

The order that works

Sequence matters, because most of these steps write new metadata as a side effect of saving.

  • Redact properly first, using a tool that removes content, and flatten the result.
  • Remove attachments, comments and hidden layers in a PDF editor.
  • Remove the metadata last, with a tool that rewrites the document rather than appending to it — that step also drops the revisions the earlier steps just created.
  • Verify by re-reading the finished file, not by re-opening the dialog you edited in.

What this site's tool does and does not cover

Being specific matters more than being reassuring. The PDF tool here removes the document information dictionary, the XMP packet and page-level application data, and it rewrites the file so superseded revisions are not carried into the output. It then re-reads the result and reports what remains.

It does not redact content, remove attachments, strip annotations or flatten layers. Those need a full PDF editor, and a tool that claimed otherwise would be setting you up for exactly the kind of failure this article describes.

Check what your PDF is carrying

PDFs record who made them, with what software and when — and often keep earlier drafts inside the same file. See what yours contains, then remove it.

Open Remove Metadata From a PDF

Frequently asked questions

Why can people still read text under a black box?

Because the box is drawn on top rather than removing anything. The text remains in the page's content stream and can be selected, copied or extracted programmatically. Only true redaction deletes the underlying content.

Does cropping a PDF page delete what is outside the crop?

Usually not. Cropping typically sets a CropBox that tells viewers which region to display, leaving the rest of the content in the file. Restoring the box reveals it again.

I cleared the document properties. Why is my name still in the file?

Almost certainly because the editor saved an incremental update: it wrote a new empty properties object and left the old one in place. The dialog reads the new one. Rewriting the document is what removes the old.

Does removing metadata remove comments and attachments?

No. Comments, annotations and embedded files are document objects rather than metadata, and this site's tool preserves them deliberately rather than deleting content you may need. Remove them in a PDF editor.

What is the single most common PDF leak?

The Author field, because it is populated automatically from the operating system account and almost nobody looks at it. Failed redaction is rarer but far more damaging.

Is printing to PDF a safe way to strip everything?

It removes a great deal, including hidden layers and attachments, but it can rasterise or re-flow text and usually loses links, form fields and accessibility tagging. It is a blunt instrument rather than a clean one.

How do I verify a document is actually clean?

Re-read the finished file rather than re-opening the editor. Load it into a metadata tool, search the raw bytes for names you expected to be gone, and try selecting text in any area you redacted.

Do scanned PDFs have the same problems?

They have the metadata and revision problems, plus the scanner model and driver in the Creator field. They usually avoid the redaction problem, since the page is an image — unless OCR has added a hidden text layer, in which case a black box over the image leaves that text intact.