The three-second test
Open the PDF, drag-select across the black bar as if you were highlighting a sentence, copy, and paste into any text editor. If the redacted words appear, the redaction failed. That is the whole test, and it catches the most common failure on its own. Two cautions. A negative result proves less than it looks: if the text sits under an image rather than a drawn shape, or if it is present but invisible somewhere else on the page, selection may miss it entirely. And in some viewers the selection highlight is hidden behind the black bar, so you cannot see what you have grabbed — paste anyway and look at the result rather than the page.
Why a black box is not a redaction
A PDF page is a list of drawing instructions carried out in order. "Write this text here" is one instruction; "fill this rectangle with black" is another. Drawing the rectangle after the text hides it from view, but the instruction to write the text is still in the file, unchanged. Real redaction deletes that instruction. The distinction matters because almost every tool that can draw a shape will happily draw one over text, and nothing in the interface warns you that the words survive. Highlighting text in black has the same problem, as does covering it with a white box on a white page — invisible to the eye, fully intact in the file.
The five failure patterns
These are the ones worth testing for, in rough order of how often they turn up:
- Text under a filled shape. A rectangle drawn over live text. Caught by the copy test.
- Text under a pasted image. Someone patches a screenshot over the line instead of drawing a shape. Selection often skips it, because there is no text where you are dragging — the text is beneath the picture.
- Invisible text. Text set to render invisibly. It appears nowhere on screen or in print, so there is no black bar to select, and no visual cue that anything is there.
- Text the same colour as its background. White on white, or any colour matched to the fill behind it. Looks blank, extracts perfectly.
- Content left in document properties. The page is genuinely clean, but the name is still in Title, Author, Subject or Keywords, where it travels with the file.
The scanned-document trap
This one deserves its own section because it defeats the intuition that a scan is safe. A searchable scan is a picture of a page with an invisible text layer laid over it — that layer is what makes the text selectable. If someone redacts the scan by drawing a black box on the picture, the picture is covered but the invisible layer underneath is untouched. The words are still extractable, and because the visible page is an image, nothing about it looks like live text. We built exactly this case as a test fixture: a scanned page with a full OCR layer and a black box drawn over one line. The canary text was still recoverable with an independent PDF library. If you are redacting anything that came from a scanner, assume the text layer exists until you have checked.
What a search test can and cannot prove
Searching the finished PDF for a redacted word is a reasonable second check, and it catches cases that selection misses — including some invisible text. But it only finds what you think to search for. You can search for a name you know you removed; you cannot search for the account number in an exhibit you did not read closely, or for the third occurrence of a term you only remembered twice. Search confirms specific removals. It does not survey the document.
Checking the whole document at once
The manual tests each cover part of the problem and none covers all of it. Selection misses images and invisible text. Search misses anything you do not think to type. Neither looks at document properties. The PDF Redaction Checker runs the equivalent checks across every page and reports what is still recoverable, along with the page it came from. It runs inside your browser — the file is never uploaded, which matters because a document being redacted is sensitive by definition. It is not a complete security audit, and the page says so plainly. It reads text, so a signature, photograph or chart hidden under a box is outside what it can see. It cannot tell you whether text that is plainly visible should have been removed. A clear result means these checks found nothing recoverable — not that the document is safe to publish.
If the check finds something
Redact the document again with a tool that removes the underlying text rather than covering it, then check the result. Verifying the finished file is the step people skip, and it is the only one that would have caught any of the published failures. Work on a copy and keep the original intact — redaction is meant to be irreversible, so there is no undo once it is applied correctly. Clear the document properties in the same pass. Then run the finished file through the checker before it leaves your machine.