You scan a multi-page report and save it as a PDF. When you open it and press Ctrl+F (or ⌘+F on Mac), nothing is found - not even a word you can see clearly on the first page. The document is an image, and images do not contain searchable text.
Making a scanned PDF searchable adds an invisible text layer behind the images. The text layer is what Ctrl+F, document management systems, and search engines read. The visual appearance of the document stays exactly the same.
What OCR does
Optical Character Recognition (OCR) software analyses each pixel of each page and identifies characters, words, and layout structure. It then places those characters in a hidden layer at the same position as the corresponding image text. The result is a PDF that looks identical to the original scan but has real, selectable, searchable text underneath.
How to make your PDF searchable with Pagivo
- Open Ask PDF on Pagivo and log in if prompted.
- Upload your scanned PDF.
- Type a request such as make this PDF searchable with OCR.
- Click Process and download. Pagivo runs OCR on pages that need it and returns the document with a searchable text layer added.
- Open the result in any PDF viewer and press Ctrl+F. Your search terms will now be found.
Factors that affect accuracy
OCR works best on clean, well-scanned documents. Several factors affect how accurate the text layer will be:
- Scan resolution. 300 DPI is the practical minimum for reliable OCR. Lower resolutions produce unreliable character recognition, especially for small text.
- Page straightness. A skewed or curved page confuses character boundary detection. Flatbed scanners consistently produce better results than phone cameras for this reason.
- Typeface clarity. Standard printed fonts are recognised with very high accuracy. Decorative fonts, very small text (under 8pt equivalent at 300 DPI), and heavily compressed scans produce more recognition errors.
- Handwriting. Most OCR engines are trained on printed text. Expect lower accuracy or mostly unusable output on handwritten pages.
When you need to edit the content, not just search it
OCR creates a searchable PDF - the text is real but the page still looks like a scan. If you need to edit the content (correct a name, change a date, reformat a table), you need to convert the PDF to an editable format. The recommended workflow is: run OCR first to create a searchable PDF, then convert that PDF to Word using PDF to Word for the best conversion quality.
Using Google Drive as a free alternative
Google Drive has built-in OCR at no cost. Upload the PDF to Google Drive, right-click it, and choose Open with Google Docs. Google runs OCR and opens the result as an editable document. The page formatting will not be preserved, but the text will be extracted. This works well for short documents where you only need the text content, not the original layout.