OCR PDF: make a scanned PDF searchable
Typed letters and printed records become text you can find, copy and reuse.
Private: your files never leave this deviceNothing is uploaded or stored on a serverWorks right here in your browserAdd a scanned PDF to start.
How to make a scanned PDF searchable
Add the scanned PDF
Choose the file or drop it on the box. The page count appears once it is read.
Pick the language
English is the default. Add the languages the pages are written in, and type a page range if you only need some.
Make it searchable
Press Make searchable. Each page is read in turn; Stop cancels. Download the PDF, or copy the text on its own.
Good to know
- 300 dpi suits normal print. Use 400 dpi for small type such as footnotes, and 200 dpi for large, clear text to save time.
- A page takes a few seconds; a 50 page scan takes a few minutes on a laptop. The page stays usable meanwhile.
- Crooked or dark phone photos read badly. Scan to PDF can straighten and clean them up before you OCR.
- Search in the result with Ctrl+F in any PDF viewer; copied text follows the reading order the engine found.
Questions people ask
What does OCR PDF do?
A scanned PDF is a stack of pictures, so you cannot search it or copy from it. OCR reads the words on each page and adds them as invisible text in exactly the right place. The pages look the same, but Ctrl+F, selecting and copying now work.
Does it change how my pages look?
No. The original pages are kept as they are and the text layer is drawn on top, invisible. The file grows by only a few kilobytes per page.
Does it work for Tamil and other Indian languages?
Yes. Pick Tamil, Hindi, Telugu, Kannada, Malayalam, Bengali or any other listed language, or up to three together for mixed pages. The hidden text uses a font that carries no shapes, so every script works. Accuracy depends on the print: clean scans do well, handwriting does not.
Why were some pages skipped?
With Skip pages with text on, pages that already contain real text (typed documents, or pages OCR'd before) are left alone. Untick it to read every page anyway.
Are my PDF files uploaded?
No. Pages are drawn with pdf.js and read by the Tesseract OCR engine inside your browser. The first time, the browser downloads the engine and the language data once and keeps them for next time; your PDF never leaves your device.