OCR PDF

Make scanned PDFs searchable and copyable — text recognition runs entirely in your browser.

How to make a scanned PDF searchable

Select your scanned PDF

Drop in the scan you want to make searchable, or click to browse.

Choose the language

English, Indonesian, or both at once for mixed-language documents.

Recognize & download

The searchable PDF downloads automatically — grab the plain text too.

Why PDFMatic?

Private by Design

Text recognition runs in your browser. Contracts, invoices, and archives never touch our servers.

Searchable & Copyable

Your pages keep their original layout and size, with an invisible text layer added on top that you can search, select, and copy.

Plain Text Included

Besides the searchable PDF, download the recognized text as a .txt file for further editing.

Getting a good result from OCR

First check whether you need OCR at all

Open the PDF and try to drag-select a line of text. If the words highlight, the file already carries a real text layer and OCR would only make it worse — you can copy from it as it is. If your cursor draws a rectangle across the page instead, the page is a picture of text and OCR is exactly what it needs.

Some PDFs are a mix: a digital document with a scanned receipt pasted in, for example. Running OCR processes every page uniformly, so it is worth extracting just the scanned pages with the Split PDF tool first, recognising those, and merging the result back.

What actually changes accuracy

Recognition quality is decided almost entirely by the input, not by settings:

  • Scan resolution. Around 300 DPI is the sweet spot. Photos of a page taken with a phone tend to do worse than a flatbed scan, mostly because of uneven lighting.
  • Straightness. Even a few degrees of skew measurably hurts results, and the fix for that is re-scanning the page squarely. A scan that came out a full quarter turn sideways is easier: rotate it in the Organize PDF tool before recognising it.
  • Contrast. Crisp black on white beats a grey photocopy, a coloured background, or a page with a stamp across the text.
  • Language choice. Pick the languages the document actually uses. Selecting both English and Indonesian for a document that is only one of them gives the recogniser more ways to be wrong, and takes longer.

Handwriting is out of scope — the engine is trained on printed type, and cursive notes in a margin will not come through. Very small print and heavily stylised display fonts are similarly unreliable. Words the engine is not reasonably confident about are left out of the text layer rather than guessed at, which keeps the searchable text trustworthy at the cost of missing the occasional word.

What you get back

The searchable PDF is rebuilt rather than edited in place: each page is rendered at up to twice its nominal size, stored as a JPEG, and given a transparent text layer positioned over the words the engine found. Page dimensions are preserved, so the document looks and prints the same and every viewer can search it — but because the page is re-encoded, this is not the tool to reach for if you need the original scan preserved untouched. Keep your original file if that matters.

Recognition is the slow part, and it runs on your own processor: expect a few seconds per page, longer on a phone. Once it finishes, the searchable PDF downloads on its own and a separate button offers the recognised words as a plain .txt file — the same text without any layout, which is the more convenient form to paste into a document or a spreadsheet.

OCR PDF without uploading — FAQ

What does OCR do to my PDF?

OCR (Optical Character Recognition) reads the text in your scanned pages and adds an invisible text layer on top of the original image. The page looks exactly the same, but you can now search the document, select text, and copy it into other apps.

Is it safe to OCR confidential documents online?

Yes — your PDF is never uploaded. The recognition engine (Tesseract) runs as WebAssembly inside your browser, so contracts, medical records, and other sensitive scans never leave your device.

Which languages are supported?

English and Indonesian, and you can select both at once for mixed-language documents. Each language model is downloaded to your browser once and reused afterwards.

Why is the recognized text not perfect?

Accuracy depends on scan quality. Clean, straight, 200+ DPI scans typically recognize very well; skewed pages, low-resolution photos, and handwriting are harder. For best results, scan documents flat and in good lighting.