How to extract text from a scanned PDF
- 1
Click Choose file or drag a scanned PDF or an image (JPG, PNG, WebP) into the box above.
- 2
Choose the Document language — English, French, Spanish, German, Portuguese or Italian.
- 3
Click Convert. The first time, the OCR engine and language data are downloaded and cached by your browser.
- 4
When recognition is finished, the text file downloads automatically.
Scanned PDFs need OCR
When you scan paper or photograph a page, the result is an image — even if it is saved as a PDF. You cannot search it, copy from it or convert it to Word. OCR analyses the shapes of the letters and rebuilds the text.
TurboConvert uses Tesseract, a widely used open-source OCR engine, compiled to run in your browser. Choosing the right language matters: it lets the engine recognise accented characters (é, ñ, ü, ç…) and common words of that language.
Get the best recognition
| Factor | Good | Problematic |
|---|---|---|
| Resolution | 300 dpi scans, sharp phone photos | Small, compressed or blurry images |
| Alignment | Straight pages | Skewed or rotated pages |
| Contrast | Black text on white paper | Faded print, coloured backgrounds, shadows |
| Content | Printed body text | Handwriting, decorative fonts, text over images |
Practical tips:
- Rotate sideways scans first with Rotate PDF; OCR works best on upright pages.
- Photographing a document? Hold the phone parallel to the page, in good light, and fill the frame with the text.
- Only need a few pages? Extract them with Split PDF to save time on long documents.
- Choose the document’s language, not your own: an English letter scanned in Paris should still be read with English.
What next?
The recognised text is ready to paste into an email or a document. If you need a formatted document, paste it into Word or Google Docs. If your PDF turns out to have real text after all, PDF to Word keeps the full layout and formatting.
OCR, PDF to Text or PDF to Word?
| Your file | Best tool | Why |
|---|---|---|
| Scanned PDF, photo of a page, screenshot | OCR PDF | The text only exists as pixels |
| Digital PDF, you need the words only | PDF to Text | Exact text, instant, no recognition errors |
| Digital PDF, you want to edit with formatting | PDF to Word | Keeps headings, styles and images |
Common problems and fixes
The result is full of nonsense characters. The page is probably rotated, very low resolution, or the wrong language is selected. Straighten the page, use a sharper scan and pick the document’s language.
Some words are wrong. Check for look-alike characters — “l” and “1”, “O” and “0”, “rn” and “m” are classic OCR confusions, especially in small print. A quick proofread fixes them.
Columns are mixed together. On multi-column pages such as newspapers, lines from neighbouring columns can be joined. Crop each column into its own image before running OCR.
It is taking a long time. Every page is analysed individually. Process the pages you actually need, and prefer a computer for documents of more than a few dozen pages.