OCR — extract text from scanned PDFs

Recognise the text in scanned PDFs, photos of documents and screenshots, and get it as editable text. OCR runs on your device with the Tesseract engine — your files are never uploaded.

Drop your scan here

or drag and drop · PDF, JPG, PNG, WEBP · up to 100 MB

    Options

    Working…

    First use downloads the conversion engine (≈ 5 MB). It is then cached by your browser.

    Done!

      Continue with

      Processed on your device. Nothing is uploaded. · Free · No sign-up · No watermark

      How to extract text from a scanned PDF

      1. 1

        Click Choose file or drag a scanned PDF or an image (JPG, PNG, WebP) into the box above.

      2. 2

        Choose the Document language — English, French, Spanish, German, Portuguese or Italian.

      3. 3

        Click Convert. The first time, the OCR engine and language data are downloaded and cached by your browser.

      4. 4

        When recognition is finished, the text file downloads automatically.

      Scanned PDFs need OCR

      When you scan paper or photograph a page, the result is an image — even if it is saved as a PDF. You cannot search it, copy from it or convert it to Word. OCR analyses the shapes of the letters and rebuilds the text.

      TurboConvert uses Tesseract, a widely used open-source OCR engine, compiled to run in your browser. Choosing the right language matters: it lets the engine recognise accented characters (é, ñ, ü, ç…) and common words of that language.

      Get the best recognition

      FactorGoodProblematic
      Resolution300 dpi scans, sharp phone photosSmall, compressed or blurry images
      AlignmentStraight pagesSkewed or rotated pages
      ContrastBlack text on white paperFaded print, coloured backgrounds, shadows
      ContentPrinted body textHandwriting, decorative fonts, text over images

      Practical tips:

      • Rotate sideways scans first with Rotate PDF; OCR works best on upright pages.
      • Photographing a document? Hold the phone parallel to the page, in good light, and fill the frame with the text.
      • Only need a few pages? Extract them with Split PDF to save time on long documents.
      • Choose the document’s language, not your own: an English letter scanned in Paris should still be read with English.

      What next?

      The recognised text is ready to paste into an email or a document. If you need a formatted document, paste it into Word or Google Docs. If your PDF turns out to have real text after all, PDF to Word keeps the full layout and formatting.

      OCR, PDF to Text or PDF to Word?

      Your fileBest toolWhy
      Scanned PDF, photo of a page, screenshotOCR PDFThe text only exists as pixels
      Digital PDF, you need the words onlyPDF to TextExact text, instant, no recognition errors
      Digital PDF, you want to edit with formattingPDF to WordKeeps headings, styles and images

      Common problems and fixes

      The result is full of nonsense characters. The page is probably rotated, very low resolution, or the wrong language is selected. Straighten the page, use a sharper scan and pick the document’s language.

      Some words are wrong. Check for look-alike characters — “l” and “1”, “O” and “0”, “rn” and “m” are classic OCR confusions, especially in small print. A quick proofread fixes them.

      Columns are mixed together. On multi-column pages such as newspapers, lines from neighbouring columns can be joined. Crop each column into its own image before running OCR.

      It is taking a long time. Every page is analysed individually. Process the pages you actually need, and prefer a computer for documents of more than a few dozen pages.

      Questions & answers

      What is OCR?
      OCR (optical character recognition) turns a picture of text — a scan, a photo or a screenshot — into real text you can copy, search and edit. Without it, a scanned PDF is just a set of images.
      How do I know if my PDF needs OCR?
      Try to select a word in the PDF. If you can't, or the whole page highlights as one block, the PDF is scanned and needs OCR. If you can select words, PDF to Text is faster and exact.
      How accurate is the OCR?
      On clean printed documents, most text is recognised correctly, but no OCR is perfect. Always proofread names, numbers and amounts before relying on them.
      Can it read handwriting?
      Not reliably. The engine is built for printed text; neat block capitals sometimes work, but cursive handwriting usually doesn't.
      Why is it slow the first time?
      Your browser first downloads the OCR engine and the data for your language. They are cached afterwards, so later runs start much faster.
      Are my scans uploaded?
      No. Recognition runs entirely in your browser. Scans of IDs, medical letters or contracts never leave your device.