How to extract text from a PDF
- 1
Click Choose files or drag one or more PDFs into the box above.
- 2
Click Convert. The text of every page is extracted.
- 3
The .txt file downloads automatically. For several PDFs, download each file or use Download all (ZIP).
Why extract plain text?
Copy-pasting from a PDF viewer is tedious on long documents, and it often breaks lines, mixes up headers and footers or misses pages. Extracting the text in one go gives you a clean .txt file that every program can open. Useful for:
- Quoting or reusing content in an email, report or website.
- Translation — paste into a translation tool without layout getting in the way.
- Search and analysis — find terms across documents, count words, process text with scripts.
- AI assistants and summarisers — many accept plain text more reliably than PDFs.
- Accessibility — plain text works with screen readers and simple e-readers.
Is my PDF digital or scanned?
Open it and try to select a single word with your cursor:
| What happens | Type of PDF | What to use |
|---|---|---|
| You can highlight words and lines | Digital PDF with a text layer | PDF to Text (this tool) |
| The whole page is selected as one block, or nothing at all | Scanned / image PDF | OCR PDF |
Tips
- Need tables as tables? Plain text flattens them. PDF to Excel keeps rows and columns.
- Hyphenated words at line ends come out as they appear in the PDF; a quick find-and-replace of ”-” followed by a line break tidies them up.
- Only need part of a long PDF? Extract the relevant pages with Split PDF first.
Your PDF is read inside your browser and never uploaded, so extracting text from contracts or internal reports is safe.
Common problems and fixes
The .txt file is empty or nearly empty. The PDF has no text layer — typical of scans, photos saved as PDF and some faxes. OCR is the only way to get the text out.
Strange symbols instead of letters. Some PDFs use fonts with a non-standard internal encoding, so the text cannot be decoded correctly even though it looks fine on screen. OCR on the page usually gives readable text in that case.
Words or lines appear in an odd order. A PDF stores text in the order it was drawn, which on complex layouts — columns, sidebars, captions — is not always the reading order. Paragraphs are all there but may need reordering.
Headers and footers repeat on every page. Running headers, page numbers and footers are real text in the PDF, so they are extracted too. A find-and-replace in your text editor removes them quickly.
Only some pages matter. Long reports produce long text files. Extract the pages you need with Split PDF first, then convert just those.