PDF to Text

PDF to Text extract selectable text from PDF save PDF as TXT copy text from PDF pages
PDF to Text

Extract text already stored in a PDF

Choose a PDF, optionally enter page numbers, and press Extract text. Read the output before copying or downloading the TXT file. This tool retrieves existing text objects; it does not recognize words inside scanned pictures. Files are processed locally in your browser.

Advertisement

Selecting pages

Leave the page field blank to include every page, up to the processing limit. Use a comma-separated selection such as 1, 3-5 to extract pages one, three, four, and five. Page numbers refer to physical positions in the PDF, starting at one, rather than printed chapter numbers or Roman-numeral labels. Repeated selections are deduplicated and processed in document order.

Worked example: a three-page document

Suppose page one contains a title, page two contains a screenshot, and page three contains a paragraph of selectable text. Enter 1, 3 to extract the title and paragraph with page separators. Selecting all three also adds a notice that page two has no extractable text. If every selected page lacks text, the tool reports that condition instead of presenting a misleading empty conversion as success.

Text extraction versus OCR

A PDF may look like a normal document while actually containing photographs of its pages. Without a text layer, there are no words for this extractor to retrieve. OCR is a separate recognition process and is not included here. A scanned PDF with an existing OCR layer can yield text, but any recognition mistakes already present in that layer may appear in the result.

Reading order and spacing

PDF stores instructions for drawing a page, which do not always follow natural reading order. The extractor uses text sequence, line-ending hints, and positions to add basic spaces and line breaks. Columns, tables, rotated labels, ligatures, unusual fonts, and right-to-left layouts may need cleanup. Output preserves text rather than page design; check names, numbers, and paragraph order against the original before relying on it.

Privacy and limitations

Choose files up to 50 MB and no more than 500 selected pages. Extraction stops if the text exceeds 5 million characters; select a smaller range in that case. Password-required PDFs must be unlocked using the known password first. The tool does not execute PDF scripts, follow document links, or upload the file. Corrupt documents and missing character mappings can prevent complete extraction even when a PDF viewer displays the page.

PDF to Text FAQ

Will this preserve tables?

No. A table becomes plain text and may lose column alignment. Review it against the source or use a dedicated table-extraction workflow.

What does the download contain?

A UTF-8 text file with page separators and the extracted content. The filename keeps the source basename and adds -text.

Can I extract images too?

This page extracts text only. Use PDF to PNG or PDF to JPG when you need rendered page images.

Related tools and reference

Categories

Tags