needfiles
PDFs

PDF to text

Extract the text from a PDF into a plain text file.

Your files stay on your device. Nothing is uploaded. How to check

Extract text from a PDF

The PDF to text tool reads the text layer of a PDF and saves it as a plain .txt file. It uses pdf.js, the PDF engine built into Firefox, running inside the browser tab, so the document is never uploaded. That matters for the kind of files people usually extract text from: contracts, bank statements, medical records, and internal reports.

Drop one PDF or several; each becomes its own job. The output keeps the base name and swaps the extension, so "report.pdf" becomes "report.txt".

What a text layer is, and why scans come out empty

A PDF made from a word processor, a web page, or any program that exported it directly contains the actual characters, positioned on the page. pdf.js reads those characters and writes them out. This is the text layer, and it is what you select when you drag the cursor across a PDF viewer.

A PDF made by scanning paper, or by photographing a page with a phone, contains only an image of the text. There are no characters to extract, and this tool does not run OCR, so a scanned PDF produces an empty or nearly empty text file. The quick test is to open the PDF in any viewer and try to select a word. If nothing highlights, there is no text layer. Some scanners add an invisible OCR layer on top of the image; those PDFs do work here, with whatever accuracy the scanner's OCR had.

The page separator option

Page separator is on by default. It writes a line between the text of each page so you can tell where one page ends and the next begins, which helps when you need to cite a page number or check that nothing was skipped.

Turn it off to get one continuous stream, which is better when the text will be pasted into another document, fed to a search index, or processed by a script that should not see page markers.

What to expect from the output

Reading order follows the order the PDF stores its text in. That is usually correct for single-column documents and often jumbled for multi-column layouts, tables, and pages with sidebars or footnotes. Line breaks follow the PDF's own text runs, so a paragraph may arrive split into several lines, and words hyphenated at a line end stay hyphenated.

Fonts, bold, italics, headings, and links are gone, since .txt has no way to represent them. Form field values and comments are not part of the page text and are not included.

For a long document, extracting is a fast way to search the whole thing, paste it into a note, or feed it to a script. For anything that needs the layout kept, you want the pages as images instead.

If you need the pages as pictures, for example to drop a page into a slide deck or to keep a scanned page legible, PDF to images renders each page as a PNG or JPG. To pull out only part of a long document before extracting, split PDF can cut it to a page range first. To combine several PDFs into one before extracting a single text file, use merge PDF.

Frequently asked questions

Why is the text file empty?

The PDF is probably a scan. It holds a picture of the text rather than the characters themselves, and this tool does not run OCR, so there is nothing to extract.

How can I tell if a PDF has a text layer?

Open it in any viewer and try to select a word. If text highlights, the layer is there and this tool will read it.

What does the page separator option do?

It writes a divider line between the text of each page. Turn it off to get one continuous stream of text.

Is formatting kept?

No. The output is plain text, so fonts, bold, headings, and links are gone. Tables and multi-column layouts may come out in a jumbled order.

Can I extract text from several PDFs at once?

Yes. Drop them together and each one becomes its own job with its own text file.

What is the output file called?

The same base name with a .txt extension, so 'report.pdf' becomes 'report.txt'.

More tools