zerouploads

PDF to Text Extractor

Pull the text out of a PDF. Pages with a real text layer come out exactly; scanned pages go through OCR, in your browser.

  • Unlimited
  • No signup
  • Private
  • Works offline

Extract the text from a PDF

No PDF yetdrop, paste or choose a PDF to start

Choose a PDF file

Pages—

none selected

Extracted text

The text appears here, page by page.

——

Extraction settings
The text appears once a PDF is read

How it works

  1. 01

    Add a PDF

    Drop or choose one document. The tool reads the text of every page as it opens, and nothing is uploaded.

  2. 02

    Pick your pages

    Click pages in the grid, or type a range like 1-3, 12, 44-48. The grid marks scanned pages.

  3. 03

    Copy or save

    Take it as plain text, Markdown with headings, or JSON with a record per page.

Two kinds of PDF

A PDF saved from a word processor or a design tool carries a real text layer. The characters are in the file with their positions, so pulling them out is exact and fast. Every character comes out as it went in.

A PDF made by scanning holds only pictures of pages. There is no text in it until OCR reads the pictures. OCR makes a good guess, not a copy, so check the result.

Some documents are both, such as a report with three signed pages or a contract with a photographed annexe. The page grid marks which pages are scanned before you extract.

Exact extraction
Where a text layer exists, what comes out is exactly what went in.
OCR fallback
Your own machine reads scanned pages, with a confidence score for each one.
Page ranges
Take the whole file or a set of pages, such as 1-3, 12, 44-48. Click the grid, or type the range.
Output
Plain text, Markdown with headings, or JSON that says which page each piece came from.

Reading order and layout

A PDF stores where each character sits, not what it belongs to. There are no paragraphs in the file, and often no spaces either. Reading order works both out from the gaps between the characters.

That works for prose and fails for tables, because a table depends on where things sit. Keep the layout pads every line with spaces, so the columns stay under each other and a wide table stays readable.

Raw order is the last resort. It hands back the characters in the order the file stores them, which can untangle a page the other two cannot.

Frequently asked questions

Is my PDF uploaded to get the text out?
No. pdf.js opens the document in this tab, and OCR runs on your own processor. Contracts, reports and medical records never leave your device.
Why did I get no text from my PDF?
Because the pages are pictures, and a scan has no text until OCR reads it. Leave OCR scanned pages on, then press the button at the top that reads them with OCR. The tool reads the pages one at a time.
Is the extraction exact?
On a page with a text layer, yes. The characters come out as they were stored. On a scanned page OCR makes a good guess. The badge over the text says which you are looking at.
Can I extract only some pages?
Yes. Click pages in the grid, or type a range such as 1-3, 12, 44-48 in the Pages field. The grid and the field stay in step, so a choice in one shows in the other.
Does it keep tables intact?
Set Layout to Keep the layout. It pads every line with spaces so the columns stay under each other. Reading order is better for prose, and it turns a table into one long line.
Can it handle a 1,000-page file?
The text pass has no page limit, though a big file takes longer. Thumbnails are drawn only for the first 60 pages. OCR is slow, so read a long scan a few pages at a time.
What about a password-protected PDF?
It cannot be opened here. Remove the password in the app that made it, then try again. The tool tells you the file is password-protected.
Which languages can it recognise?
English, for the scanned pages. A page with a real text layer comes out in whatever language it was written in, because those characters are copied and not recognised.
Browse all PDF tools