How it works
- 01
Add a PDF
Drop or choose one document. The tool reads the text of every page as it opens, and nothing is uploaded.
- 02
Pick your pages
Click pages in the grid, or type a range like 1-3, 12, 44-48. The grid marks scanned pages.
- 03
Copy or save
Take it as plain text, Markdown with headings, or JSON with a record per page.
Two kinds of PDF
A PDF saved from a word processor or a design tool carries a real text layer. The characters are in the file with their positions, so pulling them out is exact and fast. Every character comes out as it went in.
A PDF made by scanning holds only pictures of pages. There is no text in it until OCR reads the pictures. OCR makes a good guess, not a copy, so check the result.
Some documents are both, such as a report with three signed pages or a contract with a photographed annexe. The page grid marks which pages are scanned before you extract.
- Exact extraction
- Where a text layer exists, what comes out is exactly what went in.
- OCR fallback
- Your own machine reads scanned pages, with a confidence score for each one.
- Page ranges
- Take the whole file or a set of pages, such as 1-3, 12, 44-48. Click the grid, or type the range.
- Output
- Plain text, Markdown with headings, or JSON that says which page each piece came from.
Reading order and layout
A PDF stores where each character sits, not what it belongs to. There are no paragraphs in the file, and often no spaces either. Reading order works both out from the gaps between the characters.
That works for prose and fails for tables, because a table depends on where things sit. Keep the layout pads every line with spaces, so the columns stay under each other and a wide table stays readable.
Raw order is the last resort. It hands back the characters in the order the file stores them, which can untangle a page the other two cannot.