Blog Tools 6 min read

Document Scanner Online — Extract Text from PDFs and Scanned Pages

You have a PDF. You need the text out of it. Maybe it's a contract you need to quote. Maybe it's a scanned form you need to re-enter. Maybe it's a 40-page report and you need three paragraphs. Here's how to get it.

Document scanner — extract text from PDFs and scanned document images

PDFs are designed for viewing, not editing. That's the whole point of the format — the document looks the same on every device, and nobody accidentally moves a paragraph. But when you need the actual text — to paste into another document, to search for a phrase, to feed into a spreadsheet — the format works against you. You can't just Ctrl+A, Ctrl+C from every PDF.

Try it free: Document Scanner — Extract text from documents, IDs, and forms. Runs in your browser, no signup needed.

The Document Scanner extracts text from PDF files in the browser. Upload a PDF, get the text from every page, copy it. For PDFs that are actually scanned images (photos of pages rather than digital text), the tool applies OCR to read the visible characters. Everything runs client-side — the file never leaves your device.

Two Kinds of PDFs

The first thing to understand is that not all PDFs contain text the same way. This determines how extraction works — and how well it works.

Digital PDFs contain actual text data. They were created from a word processor, a web page, a design tool, or any software that generates text-based output. The characters are stored as Unicode text with font and position data. Extraction from these PDFs is fast, accurate, and produces clean output — the tool reads the text layer directly.

Image-based PDFs contain page images. They were created by a physical scanner, a fax machine, a camera capture, or a "print to PDF" from a system that rasterized the text. The pages look like they contain text, but technically they're photographs of text. There's no text layer to extract — the tool needs OCR (Optical Character Recognition) to read the characters from the image pixels.

The Document Scanner detects which kind you have. If it finds an embedded text layer, it extracts directly. If not, it falls back to OCR. Mixed PDFs (some digital pages, some scanned) are handled page by page.

Getting Clean Extraction from Digital PDFs

For digital PDFs, extraction is straightforward — the text comes out in reading order, page by page. A few things to know about the output.

Formatting is stripped. Bold, italic, font sizes, colors, and visual layout don't carry over. You get plain text. If the original had a three-column layout, the text appears as one continuous stream in the reading order the PDF defined. This is usually left-to-right, top-to-bottom, but complex layouts (magazines, brochures, multi-column academic papers) can produce jumbled reading order.

Tables are tricky. PDF tables aren't really tables — they're text boxes positioned to look like a grid. The extracted text may have tab characters or spaces between cells, but the structure isn't guaranteed. For data-heavy tables, you might get better results converting the PDF page to an image with the PDF to Image tool, then using the Text Scanner for OCR — it sometimes preserves tabular structure better because it reads the visual layout.

Headers and footers repeat. Every page's header, footer, and page number appear in the output because they're part of the text layer. If you're extracting a long document, you'll need to clean these out of the result manually.

Extract text from PDF files and scanned pages. OCR support for image-based PDFs. Multi-page processing.

Scan a Document →

Dealing with Scanned PDFs (OCR)

Image-based PDFs are harder to extract from because OCR has to interpret pixel patterns as characters. The accuracy depends almost entirely on the scan quality.

High-resolution scans (300 DPI or above) produce good OCR results. Most characters are recognized correctly. Minor errors might appear with similar-looking characters: lowercase L and uppercase I, zero and O, 5 and S in certain fonts.

Low-resolution scans (under 200 DPI) produce noisy results. Characters blur together, thin strokes disappear, and the error rate climbs. If you're scanning a physical document yourself, scan at 300 DPI minimum. The DPI Checker can verify the resolution of your scan.

Skewed or rotated pages reduce accuracy because character shapes distort. If your scanned pages are slightly rotated (common with sheet-fed scanners), straighten them before OCR. For heavily skewed documents, use the Image Cropper to straighten and crop individual pages before processing.

Handwritten text is generally not readable by standard OCR. The Document Scanner works with printed text — machine-generated characters with consistent shapes. Handwriting recognition requires specialized AI models that this tool doesn't include.

Document Scanner vs. Text Scanner

Both tools extract text, but they're designed for different inputs.

The Document Scanner takes PDF files. It reads the embedded text layer first (fast, accurate) and falls back to OCR only when needed. It handles multi-page documents natively — upload a 50-page PDF and get text from all 50 pages. Use it when you have a PDF.

The Text Scanner takes images — photos, screenshots, scanned pages as image files. It always uses OCR because there's no text layer in an image. It's optimized for single images: a photo of a whiteboard, a screenshot of an error message, a picture of a book page. Use it when you have a photo or screenshot.

If you have a scanned PDF and the Document Scanner's OCR quality isn't good enough, try this workflow: convert the PDF pages to high-resolution images with the PDF to Image tool, then run each page image through the Text Scanner. The dedicated image OCR pipeline sometimes handles challenging scans better than the PDF OCR fallback.

What to Do with the Extracted Text

Copy to clipboard. The most common use — grab the text and paste it into a document, email, form, or editor. The tool provides a copy button for the entire output or individual pages.

Search and find. Need a specific phrase from a long document? Extract the text, then use Ctrl+F to find it. This is faster than scrolling through a 100-page PDF looking for one sentence.

Feed into other tools. Extracted text can be pasted into translation tools, grammar checkers, word counters, or AI summarizers. If the document contains images you also need to extract, the Image Splitter can separate multi-image pages into individual files. If the original PDF contains data you need to analyze, paste the text into a spreadsheet and parse it there.

Accessibility. Image-based PDFs are inaccessible to screen readers. Extracting the text and providing it alongside the PDF makes the content available to people who use assistive technology.

Common Questions

What's the difference between the Document Scanner and the Text Scanner? Document Scanner is for PDFs (reads text layers, handles multi-page). Text Scanner is for images (always uses OCR). If you have a PDF, start with Document Scanner.

Why did the scanner return no text? Your PDF is image-based — the pages are scanned photos with no text layer. The tool attempts OCR, but results depend on scan quality. For stubborn PDFs, convert pages to images first with PDF to Image, then OCR each page with the Text Scanner.

Is the formatting preserved? No. You get plain text in reading order. Bold, italic, fonts, colors, and layout positioning are stripped. Tables may appear as space-separated values.

Is there a page limit? No hard limit. Processing is browser-based, so very large PDFs may be slow. Close other tabs to free up memory for large documents.

Text Locked in a PDF Is Still Just Text

A PDF is a box. The text inside it belongs to you — you just need a way to get it out. The Document Scanner reads the text layer when there is one and applies OCR when there isn't. Upload, extract, copy. The document stays on your machine, the text goes where you need it.

Scan a Document
Share:
S

Scanly.co — 92 free image analysis tools

Photo forensics, metadata, privacy, OCR, and utilities. All client-side.

Advertisement