How to turn a scanned PDF or image into editable text
If you can't select the text in a document, it's a picture — not text. OCR fixes that. Here's how it works and how to get the cleanest results.
You have a scanned contract, a photo of a page from a book, or a PDF where you can see the words but can't select them. You need that text as something you can edit, search, or paste. This guide explains what OCR is, why some PDFs behave like pictures, and how to turn those images back into real, editable text — honestly, including where the limits are.
What OCR is, and when you need it
OCR stands for optical character recognition. It is the process of looking at an image of text — a scan, a photo, or a page that is really just a picture — and working out which letters and words are in it, so you end up with text you can copy, edit, and search.
You need OCR whenever the text you want lives inside an image rather than as actual characters. Common cases:
- A document you scanned on a printer or phone and saved as a PDF or image.
- A photo of a page, a receipt, a whiteboard, or a sign.
- A PDF where you can read the words on screen but cannot highlight or select them with your cursor.
In all of these, the words exist only as pixels. OCR is what converts those pixels into text your computer understands.
Text PDF vs scanned (image) PDF — and how to tell
Not every PDF needs OCR. There are really two kinds. A text PDF was created digitally — exported from Word, a browser, or design software — and it stores the actual characters. A scanned or image PDF is a photograph of a page wrapped in a PDF container; it stores pixels, not letters.
Here is a quick test. Open the PDF and try to select a line of text with your mouse, or press Ctrl+F (Cmd+F on a Mac) and search for a word you can clearly see on the page. If you can highlight individual words, or the search finds them, it is a text PDF and you may not need OCR at all — you can just copy the text. If nothing highlights, or the search finds nothing, the page is an image and OCR is the way to get editable text out of it.
How OCR works, in plain terms
You do not need the technical detail to use OCR well, but a rough picture helps. The software first studies the image, finding the regions that look like text and separating them from pictures and blank space. It works out the layout — where the lines and blocks sit. Then it looks at each character shape and compares it against what it knows letters and numbers look like, choosing the best match. Finally, modern OCR checks its guesses against a language model, so if it reads something almost a real word, it can correct an ambiguous letter — telling a lowercase "l" from the number "1", for example.
Because that last step leans on language, telling the tool which language the document is in genuinely improves accuracy.
What affects accuracy
OCR quality depends far more on the input image than on any setting. The things that matter most:
- Resolution. Blurry or low-resolution images give the software less to work with. Scans around 300 DPI, or a sharp, close photo, work best.
- Contrast. Dark text on a clean, light background is ideal. Faded print, grey scans, or coloured backgrounds make characters harder to separate.
- Straightness. Pages that are skewed, curved, or photographed at an angle confuse line detection. Keep the text as level and flat as you can.
- Handwriting. Most OCR is built for printed text. Neat handwriting sometimes works, but expect it to struggle, and cursive especially so. Treat handwritten results as a rough draft.
- Language and characters. Unusual fonts, heavy styling, tables, maths, and mixed languages are all harder. Matching the tool's language to the document helps.
Step by step with the Toollapse Image to Text tool
The Image to Text (OCR) tool runs entirely in your browser. Nothing is uploaded — the recognition happens on your own device — which matters when the document is private.
- Open the Image to Text (OCR) tool and add your image. If your source is a PDF, first convert a PDF page to an image, then feed that image to the OCR tool.
- Choose the language of the document if you are prompted to, so the software checks against the right words.
- Let it process. On device, this can take a few seconds per page depending on your computer and the image size.
- Review the extracted text in the output area.
- Copy the text, or save it, and move it into wherever you need it.
Tips for cleaner results
A minute spent on the image saves several on fixing text. Before you run OCR:
- Scan or photograph in good, even light with no shadow falling across the page.
- Fill the frame with the page and hold the camera parallel to it so the text is not slanted.
- Crop out margins, edges, and anything that is not the text you want.
- Aim for strong contrast — clear dark text on a plain light background.
- Do one page at a time for a document. It keeps the layout tidy and the output easy to check.
Realistic expectations — always proofread
OCR is genuinely good on clean, printed pages, and often you will get text that needs only a light touch. But it is not perfect, and it is important to be honest about that. It can misread similar-looking characters, drop punctuation, scramble columns and tables, and invent odd spacing. The messier the original, the more mistakes you will see.
So treat the output as a strong first draft, not a finished document. Always read it against the original, and pay special attention to numbers, names, dates, and anything where a single wrong character matters — OCR has no way of knowing that an account number or a price is wrong.
What to do next
Once you have the text, it behaves like any other text. You can paste it straight into Word, Google Docs, or an email and format it there. You can drop it into a spreadsheet, a notes app, or a form. Because it is now real text, it is also searchable, which is often the whole point of digitising an old document.
If your goal was specifically an editable Word document from a PDF that already contained selectable text, you may not need OCR at all — a direct PDF to Word conversion will keep more of the original layout. Reach for OCR when the source is genuinely an image; reach for a direct conversion when the text is already there.
Frequently asked questions
Is my document uploaded to a server?
No. The Toollapse Image to Text tool runs the OCR in your browser on your own device, so the file and the recognised text stay with you and are not sent anywhere.
Can OCR read my handwriting?
Sometimes, if it is neat and printed rather than joined-up, but OCR is built mainly for printed text. Expect more errors with handwriting, and check the result carefully.
Why can I see the text but not select it in my PDF?
Because the page is an image — a picture of the text rather than the characters themselves. That is exactly the situation OCR is for. Convert the page to an image, then run it through the OCR tool.
What image quality do I need for good results?
Sharp, well-lit, high-contrast, and straight. A scan around 300 DPI or a clear, close photo works well. Blurry, dim, or skewed images produce more mistakes.
Will the formatting and tables come out perfectly?
Not always. OCR focuses on the words, and complex layouts, columns, and tables can shift or lose structure. You will usually need to tidy up the formatting after pasting the text.