OCR: extract text from a scanned PDF
Optical character recognition 100% in your browser (tesseract). Extract text from scans and, optionally, generate a searchable PDF.
Your files never leave your device
How to OCR a PDF
- 1
Upload the scanned PDF
And pick the document language.
- 2
Recognition
Every page is processed in your browser with tesseract.
- 3
Copy or download
The text as .txt and, optionally, a searchable PDF.
Why use this tool
100% local
OCR runs in your browser: scans are never uploaded.
Spanish and English
Trained models for both languages, even combined.
Searchable PDF
Optional invisible text layer over the original scan.
Real progress
Page-by-page progress, no fake bars.
Frequently asked questions
How accurate is the OCR?
With sharp scans (300 dpi, printed text) accuracy is high. Handwriting, skewed photos or low resolution reduce it. The engine is tesseract, the reference open-source OCR.
Why is the first run slower?
The WASM engine and the language model (about 15 MB) are downloaded from this same site. They stay cached afterwards.
What is the searchable PDF?
Your original scan untouched plus an invisible text layer aligned with each word: you can Ctrl+F and select text. Alignment is approximate.
Is my document uploaded anywhere?
No. All recognition happens on your device; neither the PDF nor the text leaves your browser.