Image to Text (OCR)
Pull the text out of a photo, screenshot or scan. Runs entirely on your device — which, for OCR, almost nothing else does.
- Select an image containing text — a scan, a screenshot, or a photo of a page.
- Choose the language. English and Hindi are available, or both at once for mixed documents.
- Click Extract Text. The first run downloads the language model (a few MB, cached afterwards), then reads the image.
The one OCR tool that doesn't take your document
Optical character recognition normally happens on a server, because it needs a real recognition engine. That means every mainstream free OCR site receives a full copy of whatever you feed it — and the documents people run through OCR are rarely trivial. Contracts. Bank statements. Medical letters. ID cards. Handwritten notes.
This runs Tesseract, the long-established open-source OCR engine, compiled to WebAssembly and executed inside your browser tab. The image never leaves your device. Load the page, switch your internet off after the language model has downloaded, and it still works — which is the proof, not a promise.
The trade is speed: a page takes a few seconds rather than being instant, and the first run downloads a few megabytes of language data. After that it is cached.
Getting a result worth having
OCR quality depends almost entirely on the input. In rough order of impact:
- Resolution. Aim for around 300 DPI — roughly 2500 pixels across an A4 page. Below about 600 pixels wide, accuracy falls off sharply, and the tool will warn you.
- Straightness. Even a few degrees of tilt hurts. Your phone's document scanner (Notes on iPhone, Google Drive on Android) corrects perspective automatically and is much better than the plain camera.
- Contrast. Black text on white paper. Shadows across the page, or a photo taken at an angle to a window, confuse the engine more than low resolution does.
- Plain layout. Single-column text reads well. Multi-column layouts, tables and text over images produce jumbled results.
- Printed, not handwritten. Tesseract is built for printed type. Handwriting will mostly fail — that is a limitation of the engine, not of this page.
Hindi and mixed documents
Devanagari is supported alongside English, and the combined option handles documents that mix the two — common in Indian official forms, where headings are in Hindi and details are filled in English. Selecting both languages is slower and slightly less accurate per language than picking one, so choose the single language when a document is genuinely all one script.
Always check the output
No OCR is perfect. The confidence score shown with the result is a useful signal — below about 70% you should expect real errors. The characters most often confused are 0 and O, 1 and l and I, and 5 and S. If you are extracting an account number, a registration number or an amount, read it back against the image character by character before using it anywhere that matters.