Image to Text (OCR)

🔒 Runs in your browser — files are never uploaded

Pull the text out of a photo, screenshot or scan. Runs entirely on your device — which, for OCR, almost nothing else does.

How it works
  • Select an image containing text — a scan, a screenshot, or a photo of a page.
  • Choose the language. English and Hindi are available, or both at once for mixed documents.
  • Click Extract Text. The first run downloads the language model (a few MB, cached afterwards), then reads the image.

The one OCR tool that doesn't take your document

Optical character recognition normally happens on a server, because it needs a real recognition engine. That means every mainstream free OCR site receives a full copy of whatever you feed it — and the documents people run through OCR are rarely trivial. Contracts. Bank statements. Medical letters. ID cards. Handwritten notes.

This runs Tesseract, the long-established open-source OCR engine, compiled to WebAssembly and executed inside your browser tab. The image never leaves your device. Load the page, switch your internet off after the language model has downloaded, and it still works — which is the proof, not a promise.

The trade is speed: a page takes a few seconds rather than being instant, and the first run downloads a few megabytes of language data. After that it is cached.

Getting a result worth having

OCR quality depends almost entirely on the input. In rough order of impact:

  • Resolution. Aim for around 300 DPI — roughly 2500 pixels across an A4 page. Below about 600 pixels wide, accuracy falls off sharply, and the tool will warn you.
  • Straightness. Even a few degrees of tilt hurts. Your phone's document scanner (Notes on iPhone, Google Drive on Android) corrects perspective automatically and is much better than the plain camera.
  • Contrast. Black text on white paper. Shadows across the page, or a photo taken at an angle to a window, confuse the engine more than low resolution does.
  • Plain layout. Single-column text reads well. Multi-column layouts, tables and text over images produce jumbled results.
  • Printed, not handwritten. Tesseract is built for printed type. Handwriting will mostly fail — that is a limitation of the engine, not of this page.
Scanned PDF? Convert it first with PDF to JPG, then run the pages through here one at a time. The result is plain text — this does not produce a searchable PDF, which needs a different process.

Hindi and mixed documents

Devanagari is supported alongside English, and the combined option handles documents that mix the two — common in Indian official forms, where headings are in Hindi and details are filled in English. Selecting both languages is slower and slightly less accurate per language than picking one, so choose the single language when a document is genuinely all one script.

Always check the output

No OCR is perfect. The confidence score shown with the result is a useful signal — below about 70% you should expect real errors. The characters most often confused are 0 and O, 1 and l and I, and 5 and S. If you are extracting an account number, a registration number or an amount, read it back against the image character by character before using it anywhere that matters.

Frequently asked questions

Is my image uploaded for OCR?
No, and that is unusual — almost every other free OCR tool uploads. This runs the Tesseract engine compiled to WebAssembly inside your browser. Once the language model has downloaded you can go offline and it still works.
Why does the first run take longer?
It downloads the language model, a few megabytes, once. Your browser caches it, so later runs start immediately.
Does it work with Hindi?
Yes. English, Hindi, or both together for mixed documents. Pick a single language when the document is genuinely all one script — it is faster and slightly more accurate.
Can it read handwriting?
Generally no. Tesseract is built for printed text. Neat block capitals sometimes work, but cursive handwriting will mostly fail.
Why is the extracted text full of errors?
Almost always the input. Aim for about 300 DPI, keep the page straight and evenly lit, and make sure the language matches. A phone document scanner gives far better results than a plain camera photo.
Can I OCR a scanned PDF?
Convert it to images first with the PDF to JPG tool, then run the pages through here. The output is plain text, not a searchable PDF.
How accurate is it?
On a clean 300 DPI scan of printed text, very good. The confidence score is shown with each result — below about 70% expect real errors and check the text against the image.