OCR for Indian Languages,
Right in Your Browser
Extract text from PDFs and images in Hindi, Tamil, Telugu, Bengali and 9 more scripts — entirely on your device. No uploads. No cloud. Complete privacy.
🔒 No sign-up required · Works offline · MIT licensed
How it works
From document to text
in a few clicks
No accounts, no configuration. Drop your file and get clean, searchable text.
Drop or browse your file
Upload a PDF, PNG, JPEG, TIFF, or WebP. The file never leaves your browser — everything runs locally.
Auto-detect language & OCR
The engine identifies the script in milliseconds. Choose on-device WASM or AI-accelerated mode for ~1 s/page.
Review inline
Bounding-box overlays show confidence per word. Edit results directly in the studio before exporting.
Export in your format
Download searchable PDF, editable DOCX, plain TXT, or a rich JSON audit file with word-level coordinates.
Features
Everything you need,
nothing you don't
100% Client-Side Privacy
Zero bytes of your document leave the device. Image preprocessing, deskewing, binarization, and OCR all run inside your browser via WebAssembly.
Zero data transmissionDual-Engine Acceleration
On-Device WASM for full offline operation, or Gemini AI Fast Mode for ~1 second per page when speed matters. Switch instantly.
Tesseract.js · Gemini AIInstant Language Auto-Detection
Tier 1 detects digital PDFs in 2 ms via Unicode classification. Tier 2 handles scanned pages in ~300 ms via AI multimodal analysis.
13 Indian scriptsMulti-Format Export
Searchable PDF with invisible text overlay, Word DOCX with structure, plain UTF-8 TXT, and a full audit JSON with word coordinates.
PDF · DOCX · TXT · JSONPWA & Offline Support
Installs as a native-like app on any device. Service Worker cache-first architecture means it works without internet after first load.
Works offlineConfidence Bounding Boxes
Every recognized word is overlaid with a color-coded box — green (high), amber (medium), red (low) — so you instantly know what to review.
Word-level confidenceOCR Engines
Choose your processing mode
Both engines produce verbatim transcription — zero translation, zero transliteration.
Tesseract.js v5
Runs entirely inside your browser via WebAssembly. No API key. Fully offline once loaded.
- Absolute privacy — zero network requests
- Works offline after first visit
- No API key or account required
- Supports all 13 Indian scripts
Google Gemini Flash / Pro
Multimodal AI transcription via the Gemini API. 8× faster with superior accuracy on poor scans.
- ~1 second per page throughput
- Superior accuracy on degraded scans
- Handles complex mixed-script documents
- Free tier available with a Gemini API key
Language Support
13 scripts, fully supported
Strict verbatim transcription — the exact script of the original document, always preserved.
Privacy First
Your documents stay
on your device
In on-device mode, OCR Studio processes everything inside your browser's sandbox. No telemetry, no analytics, no cloud storage. There is no server that could receive your files.
In Gemini AI mode, only the page image is sent to the Gemini API using your own API key.
in on-device mode
supported
Gemini AI mode
available
Ready to extract text
from your documents?
No account. No upload. Works in any modern browser. Free and open source.
Drop your document here
Supports PDF, PNG, JPEG, TIFF, WebP · 100% on-device processing
Pre-flight Analysis
Text will appear here as pages are processed.
Click any block to highlight its scan region.
Proposed Corrections
No corrections proposed for this page.