Convert JPG, PNG and WEBP photos, scans or screenshots into editable text. Choose the document language and layout, then review the result before using it.
📂 Upload Image
🔤
Drop an image here or click to browse
JPG, PNG or WEBP · one image · up to 30 MB and 40 megapixels
Choose Image
One image is processed at a time. Original mode sends the selected image pixels straight to the local OCR worker without silent resizing. For a mixed English/Urdu or English/Arabic page, choose the matching combined option.
No account or subscription is required.
🤖 Recognizing Text
Loading OCR engine…
The first run downloads OCR code and the selected language data. Later runs may be faster when browser caching is available.
✅ Extracted Text
Source Image
Extracted Text
OCR confidence estimate
—
Words
—
Characters
—
How can I extract text from an image? Add one JPG, PNG or WEBP image, select its language and layout, then choose Extract Text. Edit the OCR result if needed before copying or downloading it. Recognition quality depends on sharpness, page structure, script and the selected model, so important text should always be checked against the source.
PDFdukan uses Tesseract.js, a JavaScript/WebAssembly port of the open-source Tesseract OCR engine, to recognise printed text in photos, screenshots and scans. The selector includes 15 individual language models plus English–Urdu and English–Arabic mixed modes. Recognition runs in a browser worker; OCR code and language data are downloaded when needed from configured third-party CDNs.
🌍
Urdu, Arabic & Mixed Text
Choose one of 15 single-language models or a combined English–Urdu or English–Arabic mode for bilingual pages.
🎯
Quality Controls
Keep original pixels or try grayscale/document contrast, then match the page layout for a more suitable OCR pass.
📋
Edit, Copy & Export
Correct recognised text in the result box, copy it, or download a UTF-8 TXT file that preserves Urdu and Arabic characters.
How to Extract Text from an Image in 3 Steps
Choose one image — add a JPG, PNG or WEBP photo, screenshot, receipt or scanned page. The tool does not accept PDF files on this page.
Match the language and layout — choose the primary language or a listed mixed mode. Keep “Automatic page” unless the image is one column, a single block, sparse text or one line.
Extract, compare and correct — processing time depends on image size, device and whether code/model files are cached. Edit the result before copying or downloading it.
Common Uses for OCR
Digitising printed documents: Convert photographed letters, certificates, contracts, and reports into editable, searchable text instead of retyping.
Extracting text from screenshots: Grab text from a screenshot of a webpage, chat, error message, or app where copying is disabled.
Book and printed notes: Students can extract text from photographed textbook pages or clearly printed notes, subject to copyright and their permitted use.
Receipts and invoices: Pull amounts, dates, and details from photographed receipts for record-keeping and expense reports.
Old archives: Convert old printed family documents, certificates, and records into digital text before they fade.
Translating foreign text: Extract text from an image first, then paste it into a translator — useful for documents, signs, and labels.
Tips for the Best OCR Accuracy
Use a sharp, well-lit photo — blur is the biggest cause of recognition errors.
Hold the camera straight above the text — tilted text reduces accuracy. For paper documents, scan with CamMaster Scanner first to flatten and enhance the page.
Select the correct language before running OCR — the engine loads language-specific recognition models.
Keep enough detail for readable characters. Tesseract's own guidance says images around 300 DPI often work well; avoid heavily compressed forwards when the original is available.
Try Document contrast for faded paper or uneven lighting, but compare it with Original mode because preprocessing does not improve every image.
Use Single column for a simple vertical article and Sparse text for labels or screenshots with separated text.
Printed text recognises far better than handwriting — Tesseract is built for printed characters.
Frequently Asked Questions
There is no honest fixed accuracy percentage for every image. Clear, upright printed text normally works better than blur, handwriting, tables, decorative fonts or low-contrast scans. Select the correct language and compare names, numbers and other important text with the original.
Yes. Urdu and Arabic are included among the 15 single-language choices. English + Urdu and English + Arabic mixed modes are also available for bilingual pages. The selected model data may download on first use. Review right-to-left text, names, numbers and punctuation against the original.
Tesseract is designed for printed text and performs poorly on cursive or rough handwriting. Very neat, print-style handwriting may partially work. For handwritten notes, expect to correct errors manually.
The OCR workflow processes the selected image in your browser and does not intentionally upload its pixels. The browser does contact third-party CDNs to download Tesseract.js code and selected language data, so those providers receive normal request information such as an IP address. Avoid processing material you are not authorised to handle.
Convert the PDF pages to images first with the PDF to JPG tool, then run OCR on each page image. Alternatively, use the Searchable PDF tool to embed recognised text directly back into the PDF.
On first use, the browser downloads the OCR engine and selected language model data. File sizes vary by model. If the browser keeps those assets in cache, later extractions may start faster.
The listed OCR workflow is currently available without a subscription or account. Practical use is limited by browser memory, language-data downloads and device performance.
Welcome back
Connect Google Drive to save selected outputs to your own Drive storage. Review the destination before uploading sensitive documents.