MindT
Document AI · Perso-Arabic (Nastaliq)

Urdu OCR services

Scanned Urdu books, registers, court files and government records converted into accurate, searchable XML, EPUB or plain text — with a measured accuracy report for every batch.

Used in Telangana, Uttar Pradesh, Bihar, Jammu & Kashmir and the Deccan.

سرکاری دستاویزات اور عدالتی ریکارڈ
اردو · Perso-Arabic (Nastaliq)
10 lakh+
records extracted
70,000+
scanned PDFs processed
3 days
end-to-end, including QA
10+
constituencies indexed

Figures from our 2002 electoral-roll project, run on Telugu and English scans.

Why Urdu OCR is harder than English

Generic OCR tools are built and benchmarked on Latin script. These are the specific things that break them on Urdu — and the things we had to solve to process 70,000 scanned PDFs into 10 lakh+ structured records.

  • Nastaliq is the hardest script we handle

    Words cascade diagonally down and to the left rather than sitting on a straight baseline. OCR built for horizontal text fails immediately.

  • Letters change shape by position

    The same letter has initial, medial, final and isolated forms, and joins its neighbours. Segmenting a word into letters is itself a modelling problem.

  • Dots carry the meaning

    Several letters differ only by the number and position of dots. Faded scans and photocopy speckle destroy exactly that information.

  • Nastaliq versus Naskh

    Older Deccan and Nizam-era printing uses Nastaliq; newer material often uses Naskh. They need different handling, and a single archive frequently contains both.

Urdu documents we handle

  • Nizam-era administrative and revenue records
  • Court records and legal documents
  • Urdu newspapers and literary archives
  • Madrasa and educational records
  • Land and property deeds

Typically for Hyderabad archives, Urdu academies, courts, libraries and research institutions.

How a Urdu batch runs

  1. 1
    Ingest

    Bulk PDF intake from any source — scans, photos, government portals. Page splitting, de-skew, de-noise.

  2. 2
    OCR in the right script

    Models for Devanagari, Telugu, Tamil, Kannada, Malayalam, Bengali, Gujarati, Gurmukhi, Odia, Urdu and Latin scripts, tuned on Indian government and publishing layouts: tables, multi-column, mixed scripts.

  3. 3
    Structure to XML

    Fields, tables and hierarchy mapped to your schema — JATS, TEI, DocBook, ONIX, or a custom XSD. Validated on every file.

  4. 4
    Quality checks

    Automated confidence scoring plus human review of low-confidence pages. You get an accuracy report, not a promise.

  5. 5
    Deliver & search

    XML/EPUB/JSON/CSV delivery, or a ready-made search interface like the one we built for 2002 voter rolls.

Urdu OCR questions

How accurate is Urdu OCR?
It depends far more on scan quality than on the language. On clean modern print we reach high accuracy; on twenty-year-old photocopies the honest answer is that some pages need human review. We run a free 50-page sample of your own documents first and send you the measured accuracy before you commit to anything.
Can you read handwritten Urdu?
Handwriting is a separate and harder problem than print, and results vary with the hand. We assess it on your sample and tell you plainly whether it is viable, rather than promising and disappointing you later.
What formats do I get back?
Validated XML in your schema, EPUB 3, JSON, CSV or plain Unicode text — plus an accuracy report per batch. You own all output.
Do you keep the table and page structure?
Yes. For records like registers and land documents the structure matters as much as the text, so we map fields, rows and columns into your schema rather than dumping a wall of characters.
Is my Urdu data sent outside India?
No. Processing runs on our own Python pipeline on servers in India, and we do not upload your documents to third-party AI services without your written agreement.