MindT
Document AI

OCR & XML conversion for Indian-language documents

We turn scanned PDFs — in Telugu, Hindi, Tamil, Kannada, Marathi, Bengali, Urdu, English and other Indian languages — into structured, searchable XML, EPUB or JSON. Built in Python, in-house, and proven on over a million records.

Publishers, archives, courts and government departments send us scanned PDFs. They get back validated XML, EPUB or JSON — in Telugu, Hindi or English — with a measured accuracy report.

10 lakh+
records extracted
70,000+
scanned PDFs processed
3 days
end-to-end, including QA
10+
constituencies indexed

Publisher? See JATS, BITS and EPUB services. From one project: 70,000 scanned voter-roll PDFs → a search engine in 3 days

How a batch moves through the pipeline

  1. Step 1

    Ingest

    Bulk PDF intake from any source — scans, photos, government portals. Page splitting, de-skew, de-noise.

  2. Step 2

    OCR in the right script

    Models for Devanagari, Telugu, Tamil, Kannada, Malayalam, Bengali, Gujarati, Gurmukhi, Odia, Urdu and Latin scripts, tuned on Indian government and publishing layouts: tables, multi-column, mixed scripts.

  3. Step 3

    Structure to XML

    Fields, tables and hierarchy mapped to your schema — JATS, TEI, DocBook, ONIX, or a custom XSD. Validated on every file.

  4. Step 4

    Quality checks

    Automated confidence scoring plus human review of low-confidence pages. You get an accuracy report, not a promise.

  5. Step 5

    Deliver & search

    XML/EPUB/JSON/CSV delivery, or a ready-made search interface like the one we built for 2002 voter rolls.

Numbered because the order matters: you cannot get good XML from bad OCR, and you cannot get good OCR without clean page images.

Who uses this

Anyone sitting on paper or scanned PDFs that people need to search. The scripts are the hard part; we've done the hard part.

  • PublishersBack-list books and journals → EPUB 3 and JATS XML, with Indian-language typesetting preserved.
  • Government & archivesDecades of scanned registers, gazettes and rolls → searchable databases in the original language.
  • Courts & law firmsCase files and judgments → indexed, searchable, citation-linked text.
  • Universities & librariesTheses, manuscripts and regional-language collections → digital archives with full-text search.
  • Banks, insurers, hospitalsForms and KYC documents → structured data without manual entry.

Questions

Which languages can you OCR?
All major Indian languages — Hindi, Telugu, Tamil, Kannada, Malayalam, Marathi, Gujarati, Bengali, Odia, Punjabi, Urdu — plus English, including mixed-script pages. We train on a sample of your documents first and show you the accuracy before you commit.
What accuracy do you get?
It depends on scan quality. On the 2002 electoral rolls (poor-quality 20-year-old scans) we reached search-grade accuracy with human review on flagged pages. We report measured accuracy per batch rather than quoting a single number.
How fast can you go?
The voter-roll project processed 70,000+ PDFs into 10 lakh+ structured records in three days. Throughput scales with machines, not people.
What do I receive?
Your documents as validated XML (or EPUB, JSON, CSV), an accuracy report, and optionally a hosted search interface. You own all output.
Do you use ChatGPT or send my documents abroad?
Processing runs on our own Python pipeline on servers in India. We don't upload your documents to third-party AI services unless you ask us to.

Try it on 50 of your pages

Send a sample. We return XML and an accuracy report within a week, free. Then you decide.