OCR & XML conversion for Indian-language documents
We turn scanned PDFs — in Telugu, Hindi, Tamil, Kannada, Marathi, Bengali, Urdu, English and other Indian languages — into structured, searchable XML, EPUB or JSON. Built in Python, in-house, and proven on over a million records.
Publishers, archives, courts and government departments send us scanned PDFs. They get back validated XML, EPUB or JSON — in Telugu, Hindi or English — with a measured accuracy report.
Publisher? See JATS, BITS and EPUB services. From one project: 70,000 scanned voter-roll PDFs → a search engine in 3 days
How a batch moves through the pipeline
- Step 1
Ingest
Bulk PDF intake from any source — scans, photos, government portals. Page splitting, de-skew, de-noise.
- Step 2
OCR in the right script
Models for Devanagari, Telugu, Tamil, Kannada, Malayalam, Bengali, Gujarati, Gurmukhi, Odia, Urdu and Latin scripts, tuned on Indian government and publishing layouts: tables, multi-column, mixed scripts.
- Step 3
Structure to XML
Fields, tables and hierarchy mapped to your schema — JATS, TEI, DocBook, ONIX, or a custom XSD. Validated on every file.
- Step 4
Quality checks
Automated confidence scoring plus human review of low-confidence pages. You get an accuracy report, not a promise.
- Step 5
Deliver & search
XML/EPUB/JSON/CSV delivery, or a ready-made search interface like the one we built for 2002 voter rolls.
Numbered because the order matters: you cannot get good XML from bad OCR, and you cannot get good OCR without clean page images.
Who uses this
Anyone sitting on paper or scanned PDFs that people need to search. The scripts are the hard part; we've done the hard part.
- PublishersBack-list books and journals → EPUB 3 and JATS XML, with Indian-language typesetting preserved.
- Government & archivesDecades of scanned registers, gazettes and rolls → searchable databases in the original language.
- Courts & law firmsCase files and judgments → indexed, searchable, citation-linked text.
- Universities & librariesTheses, manuscripts and regional-language collections → digital archives with full-text search.
- Banks, insurers, hospitalsForms and KYC documents → structured data without manual entry.
Questions
- Which languages can you OCR?
- All major Indian languages — Hindi, Telugu, Tamil, Kannada, Malayalam, Marathi, Gujarati, Bengali, Odia, Punjabi, Urdu — plus English, including mixed-script pages. We train on a sample of your documents first and show you the accuracy before you commit.
- What accuracy do you get?
- It depends on scan quality. On the 2002 electoral rolls (poor-quality 20-year-old scans) we reached search-grade accuracy with human review on flagged pages. We report measured accuracy per batch rather than quoting a single number.
- How fast can you go?
- The voter-roll project processed 70,000+ PDFs into 10 lakh+ structured records in three days. Throughput scales with machines, not people.
- What do I receive?
- Your documents as validated XML (or EPUB, JSON, CSV), an accuracy report, and optionally a hosted search interface. You own all output.
- Do you use ChatGPT or send my documents abroad?
- Processing runs on our own Python pipeline on servers in India. We don't upload your documents to third-party AI services unless you ask us to.
OCR by language
Each script breaks generic OCR in a different way. Pick yours to see what we had to solve:
- Telugu · తెలుగు
- Hindi · हिन्दी
- Urdu · اردو
- Tamil · தமிழ்
- Kannada · ಕನ್ನಡ
- Malayalam · മലയാളം
- Marathi · मराठी
- Bengali · বাংলা
- Gujarati · ગુજરાતી
- Punjabi · ਪੰਜਾਬੀ
- Odia · ଓଡ଼ିଆ
Serving every major city
Languages, document types and institutions differ by city. Pick yours:
Try it on 50 of your pages
Send a sample. We return XML and an accuracy report within a week, free. Then you decide.