70,000 scanned voter-roll PDFs → a search engine in 3 days
The 2002 electoral rolls exist only as low-quality scanned PDFs in Telugu and English. We built a Python OCR-to-XML pipeline that turned them into 10 lakh+ searchable voter records across 10+ constituencies, with sub-second search by name, EPIC ID, door number or booth.
Internal project, published as mindcap.in and checksir.com
The problem
Campaign workers, researchers and families trying to trace 2002 records had to scroll through thousands of scanned pages by hand. The official portal doesn't expose this year. No existing OCR tool handled the Telugu tabular layout reliably.
What we built
- 1Bulk-fetched and normalised 70,000+ PDFs; split into pages, de-skewed, cleaned.
- 2Trained OCR for the Telugu/English mixed roll layout: name, relative's name, house number, age, sex, EPIC ID, part/booth.
- 3Mapped every page into a validated XML schema, then into a search index.
- 4Phonetic matching so 'Ramesh', 'Ramesha' and రమేష్ all find the same record.
- 5Confidence scoring with human review on flagged pages.
Steps are in pipeline order — each depends on the one before.
Result
- 10 lakh+ records, 10+ constituencies, 1,000+ booth parts indexed.
- Three days from raw PDFs to live search, including QA.
- Sub-second search in Telugu or English on any phone.
- Live at mindcap.in (multi-constituency) and checksir.com (Proddatur, 2.14 lakh voters).
Stack
- Python
- Custom OCR models
- XML / XSD validation
- Search index
- Next.js front-end
Everything runs on our own Python pipeline on servers in India. No documents were sent to third-party AI services.
Have scanned documents in Telugu, Hindi or English?
Books, journals, registers, case files, gazettes. Same pipeline, your schema.
mindcap.in and checksir.com are independent MindT projects built from publicly available historical electoral rolls. They are not affiliated with or endorsed by the Election Commission of India.