XML, EPUB and Indian-language localisation for publishers
JATS and BITS XML, accessible EPUB 3, DITA and ONIX — produced by an in-house Python pipeline with measured accuracy, at an Indian cost base. And the one thing global vendors can't offer: native localisation into 11 Indian languages.
Quotes in USD, GBP or INR · NDA standard · SFTP delivery
<article dtd-version="1.3" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">JDH</journal-id>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.1234/jdh.2026.017</article-id>
<title-group>
<article-title>Dairy herd health records in rural
Andhra Pradesh</article-title>
<trans-title-group xml:lang="te">
<trans-title>గ్రామీణ ఆంధ్రప్రదేశ్లో పాడి పశువుల
ఆరోగ్య రికార్డులు</trans-title>
</trans-title-group>
</title-group>
</article-meta>
</front>
</article>Bilingual title groups, MathML, tables and reference linking — validated against the NISO JATS DTD on every file.
Formats we deliver
JATS 1.2 / 1.3
Journal articles — front, body, back, references tagged and validated against NLM/NISO DTDs.
BITS 2.0
Books and book chapters for platforms like Silverchair, Atypon and PubMed Bookshelf.
EPUB 3 (accessible)
Reflowable and fixed-layout, WCAG 2.2 / EPUB Accessibility 1.1, validated with EPUBCheck and Ace.
DocBook / TEI / DITA
Technical documentation, scholarly editions and structured authoring.
ONIX 3.0
Product metadata feeds for distributors and retailers.
Custom XSD / JSON
Your schema, your platform — we map to it and validate every file.
What publishers send us
From one-off back-list projects to ongoing journal production.
Reference project: 70,000 scanned voter-roll PDFs → a search engine in 3 days
Back-list digitisation
Scanned or print-only titles → structured XML and EPUB. Indian-language and mixed-script back-lists are our speciality.
Born-digital conversion
Word, InDesign, LaTeX and PDF manuscripts → JATS/BITS with reference linking (Crossref, PubMed).
Accessibility remediation
Existing EPUB/PDF libraries brought up to WCAG 2.2 with alt text, reading order and MathML.
Localisation into Indian languages
Translate and typeset into Telugu, Hindi, Tamil, Kannada, Malayalam, Marathi, Bengali and more — with native-script QA, not machine output pasted in.
Hosted search
A searchable reader or archive on top of your XML, like the one we built for 10 lakh+ voter records.
Questions from publishers
- Do you work with publishers outside India?
- Yes. We work in English with publishers and typesetting vendors in the US, UK, Europe and Asia-Pacific. Files move over SFTP or your DAM; processing runs on servers in India; quotes are in USD, GBP or INR.
- What is your turnaround and pricing model?
- Per page or per article, with a fixed quote after a free sample. Typical journal article: 24–48 hours. Back-list books: batched per week. Indian cost base, measured accuracy — we send the validation and accuracy report with every delivery.
- How do you ensure quality?
- Automated DTD/schema validation, EPUBCheck and Ace reports, plus human QA on every file in the first batch and on flagged pages after. You see the accuracy numbers before you scale.
- What makes you different from existing XML vendors?
- Two things: we built our own Python OCR-to-XML pipeline (proven on 70,000 scanned PDFs), and we natively handle Indian scripts — so publishers can localise into India's languages through the same vendor that does their JATS.
- Is my content secure?
- Files are processed on our own infrastructure in India, not uploaded to third-party AI services. NDAs as standard; deletion on delivery if you require it.
Send us one article or one chapter
You get back valid XML or EPUB plus the validation report, free, within five working days.