Lipi Print is the print reader behind both archive digitization and the corpus pipeline described on Research. It is trained for Assamese letterforms in books, government gazettes, forms, and newsprint rather than Latin-first multilingual OCR.

The release cut improves on degraded newsprint and multi-column layouts that generic engines routinely scramble.

Character error rate by document class CER % · print eval set · April 2025
Books 1.6%
Gazettes 2.9%
Forms 3.4%
Newsprint 4.7%

Deep model documentation stays on the Models page. This post is the release record.