Eri is the open proof of process: pretrained from scratch on Navdyut’s scanned Assamese corpus, not adapted from a multilingual checkpoint that barely saw the script. Weights are on Hugging Face for research use.

The curve below is validation perplexity on a held-out Assamese literary and gazette mix during the final pretraining run.

Validation perplexity during pretraining PPL on held-out Assamese mix · January 2025 run
10 20 30 40 50 Training progress %

Deep model documentation stays on the Models page. This post is the release record.