The tokenizer shipped first, before any language-model weights. Multilingual tokenizers trained on Latin-heavy crawls routinely shatter Assamese into byte-level noise; this vocabulary was fit on our scanned corpus so words stay whole enough to train on. Eri and Muga both sit on it.

Fertility (tokens per word) is the blunt metric we watch: lower means denser context windows and cleaner morphological signal.

Tokenizer fertility on Assamese text Mean tokens per whitespace word · Assamese literary sample · November 2024
Navdyut 32K 1.42
mBERT 2.31
XLM-R 2.18
IndicBERT 1.87

Deep model documentation stays on the Models page. This post is the release record.