The tokenizer shipped first, before any language-model weights. Multilingual tokenizers trained on Latin-heavy crawls routinely shatter Assamese into byte-level noise; this vocabulary was fit on our scanned corpus so words stay whole enough to train on. Eri and Muga both sit on it.
Fertility (tokens per word) is the blunt metric we watch: lower means denser context windows and cleaner morphological signal.
Deep model documentation stays on the Models page. This post is the release record.