Bani is not a fourth speech model. It is the wiring: microphone audio into Bani STT, language through Muga, speech out through Bani TTS, with barge-in and silence handling tuned for Assamese service desks.

We measured full turns on a held-out set of citizen queries recorded across Upper Assam, Lower Assam, and Barak Valley. The numbers below are lab median on a single A100 node; production times depend on the host you control.

Turn latency by stage Median milliseconds on the Bani evaluation set · March 2026
Bani STT 420ms
Muga 890ms
Bani TTS 310ms
Orchestration 180ms

Deep model documentation stays on the Models page. This post is the release record.