arXiv 2607.09598

Tokenizer Transplantation: Mitigating Autoregressive Collapse in Edge-Efficient Bengali ASR

By Sanjid Hasan and Md. Abdur Rahman

Published 2026-07-10

Wiki summary

Explore the paper's summary, context, and related research on Papiers.

Lightweight speech recognition models are critical for edge deployment, yet highly optimized architectures like Moonshine often fail on morphologically rich, non-Latin languages such as Bengali. This study identifies the root cause of this failure as the model's English-centric byte-level tokenizer, which fragments Bengali words into high-fertility byte chains and triggers catastrophic autoregressive collapse during…

View the original paper on arXiv