arXiv 2607.09598
Tokenizer Transplantation: Mitigating Autoregressive Collapse in Edge-Efficient Bengali ASR
By Sanjid Hasan and Md. Abdur Rahman
Published 2026-07-10
Mindmap
Browse the paper's core ideas, clusters, and relationships in a structured outline.
Lightweight speech recognition models are critical for edge deployment, yet highly optimized architectures like Moonshine often fail on morphologically rich, non-Latin languages such as Bengali. This study identifies the root cause of this failure as the model's English-centric byte-level tokenizer, which fragments Bengali words into high-fertility byte chains and triggers catastrophic autoregressive collapse during…