arXiv 2607.09598

Tokenizer Transplantation: Mitigating Autoregressive Collapse in Edge-Efficient Bengali ASR

By Sanjid Hasan and Md. Abdur Rahman

Published 2026-07-10

Mindmap

Browse the paper's core ideas, clusters, and relationships in a structured outline.

Lightweight speech recognition models are critical for edge deployment, yet highly optimized architectures like Moonshine often fail on morphologically rich, non-Latin languages such as Bengali. This study identifies the root cause of this failure as the model's English-centric byte-level tokenizer, which fragments Bengali words into high-fertility byte chains and triggers catastrophic autoregressive collapse during…

View the original paper on arXiv