arXiv 2605.31483
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
By Shefayat E Shams Adib, Ahmed Alfey Sani, et al.
Published 2026-05-29
Discussion
Read the public discussion and references gathered around this paper.
Despite Bengali being the sixth most spoken language in the world, no prior work has systematically evaluated hallucination in large language models (LLMs) for Bengali. We introduce BenHalluEval, a fine-grained hallucination evaluation framework for Bengali covering four tasks: Generative Question Answering (GQA), Bangla-English Code-Mixed QA, Summarization, and Reasoning. We construct 12,000 hallucinated candidates…