arXiv 2605.31483

BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali

By Shefayat E Shams Adib, Ahmed Alfey Sani, et al.

Published 2026-05-29

Discussion

Read the public discussion and references gathered around this paper.

Despite Bengali being the sixth most spoken language in the world, no prior work has systematically evaluated hallucination in large language models (LLMs) for Bengali. We introduce BenHalluEval, a fine-grained hallucination evaluation framework for Bengali covering four tasks: Generative Question Answering (GQA), Bangla-English Code-Mixed QA, Summarization, and Reasoning. We construct 12,000 hallucinated candidates…

View the original paper on arXiv