arXiv 2605.31483

BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali

By Shefayat E Shams Adib, Ahmed Alfey Sani, et al.

Published 2026-05-29

Mindmap

Browse the paper's core ideas, clusters, and relationships in a structured outline.

Despite Bengali being the sixth most spoken language in the world, no prior work has systematically evaluated hallucination in large language models (LLMs) for Bengali. We introduce BenHalluEval, a fine-grained hallucination evaluation framework for Bengali covering four tasks: Generative Question Answering (GQA), Bangla-English Code-Mixed QA, Summarization, and Reasoning. We construct 12,000 hallucinated candidates…

View the original paper on arXiv