This study presents HalluScope, a benchmark specifically developed for evaluating hallucinations in large language models. HalluScope comprises 800 adversarially designed questions spanning multiple domains, systematically categorized into selective, temporal, imitative, factual, and overconfidence hallucinations. The dataset was constructed through automated question generation with mutual supervision between models, enabling both generation and evaluation. The evaluation adopts a multiple-choice format, requiring models to select the correct answers from options containing multiple correct choices, thereby providing a more nuanced assessment of model confidence and judgment under uncertainty. Extensive experiments were conducted on 12 large language models, including ERNIE-Bot, ChatGLM, Qwen, and XVerse, with nine models exhibiting hallucination-free rates below 50%, underscoring the benchmark’s difficulty. Furthermore, HalluScope offers insights into hallucination-prone domains and hallucination types, providing guidance for fine-tuning models to mitigate hallucinations effectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

HalluScope: A Comprehensive Dataset for Evaluating Hallucination in Large Language Models Across Multiple Domains

  • Chen Zhao,
  • BiaoJie Zeng,
  • Kedi Chen,
  • Xin Lin

摘要

This study presents HalluScope, a benchmark specifically developed for evaluating hallucinations in large language models. HalluScope comprises 800 adversarially designed questions spanning multiple domains, systematically categorized into selective, temporal, imitative, factual, and overconfidence hallucinations. The dataset was constructed through automated question generation with mutual supervision between models, enabling both generation and evaluation. The evaluation adopts a multiple-choice format, requiring models to select the correct answers from options containing multiple correct choices, thereby providing a more nuanced assessment of model confidence and judgment under uncertainty. Extensive experiments were conducted on 12 large language models, including ERNIE-Bot, ChatGLM, Qwen, and XVerse, with nine models exhibiting hallucination-free rates below 50%, underscoring the benchmark’s difficulty. Furthermore, HalluScope offers insights into hallucination-prone domains and hallucination types, providing guidance for fine-tuning models to mitigate hallucinations effectively.