VietPS-Hallu: A Vietnamese Dataset for Hallucination Detection in Large Language Models Within the Public Services
摘要
Large Language Models (LLMs), such as ChatGPT, are increasingly explored in public administration for tasks like information retrieval, citizen support, and legal guidance. Yet, a key barrier to their adoption is the phenomenon of hallucination, where models produce content that is factually incorrect or legally misleading. In sensitive domains such as public services, even minor errors can undermine citizen trust, cause procedural failures, or create legal risks. To address this challenge, we introduce VietPS-Hallu, the first benchmark dataset dedicated to hallucination detection in Vietnamese public service scenarios. Unlike existing English-centric and domain-generic resources, VietPS-Hallu is uniquely grounded in authoritative government procedures, ensuring legal traceability and domain specificity. The dataset is constructed through a novel framework that combines large-scale data acquisition from the Vietnamese National Public Service Portal, structured hallucination generation guided by linguistic and factual error patterns, and rigorous human-in-the-loop validation. We further evaluate a range of state-of-the-art LLMs, providing the first systematic insights into their strengths and limitations in low-resource, legally critical environments. By establishing this benchmark, VietPS-Hallu not only pioneers hallucination research for Vietnamese but also lays the groundwork for developing safer, more trustworthy, and policy-compliant AI systems in public administration.