Evaluation of Ancient Chinese Natural Language Understanding in Large Language Models Based on ACHNLU
摘要
The remarkable performance of large language models (LLMs) has garnered widespread attention across multiple research domains. The field of ancient Chinese information processing also requires the incorporation of cutting-edge technologies to meet the substantial demands for data processing. To facilitate the application of large language models in the context of ancient Chinese text processing, this study introduces the Ancient Chinese Natural Language Understanding (ACHNLU) benchmark. The benchmark includes a comprehensive evaluation mechanism for assessing the performance of LLMs in 6 typical tasks related to natural language understanding of ancient Chinese. With its high level of comprehensiveness and granularity, the system also demonstrates good scalability, offering a solid evaluation framework for related research. Based on this benchmark, the study evaluates the ancient Chinese understanding capabilities of 13 mainstream LLMs. It was found that while LLMs perform well in tasks such as sentence segmentation and punctuation, they still face significant challenges in word segmentation, part-of-speech tagging, and named entity recognition for ancient Chinese. Among all evaluated models, GPT-4 outperformed the rest. In the realm of open-source LLMs, Ziya-LLaMA-13B-v1.1 and Baichuan-13B performed relatively well, although they still lag behind highly optimized, closed-source LLMs. The LLMs evaluation based on ACHNLU provides empirical support for model selection in the domain of ancient Chinese information processing and offers valuable insights for subsequent research and practical applications.