错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluation of Ancient Chinese Natural Language Understanding in Large Language Models Based on ACHNLU

  • Die Hu,
  • Guangyao Sun,
  • Liu Liu,
  • Chang Liu,
  • Dongbo Wang

摘要

The remarkable performance of large language models (LLMs) has garnered widespread attention across multiple research domains. The field of ancient Chinese information processing also requires the incorporation of cutting-edge technologies to meet the substantial demands for data processing. To facilitate the application of large language models in the context of ancient Chinese text processing, this study introduces the Ancient Chinese Natural Language Understanding (ACHNLU) benchmark. The benchmark includes a comprehensive evaluation mechanism for assessing the performance of LLMs in 6 typical tasks related to natural language understanding of ancient Chinese. With its high level of comprehensiveness and granularity, the system also demonstrates good scalability, offering a solid evaluation framework for related research. Based on this benchmark, the study evaluates the ancient Chinese understanding capabilities of 13 mainstream LLMs. It was found that while LLMs perform well in tasks such as sentence segmentation and punctuation, they still face significant challenges in word segmentation, part-of-speech tagging, and named entity recognition for ancient Chinese. Among all evaluated models, GPT-4 outperformed the rest. In the realm of open-source LLMs, Ziya-LLaMA-13B-v1.1 and Baichuan-13B performed relatively well, although they still lag behind highly optimized, closed-source LLMs. The LLMs evaluation based on ACHNLU provides empirical support for model selection in the domain of ancient Chinese information processing and offers valuable insights for subsequent research and practical applications.