<p>Large language modeling has permeated various fields of society and considerably impacted the humanities and social sciences. The emergence of domain-specific large language models (LLMs) has opened new avenues for research in these fields. For instance, in the area of intangible cultural heritage (ICH), in which strong regional and ethnic roots and existing transmission methods hinder its transmission in modern society, domain-specific LLMs provide a digital approach to preserving and transmitting ICH-related knowledge. As such, a comprehensive model evaluation benchmark is essential to develop an effective ICH-oriented LLM. This study develops a domain-specific LLM evaluation benchmark for the ICH domain and evaluates the performance of current mainstream LLMs utilizing relevant data sourced from the China ICH network. The results demonstrated that the linguistic capability of LLMs affected their performance, even in vertical domain evaluations.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluation of large language models for the intangible cultural heritage domain

  • Zhixiao Zhao,
  • Dongbo Wang

摘要

Large language modeling has permeated various fields of society and considerably impacted the humanities and social sciences. The emergence of domain-specific large language models (LLMs) has opened new avenues for research in these fields. For instance, in the area of intangible cultural heritage (ICH), in which strong regional and ethnic roots and existing transmission methods hinder its transmission in modern society, domain-specific LLMs provide a digital approach to preserving and transmitting ICH-related knowledge. As such, a comprehensive model evaluation benchmark is essential to develop an effective ICH-oriented LLM. This study develops a domain-specific LLM evaluation benchmark for the ICH domain and evaluates the performance of current mainstream LLMs utilizing relevant data sourced from the China ICH network. The results demonstrated that the linguistic capability of LLMs affected their performance, even in vertical domain evaluations.