Large language models (LLMs) have demonstrated remarkable capabilities across various domains, fostering growing interest in their applications within the manufacturing industry. This paper presents a comprehensive evaluation of LLMs across 22 subcategories spanning 8 major sectors in the manufacturing industry, including machinery, automobile, electronics, chemical, light industry, pharmaceutical, transportation, and food manufacturing. To establish benchmark results on these tasks, we conduct a thorough evaluation of 20 top-performing LLMs. We anticipate that this evaluation will help analyze important strengths and shortcomings of both general domain models and domain-specific models, fostering their continued development and growth to better serve the Chinese manufacturing industry, thereby enabling AI-driven smart manufacturing. The dataset is available at https://github.com/KnowdeeAI/Manu-Eval .

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Manu-Eval: A Chinese Language Understanding Benchmark for Manufacturing Industry

  • Xiaoyi Liu,
  • Shuangtao Yang,
  • Xiaozheng Dong,
  • Honghui Rong,
  • Bo Fu

摘要

Large language models (LLMs) have demonstrated remarkable capabilities across various domains, fostering growing interest in their applications within the manufacturing industry. This paper presents a comprehensive evaluation of LLMs across 22 subcategories spanning 8 major sectors in the manufacturing industry, including machinery, automobile, electronics, chemical, light industry, pharmaceutical, transportation, and food manufacturing. To establish benchmark results on these tasks, we conduct a thorough evaluation of 20 top-performing LLMs. We anticipate that this evaluation will help analyze important strengths and shortcomings of both general domain models and domain-specific models, fostering their continued development and growth to better serve the Chinese manufacturing industry, thereby enabling AI-driven smart manufacturing. The dataset is available at https://github.com/KnowdeeAI/Manu-Eval .