Manu-Eval: A Chinese Language Understanding Benchmark for Manufacturing Industry
摘要
Large language models (LLMs) have demonstrated remarkable capabilities across various domains, fostering growing interest in their applications within the manufacturing industry. This paper presents a comprehensive evaluation of LLMs across 22 subcategories spanning 8 major sectors in the manufacturing industry, including machinery, automobile, electronics, chemical, light industry, pharmaceutical, transportation, and food manufacturing. To establish benchmark results on these tasks, we conduct a thorough evaluation of 20 top-performing LLMs. We anticipate that this evaluation will help analyze important strengths and shortcomings of both general domain models and domain-specific models, fostering their continued development and growth to better serve the Chinese manufacturing industry, thereby enabling AI-driven smart manufacturing. The dataset is available at https://github.com/KnowdeeAI/Manu-Eval .