As Artificial Intelligence Generated Content (AIGC) continues to deepen its application in the field of scientific research, this study aims to explore the current quality of AIGC in completing research tasks, providing insights for improving AIGC in the scientific research domain. This study first reviews and summarizes existing information quality evaluation frameworks and AIGC-related research to propose quality evaluation criteria for AIGC in the research context. Then, by setting research tasks with different cognitive complexities, user experiments were conducted on the ChatGPT and ERNIE Bot platforms to select appropriate AIGC quality evaluation criteria for these tasks. The quality of AIGC generated by ChatGPT and ERNIE Bot was evaluated based on the selected criteria, revealing the strengths and weaknesses of current AIGC in meeting users’ research information needs. The results show that users generally value relevance, professionalism, and readability when evaluating AIGC for research tasks. However, attention to specific criteria such as accuracy, diversity, coherence, and creativity varies depending on the cognitive complexity of the research tasks. Additionally, AIGC performs well in understanding, evaluating, and creating tasks but has significant shortcomings in remembering and analyzing tasks, particularly in terms of accuracy and professionalism.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluation of the Quality of AI-Generated Scientific Text Under Different Types of Cognitive Complexity Tasks

  • Hui Peng,
  • Shujun Liu,
  • Lei Li

摘要

As Artificial Intelligence Generated Content (AIGC) continues to deepen its application in the field of scientific research, this study aims to explore the current quality of AIGC in completing research tasks, providing insights for improving AIGC in the scientific research domain. This study first reviews and summarizes existing information quality evaluation frameworks and AIGC-related research to propose quality evaluation criteria for AIGC in the research context. Then, by setting research tasks with different cognitive complexities, user experiments were conducted on the ChatGPT and ERNIE Bot platforms to select appropriate AIGC quality evaluation criteria for these tasks. The quality of AIGC generated by ChatGPT and ERNIE Bot was evaluated based on the selected criteria, revealing the strengths and weaknesses of current AIGC in meeting users’ research information needs. The results show that users generally value relevance, professionalism, and readability when evaluating AIGC for research tasks. However, attention to specific criteria such as accuracy, diversity, coherence, and creativity varies depending on the cognitive complexity of the research tasks. Additionally, AIGC performs well in understanding, evaluating, and creating tasks but has significant shortcomings in remembering and analyzing tasks, particularly in terms of accuracy and professionalism.