Evaluation of the Quality of AI-Generated Scientific Text Under Different Types of Cognitive Complexity Tasks
摘要
As Artificial Intelligence Generated Content (AIGC) continues to deepen its application in the field of scientific research, this study aims to explore the current quality of AIGC in completing research tasks, providing insights for improving AIGC in the scientific research domain. This study first reviews and summarizes existing information quality evaluation frameworks and AIGC-related research to propose quality evaluation criteria for AIGC in the research context. Then, by setting research tasks with different cognitive complexities, user experiments were conducted on the ChatGPT and ERNIE Bot platforms to select appropriate AIGC quality evaluation criteria for these tasks. The quality of AIGC generated by ChatGPT and ERNIE Bot was evaluated based on the selected criteria, revealing the strengths and weaknesses of current AIGC in meeting users’ research information needs. The results show that users generally value relevance, professionalism, and readability when evaluating AIGC for research tasks. However, attention to specific criteria such as accuracy, diversity, coherence, and creativity varies depending on the cognitive complexity of the research tasks. Additionally, AIGC performs well in understanding, evaluating, and creating tasks but has significant shortcomings in remembering and analyzing tasks, particularly in terms of accuracy and professionalism.