Evaluating ChatGPT-4’s Performance in Summarizing Academic Texts: A Qualitative Analysis Using Self-Authored Publications
摘要
This study examines the performance of ChatGPT-4 in summarizing academic texts by comparing AI-generated summariesAI-generated summaries with human-generated ones. To this end, using a sample of 12 self-authored academic articles, both ChatGPT- and human-generated summaries were first created and then evaluated by two independent evaluators based on the following five criteria: (1) accuracyAccuracy; (2) conciseness; (3) coherenceCoherence; (4) comprehensivenessComprehensiveness; and (5) readabilityReadability. Inter-rater reliability was calculated using Cohen’s kappa coefficient. The results revealed that, while ChatGPT-4 excelled in producing concise and readable summaries, it frequently omitted nuanced details and supporting evidence, which resulted in moderate comprehensivenessComprehensiveness. By contrast, human summaries were evaluated as more accurate, comprehensive, and capable of capturing critical counterarguments and detailed insights. Common issues with ChatGPT-4 summaries included abrupt transitions, occasional inaccuracies, and redundancies. However, despite these limitations, ChatGPT-4 effectively captured main arguments and conclusions, suggesting its utility for quick content overviews. Taken together, the results of the present study underscore the need for human oversight of AI-generated summarizations, along with further model training to enhance AI summarization capabilities in academic contexts. These findings provide valuable insights into the integration of AI tools in scholarly research and writing.