The growing presence of artificial intelligence (AI) in education has sparked significant interest, particularly regarding its application in automated grading of student assignments. This study aims to assess the effectiveness and reliability of AI in evaluating academic performance compared to human grading. Using a dataset containing grades assigned by teachers alongside those generated by an AI system, several statistical analyses were conducted, including ANOVA, Pearson’s test, linear regression, and overall model fit tests. The results indicate that AI provides notable grading consistency but has certain limitations in interpreting contextual nuances and subjective elements that teachers naturally incorporate. On average, AI-assigned slightly higher scores than human teachers (p < 0.05) and was over 90% faster in grading. However, algorithmic biases were identified, favoring structural grading criteria over contextual comprehension. These findings suggest that AI can serve as an effective complementary tool for academic assessment, particularly in standardized evaluations. However, human oversight remains essential to ensure fairness and contextual relevance. This study contributes to the ongoing discussion on integrating AI into pedagogical practices by highlighting both its strengths and challenges. Furthermore, it advocates for a hybrid evaluation approach, leveraging AI’s efficiency while preserving the pedagogical expertise of human educators.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Artificial Intelligence in Educational Assessment: A Comparative Study of AI and Human Grading Performance

  • Hassou Kamal,
  • Karim Khaddouj

摘要

The growing presence of artificial intelligence (AI) in education has sparked significant interest, particularly regarding its application in automated grading of student assignments. This study aims to assess the effectiveness and reliability of AI in evaluating academic performance compared to human grading. Using a dataset containing grades assigned by teachers alongside those generated by an AI system, several statistical analyses were conducted, including ANOVA, Pearson’s test, linear regression, and overall model fit tests. The results indicate that AI provides notable grading consistency but has certain limitations in interpreting contextual nuances and subjective elements that teachers naturally incorporate. On average, AI-assigned slightly higher scores than human teachers (p < 0.05) and was over 90% faster in grading. However, algorithmic biases were identified, favoring structural grading criteria over contextual comprehension. These findings suggest that AI can serve as an effective complementary tool for academic assessment, particularly in standardized evaluations. However, human oversight remains essential to ensure fairness and contextual relevance. This study contributes to the ongoing discussion on integrating AI into pedagogical practices by highlighting both its strengths and challenges. Furthermore, it advocates for a hybrid evaluation approach, leveraging AI’s efficiency while preserving the pedagogical expertise of human educators.