<p>Current AI-based programming education assessment systems do not involve process-oriented learning and a multi-faceted evidence of learning other than code correctness. This paper presents a three-partnering model of AI-Teacher, AI-Student, and Aggregator agents that combines multiple sources of evidence (syntax, semantics, process, and behavior) into a centralized learner state for formative OOP assessment. The framework evaluated on 2,550 student submissions has a larger grading accuracy (Cohen’s <InlineEquation ID="IEq1"><EquationSource Format="TEX">\(\kappa = 0.90\)</EquationSource></InlineEquation>), adequate error coverage (<InlineEquation ID="IEq2"><EquationSource Format="TEX">\(85.0\%\)</EquationSource></InlineEquation>), high explanation quality (4.6/5), and substantial simulated learning gains (<InlineEquation ID="IEq3"><EquationSource Format="TEX">\(26.1\%\)</EquationSource></InlineEquation> after six feedback cycles), surpassing traditional autograders and large-effect single-agent LLMs. Handling 1,&#xa0;710 submissions per hour (<InlineEquation ID="IEq4"><EquationSource Format="TEX">\(67\%\)</EquationSource></InlineEquation> points faster than LLM-only) the framework is a scalable, pedagogically valuable option in automated formative assessment in programming education.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Agentic formative assessment for object oriented programming through multi source evidence aggregation

  • Xiaomei Ding,
  • Huaibao Ding,
  • Fei Zhou,
  • Qiongpei Wang,
  • Fang Xia,
  • Yujie Ma,
  • Jiayun Lang

摘要

Current AI-based programming education assessment systems do not involve process-oriented learning and a multi-faceted evidence of learning other than code correctness. This paper presents a three-partnering model of AI-Teacher, AI-Student, and Aggregator agents that combines multiple sources of evidence (syntax, semantics, process, and behavior) into a centralized learner state for formative OOP assessment. The framework evaluated on 2,550 student submissions has a larger grading accuracy (Cohen’s \(\kappa = 0.90\)), adequate error coverage (\(85.0\%\)), high explanation quality (4.6/5), and substantial simulated learning gains (\(26.1\%\) after six feedback cycles), surpassing traditional autograders and large-effect single-agent LLMs. Handling 1, 710 submissions per hour (\(67\%\) points faster than LLM-only) the framework is a scalable, pedagogically valuable option in automated formative assessment in programming education.