Writing high-quality multiple-choice (MC) test items is crucial for accurately assessing student’s knowledge and skills. Linguistic and formatting considerations are critical for writing high-quality items, yet they are often overlooked by item writers. Manually verifying the linguistic and formatting guidelines for every MC item is both time-consuming and prone to error. This study introduces an augmented intelligence multi-agent system that detects item-writing flaws, provides automated feedback, and offers AI-generated suggestions to help item writers produce high-quality MC items. The system leverages established item-writing guidelines for flaw detection and incorporates GPT-4 to generate feedback and recommend corrections. Results indicate the system was accurate on three criteria: 85% accuracy in producing flawless items, 90% accuracy in ensuring linguistic clarity, and 70% accuracy in maintaining content validity. These findings support the use of augmented intelligence systems for creating high-quality MC test items.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Augmented Intelligence System for Automated Quality Control and Feedback Generation of Multiple Choice Test Items

  • Tahereh Firoozi,
  • Lia Daniels,
  • Vijay Daniels,
  • Mark Gierl

摘要

Writing high-quality multiple-choice (MC) test items is crucial for accurately assessing student’s knowledge and skills. Linguistic and formatting considerations are critical for writing high-quality items, yet they are often overlooked by item writers. Manually verifying the linguistic and formatting guidelines for every MC item is both time-consuming and prone to error. This study introduces an augmented intelligence multi-agent system that detects item-writing flaws, provides automated feedback, and offers AI-generated suggestions to help item writers produce high-quality MC items. The system leverages established item-writing guidelines for flaw detection and incorporates GPT-4 to generate feedback and recommend corrections. Results indicate the system was accurate on three criteria: 85% accuracy in producing flawless items, 90% accuracy in ensuring linguistic clarity, and 70% accuracy in maintaining content validity. These findings support the use of augmented intelligence systems for creating high-quality MC test items.