Designing Essay Questions for Effective Automatic Scoring
摘要
The domain of automatic essay scoring (AES) has increasingly garnered attention, buoyed by advancements in natural language processing (NLP) and deep learning. Despite the progress, much of the existing research has narrowly focused on developing NLP models, and has conducted performance evaluations across diverse essay types without considering the influence of question characteristics on scoring accuracy. This study reviews pedagogical literature on essay assessments to introduce a set of criteria for crafting essay questions that can be scored automatically with high accuracy. Our criteria emphasize measurable learning objectives, the consolidation of questions to a single learning goal, and the stipulation of restricted answers to ensure ease and consistency in evaluation. Experiments with the ASAP dataset show variations in the scoring accuracy of BERT-based AES systems across essay types. Notably, essays aligned with our proposed criteria exhibited superior performance, showcasing an improvement in scoring accuracy by over 40%. This enhancement not only the underscores efficacy of our approach, but more broadly implies that integrating these criteria could expand the utility of AES systems, enabling their more widespread application.