Contextual Features for Automatic Essay Scoring in Portuguese
摘要
Automated Essay Scoring (AES) efficacy often varies across linguistic and contextual nuances. This study addresses this gap by proposing and evaluating a contextualized approach tailored for Portuguese. Unlike prior research, which often focused on overall scores or limited to general-purpose features, we explored devising contextualized feature extractors and investigated their impact on predictive performance. Our analysis encompassed the proposed specific features (conjunctions, syntactic quantification, and entity recognition) and two well-established baselines (i.e., TF-IDF and Coh-Metrix). Utilizing the Essay-BR dataset (n = 6,563 essays), we investigated our approach through classification and regression tasks supported by diverse machine learning algorithms and optimization techniques. Mainly, we found that our approach enhanced predictive performance when combined with existing techniques. Our findings reveal the importance of addressing and considering contextual nuances in AES, revealing insights that might help accelerate the evaluation of essays in a large-scale setting.