Automated Math Word Problem Knowledge Component Labeling and Recommendation
摘要
The accurate annotation of math exercises and word problems according to a diverse set of knowledge components is an important task for many education applications. It is an extremely complex and resource-intensive process, and has traditionally been done manually by experienced educators. There has been research work in recent years to apply machine learning to automate the process, however, due to the varied datasets used by different researchers and the private nature of these datasets, there is no good benchmark for which model performs the best for such a task. Moreover, the datasets used in literature typically comprise math exercises that follow similar templates. In this paper, we benchmark some of the best reported models on a math word problem dataset with fine-grained knowledge component annotations, and highlight the challenges in making accurate predictions. We propose models to improve the prediction accuracy, and demonstrate how they can be used to extract similar problems based on knowledge component from an unlabelled pool of questions, thus empowering educators to better identify the knowledge gaps of their students and tailor suitable practice problems for them.