A Probabilistic Graphical Model for Concept Identification from Educational Documents
摘要
The large volume of online educational materials makes it difficult for learners to find adequate resources for better learning. Understanding these materials relies on identifying key concepts essential for comprehension. Automatic concept extraction is an important task in educational data mining and is similar to keyphrase extraction in Natural Language Processing (NLP). This process helps identify key ideas, organize documents, and build an insightful learning path. We present a probabilistic approach for concept extraction. Candidate concepts are generated using Wikipedia anchor texts. We identify the necessary concepts based solely on the educational context of a particular document using a graph-based probabilistic model. Evaluation of our method on two datasets (namely, a Physics school textbook and Physics articles) outperforms existing unsupervised and supervised methods.