Automatic Population of Educational Ontology from Course Materials
摘要
In this paper, we present a procedure for automatic ontology population with the data from MIT OpenCourseWare course materials. The data was collected for ten Computer Science domains, by scraping the text from PowerPoint presentation and preprocessing it. In the next step, a word2vec neural network was trained for each domain. The obtained words with their embeddings were next used for populating the ontology. The procedure is evaluated by analyzing the clusters of word embeddings and comparing the affiliation of embeddings to clusters with their original class.