错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Collection and Automatic Analysis with Natural Language Processing on a Corpus of Andean Oral Literature Implemented on the Web

  • Ivan Soria Solis,
  • Carlos Yinmel Castro Buleje,
  • Humberto Silvera Reynaga,
  • Mauro Felix Mamani Macedo,
  • Dionicia León Soncco,
  • Alejandro Giancarlo Mautino Guillen

摘要

Oral literature is transmitted through tradition and is not typically cataloged or recorded. Analyzing this type of literature presents certain challenges as there are limited computer tools available for its collection, preservation and study. In particular, Andean oral literature in Quechua faces significant challenges as there are scarce computer resources to support its study and processing. This study aims to establish a Web-based repository for the development of a corpus comprising texts of Andean oral literature. We utilized this corpus to train a natural language processing model that enables automatic analysis and topic classification of the texts. We applied the FastText tool, which features a pre-trained model for Quechua. Due to the variability in writing that characterizes this language, Word Embeddings were preferred to represent the meaning according to the context, which can handle similar words. The model was evaluated using the corpus dataset in conjunction with the pre-trained vectors. Texts were collected and then trained on a classification model. Metrics of accuracy, precision, recall, and f1 were evaluated. The model was determined to perform best when not utilizing a pre-trained model.