错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Semantic proximity assessment in Bhojpuri and Maithili: a word embedding perspective

  • Arun Kumar Yadav,
  • Abhishek Kumar,
  • Mohit Kumar,
  • Divakar Yadav

摘要

Natural Language Processing has been extensively researched for languages with abundant resources like English and Spanish, but low-resource languages present unique challenges and opportunities. Consequently, the performance of learning algorithms for downstream tasks in low resource languages, such as the Indic languages Bhojpuri and Maithili, is unsatisfactory. To address these challenges, we collect a corpus of 20,000 sentences in both Bhojpuri and Maithili each and use them to develop word representations using popular the approaches such as Word2Vec, FastText, and GloVe. The evaluation of the word representation is accomplished through machine translation and text classification tasks. Among the various models used in the experiments, word2Vec with Bidirectional Encoder Representations from Transformers (BERT) outperforms GloVe and FastText in machine translation task, achieving a BiLingual Evaluation Understudy score of 32.95 for Bhojpuri to English and 28.95 for Maithili to English. Additionally, the BERT model provides a good text classification accuracy of 71.23% for Bhojpuri and 66.91% for Maithili. This work may serve as a reasonable beginning point for further research in low-resource Indic languages.