Tenyidie language belonging to the Tibeto-Burman family is one of the major languages of Nagaland in northeastern India. It is a low-resource, tonal, SOV, and a high agglutinative language. Stemming is an important Natural Language Processing (NLP) task wherein the stem or the root word is identified and extracted. To the best of the authors knowledge, there has been no work reported on corpus creation and stemming using deep learning for the Tenyidie language. In this paper, we build an annotated corpus of 5400 manually stemmed words in Tenyidie to apply deep learning to the stemming task. We applied state-of-the-art deep learning models, such as the Encoder-decoder models, and compared our results with the rule-based approach, which uses both affix stripping and root word matching methods. Our experimental result shows that the deep learning models outperform the rule-based approach. We achieved the highest accuracy of 85.55% with Encoder-decoder (BLSTM) with attention. We also provide an analysis of the overstemmed and understemmed words in our test dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Preliminary Work on Tenyidie Stemming Corpus Creation and Deep Learning Applications

  • Teisovi Angami,
  • Themrichon Tuithung

摘要

Tenyidie language belonging to the Tibeto-Burman family is one of the major languages of Nagaland in northeastern India. It is a low-resource, tonal, SOV, and a high agglutinative language. Stemming is an important Natural Language Processing (NLP) task wherein the stem or the root word is identified and extracted. To the best of the authors knowledge, there has been no work reported on corpus creation and stemming using deep learning for the Tenyidie language. In this paper, we build an annotated corpus of 5400 manually stemmed words in Tenyidie to apply deep learning to the stemming task. We applied state-of-the-art deep learning models, such as the Encoder-decoder models, and compared our results with the rule-based approach, which uses both affix stripping and root word matching methods. Our experimental result shows that the deep learning models outperform the rule-based approach. We achieved the highest accuracy of 85.55% with Encoder-decoder (BLSTM) with attention. We also provide an analysis of the overstemmed and understemmed words in our test dataset.