Parts-of-Speech Tagger in Assamese Using LSTM and Bi-LSTM
摘要
Parts-of-speech (POS) tagging is considered one of the most challenging fields in natural language processing (NLP). The objective of this research is to develop a POS tagger for the Assamese language. Due to the scarcity of digital linguistic resources, Assamese lacks high-performing POS taggers. To fill this gap, long short-term memory (LSTM) and bidirectional long short-term memory (Bi-LSTM) are explored in the proposed research to develop a POS tagger using an Assamese POS corpus. It is important to note that this experiment faced difficulties in understanding and managing natural language for computational linguistics, which was also anticipated. The Assamese corpus considered in this research comprises around 50,000 words. At the initial stage, while examining the first few sets of data, which are about 20,000 words, it was noticed that the taggers yielded satisfactory results. Based on the result derived, the Assamese corpus size has been enhanced to 50,000 words and better performance is noted in terms of accuracy rate, precision, recall, and F1-score. As a result, an accuracy of 91.20% is achieved for LSTM and 91.72% for Bi-LSTM. Concerned with substantial research from the NLP perspective for the Assamese language and for comparative purposes, a comparison between existing POS taggers in the Assamese language and the proposed work is also presented.