Sindhi POS Tagger Using LSTM and Pre-Trained Word Embeddings
摘要
The chapter shows the development of a part-of-speech (POS) tagger for Sindhi, which is a highly resource-poor language. For our study, we have used Sindhi in the Devanagari script. We have developed a corpus of thirty thousand POS-tagged sentences and developed Glove word embeddings for the monolingual Sindhi corpus of 1 lac sentences. We have used LSTM and Glove word embeddings for developing our POS tagger. The developed system was compared with a Hidden Markov Model (HMM)-based POS tagger developed for Sindhi (Nathani and Joshi, Part of speech tagging for a resource poor language: Sindhi in Devanagari script using HMM and CRF. In: Proceedings of the 18th international conference on natural language processing, 2021). The evaluation shows significant improvement over the previous study. The HMM POS tagger achieved an overall accuracy of 81.85%, while the proposed LSTM tagger achieved an accuracy of 96.15%.