Natural Language Processing for Tulu: Challenges, Review and Future Scope
摘要
This paper provides a comprehensive analysis of publicly-available research done to date on Natural Language Processing (NLP) in Tulu while exploring its development, challenges, and future scope. Tulu is a low-resource Dravidian language with more than 2.5 million speakers. Work done in NLP for Tulu includes code-mixed corpus generation, optical character recognition of historical manuscripts, machine translation, sentiment analysis, speech recognition, and morphological analysis. However, due to data scarcity, morphological complexity, and code-mixing, challenges arise for NLP practitioners and more research and innovation are needed. Future work in NLP for Tulu involves expanding code-mixed corpora, improving machine translation and speech recognition, cross-lingual transfer learning, specialized named entity recognition, and interdisciplinary collaborations. Unlocking Tulu’s potential as a language with a rich cultural heritage requires addressing these challenges and embracing future opportunities to enhance linguistic diversity and accessibility of NLP technologies.