Sentiment Classification in Code-Mixed Indo-Aryan Languages: A Transformer-Based Survey
摘要
Sentiment classification (SC) of user-generated comments (UGC) on social media and e-commerce websites written in code-mixed languages is a challenging task in natural language processing (NLP). In current days, Transformer-based models with attention mechanism are achieving state-of-the-art (SoTA) results for NLP tasks. When many authors have explored NLP tasks for code-mixed Dravidian languages using Transformer models, the languages belonging to the Indo-Aryan language family still have not been addressed to a large extent. The present work focuses on SC task on eight code-mixed Indo-Aryan languages such as Assamese–English, Bengali–English, Hindi–English, Marathi–English, Nepali–English, Odia–English, Punjabi–English and Urdu–English. In this present survey, we have studied in detail the existing code-mixed datasets of Indo-Aryan languages which have addressed SC tasks using Transformer-based models, the evolution of Transformer-based models for Indian Languages, research progress, analysis, challenges and research scopes during the evaluation of Indo-Aryan code-mixed dataset using Transformer-based models.