Sentiment classification (SC) of user-generated comments (UGC) on social media and e-commerce websites written in code-mixed languages is a challenging task in natural language processing (NLP). In current days, Transformer-based models with attention mechanism are achieving state-of-the-art (SoTA) results for NLP tasks. When many authors have explored NLP tasks for code-mixed Dravidian languages using Transformer models, the languages belonging to the Indo-Aryan language family still have not been addressed to a large extent. The present work focuses on SC task on eight code-mixed Indo-Aryan languages such as Assamese–English, Bengali–English, Hindi–English, Marathi–English, Nepali–English, Odia–English, Punjabi–English and Urdu–English. In this present survey, we have studied in detail the existing code-mixed datasets of Indo-Aryan languages which have addressed SC tasks using Transformer-based models, the evolution of Transformer-based models for Indian Languages, research progress, analysis, challenges and research scopes during the evaluation of Indo-Aryan code-mixed dataset using Transformer-based models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Sentiment Classification in Code-Mixed Indo-Aryan Languages: A Transformer-Based Survey

  • Saikat Roy,
  • Jatinderkumar R. Saini

摘要

Sentiment classification (SC) of user-generated comments (UGC) on social media and e-commerce websites written in code-mixed languages is a challenging task in natural language processing (NLP). In current days, Transformer-based models with attention mechanism are achieving state-of-the-art (SoTA) results for NLP tasks. When many authors have explored NLP tasks for code-mixed Dravidian languages using Transformer models, the languages belonging to the Indo-Aryan language family still have not been addressed to a large extent. The present work focuses on SC task on eight code-mixed Indo-Aryan languages such as Assamese–English, Bengali–English, Hindi–English, Marathi–English, Nepali–English, Odia–English, Punjabi–English and Urdu–English. In this present survey, we have studied in detail the existing code-mixed datasets of Indo-Aryan languages which have addressed SC tasks using Transformer-based models, the evolution of Transformer-based models for Indian Languages, research progress, analysis, challenges and research scopes during the evaluation of Indo-Aryan code-mixed dataset using Transformer-based models.