Aspect Based Sentiment Analysis on Code Mixed Languages
摘要
The mixing of two or more languages or language varieties is known as Codemixing. It is a widespread practice in India, a nation renowned for its language diversity, and it significantly affects how individuals communicate on social media and in casual discussions. For language processing and machine translation tasks, the widespread occurrence of code-mixing in social media platforms poses a significant challenge. The volume of unstructured text in code-mixed form found on these platforms draws attention to an important area of NLP research. These texts need to be processed appropriately to help monolingual users and language processing models understand them. In this paper, we present a new task in order to contribute towards code-mixed Aspect Based Sentiment Analysis research. By building a codemixed Hinglish dataset for ABSA—that is, a dataset that combines Hindi and English—and annotating it with aspect terms and their sentiment values, we provide a benchmark setting. We develop a number of deep learning-based models for sentiment analysis and aspect phrase extraction to show how the dataset can be used effectively, and we set them as the standards for future studies in this area.