Sentiment Analysis for Cross-Lingual Kannada–English Language Pair
摘要
Sentiment analysis involves analyzing text to identify the emotion behind it and has numerous applications. Kannada is a Dravidian language spoken in India that presents a challenge to monolingual NLP models due to the nature of its script and limited language technology resources, but the Kannada–English code-mixed dataset from the Dravidian-CodeMix-FIRE 2021 shared task provides an opportunity for sentiment analysis research. The dataset was collected and text processing has been carried out at different levels for sentiment analysis. Text cleaning, transliteration, spell check, and translation to monolingual Kannada are some of the stages involved. Different models are built and evaluated based on performance metrics such as precision, recall, F1-score, and accuracy. The results indicate that the proposed approach can achieve high accuracy in predicting sentiment for Kannada code-mixed sentences. The paper concludes with a discussion of the potential applications and future research directions for cross-lingual sentiment analysis.