Sentiment Analysis in Assamese: Unveiling Emotions in Native Texts
摘要
Sentiment Analysis (SA) refers to the process of detecting and classifying opinions conveyed within a given piece of text, audio, or other forms of communication, which is one of the main sub-field of Natural Language Processing (NLP). There is numerous amount of SA task in the language like English. Because of the rapid growth of social media, the use of various regional languages from different regions of India is increasing dramatically. Nowadays, people often prefer to leave comments or reviews in their own language. So, the Sentiment analysis in regional language is becoming very important. Due to the unavailability of resources, a very little amount of sentiment analysis task is done in low-resource language like Assamese. Our primary objective is to develop an Assamese-language dataset for sentiment analysis and to analyze different ML (Machine Learning) classifier on the dataset. We have created a dataset of 8000 sentences and performance of the classifiers such as Logistic Regression classifier (LR), SVC, MultinomialNB, Decision Tree classifier(DT), Random Forest classifier (RF), Extra Tree Classifier on the dataset are evaluated. The SVC (Support Vector Classification) outperformed all other classifiers, achieving an F1 score of 67.9%. Given the absence of various sentiment analysis tools, such as POS taggers, in low-resource languages, there is significant potential for improvement in this area. Developing and enhancing these tools could greatly benefit the processing and analysis of text in these languages, making sentiment analysis more accurate and accessible.