Sentiment analysis (SA) is a crucial natural language processing (NLP) tool for understanding user sentiments, particularly in business intelligence. However, the lack of labeled data in low-resource languages like Gujarati limits the effectiveness of such models. We propose a novel data augmentation pipeline specifically designed for Gujarati movie review datasets to balance the class distribution. This paper presents manually tagged SA datasets for Gujarati movie reviews, annotated with three sentiment classes: positive, negative, and neutral. IndoWordNet is used to generate new data points for under-represented classes by replacing words with their synonyms, while IndicBERT embeddings are utilized to verify the semantic similarity of the augmented sentences. After balancing the dataset, standard classifiers such as SVM, Naive Bayes, Logistic Regression, and KNN were used to evaluate improvements in sentiment classification performance. An increase of 21% in the macro F1 score was empirically observed when the augmented balanced dataset was used for the classification task.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Data Augmentation of Gujarati Texts for Sentiment Analysis

  • Nikita Desai,
  • Vipul Dabhi

摘要

Sentiment analysis (SA) is a crucial natural language processing (NLP) tool for understanding user sentiments, particularly in business intelligence. However, the lack of labeled data in low-resource languages like Gujarati limits the effectiveness of such models. We propose a novel data augmentation pipeline specifically designed for Gujarati movie review datasets to balance the class distribution. This paper presents manually tagged SA datasets for Gujarati movie reviews, annotated with three sentiment classes: positive, negative, and neutral. IndoWordNet is used to generate new data points for under-represented classes by replacing words with their synonyms, while IndicBERT embeddings are utilized to verify the semantic similarity of the augmented sentences. After balancing the dataset, standard classifiers such as SVM, Naive Bayes, Logistic Regression, and KNN were used to evaluate improvements in sentiment classification performance. An increase of 21% in the macro F1 score was empirically observed when the augmented balanced dataset was used for the classification task.