Sentiment analysis for low-resource languages, including Marathi, has garnered increased interest, driven by the rising volume of digital content generated in regional languages. This survey paper provides a comprehensive review of 30 research papers focusing on opinion mining and sentiment analysis in the Marathi language and other related domains. The reviewed works are categorized based on their primary contributions, including model development, methodology proposals, and surveys of existing techniques. Various approaches such as lexicon-based methods, classical machine learning models, and deep learning techniques have been explored, highlighting the evolving trends and challenges in handling Marathi language text, including issues with transliteration and code-mixing. Besides the survey, this paper introduces an innovative approach for sentiment analysis of transliterated Marathi text. Our approach involves manually curating a sentiment wordlist using a Marathi-to-English dictionary and assigning weighted sentiment scores. We collected user-generated content from social media platforms such as Instagram, Twitter, and YouTube, which was then pre-processed and analyzed based on the occurrence of sentiment words. The sentences were labeled and used to train machine learning models, specifically Logistic Regression (LR) and Support Vector Machines (SVM). The models’ performance was assessed using standard evaluation metrics such as accuracy, precision, recall, and F1-score. Future work will explore additional pre-processing techniques and advanced models to optimize sentiment classification in transliterated Marathi. By combining a detailed literature review with our practical implementation, this paper aims to provide insights into the state-of-the-art in Marathi sentiment analysis and set the stage for future research to address existing gaps and challenges.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Opinion Mining for Marathi Language Using Machine Learning

  • Siddhi Subhashsing Pardeshi,
  • Gurunath Keshav Salve,
  • Sayali Shyam Thakur,
  • Rishikeh J. Sutar

摘要

Sentiment analysis for low-resource languages, including Marathi, has garnered increased interest, driven by the rising volume of digital content generated in regional languages. This survey paper provides a comprehensive review of 30 research papers focusing on opinion mining and sentiment analysis in the Marathi language and other related domains. The reviewed works are categorized based on their primary contributions, including model development, methodology proposals, and surveys of existing techniques. Various approaches such as lexicon-based methods, classical machine learning models, and deep learning techniques have been explored, highlighting the evolving trends and challenges in handling Marathi language text, including issues with transliteration and code-mixing. Besides the survey, this paper introduces an innovative approach for sentiment analysis of transliterated Marathi text. Our approach involves manually curating a sentiment wordlist using a Marathi-to-English dictionary and assigning weighted sentiment scores. We collected user-generated content from social media platforms such as Instagram, Twitter, and YouTube, which was then pre-processed and analyzed based on the occurrence of sentiment words. The sentences were labeled and used to train machine learning models, specifically Logistic Regression (LR) and Support Vector Machines (SVM). The models’ performance was assessed using standard evaluation metrics such as accuracy, precision, recall, and F1-score. Future work will explore additional pre-processing techniques and advanced models to optimize sentiment classification in transliterated Marathi. By combining a detailed literature review with our practical implementation, this paper aims to provide insights into the state-of-the-art in Marathi sentiment analysis and set the stage for future research to address existing gaps and challenges.