错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deep Learning Based Sentiment Analysis of Tamil–English YouTube Comments

  • Malliga Subramanian,
  • S. V. Kogilavani,
  • D. Gowthesh,
  • S. Lohith,
  • S. Mithunajha

摘要

With the exponential growth of user-generated content on YouTube, understanding the sentiments expressed in comments has become crucial. The research work aims to analyze the impact of online sentiments on social welfare through a comprehensive analysis. While sentiment analysis tools and resources may be more readily available for widely spoken languages, the increasing demand for extending sentiment analysis to low-resource languages like Tamil is also a pressing issue. In this work, transformer-based models such as BERT, RoBERTa, and XLMRobertA are used to extract the features and detect the sentiment of the Tamil-English YouTube comments and these models have been specifically designed for text-based multilingual aspects. The combined approach leveraging Support Vector Machine (SVM) and transformer models is enhanced through the application of SMOTE (Synthetic Minority Over-sampling Technique) and Random OverSampling techniques, contributing to the increased performance. These oversampling techniques are employed to address class imbalance in the dataset, ensuring a more robust and balanced model performance. Of all the proposed models, XLMRoBERTa outperforms with 63.60% of accuracy. In the study, a multilingual dataset of YouTube comments was collected, with a particular focus on low-resource languages. The challenges in sentiment analysis include limited code-mixing, and unique linguistic characteristics of these languages, which pose difficulties for traditional sentiment analysis approaches.