In the era of big data, sentiment analysis on social media platforms presents unique challenges due to the noisy and unstructured nature of the data. This study introduces a novel deep learning approach for discovering causal relationships within such noisy web data using an attention-enhanced architecture. The proposed model leverages an embedding layer, bidirectional LSTM, attention mechanisms, global average pooling, and dropout techniques, followed by a dense layer with sigmoid activation for binary classification. We validated our model on the Sentiment140 dataset, achieving a test accuracy of 82.36% and an F1-score of 0.8246, outperforming traditional machine learning models such as Logistic Regression, Support Vector Machines (SVM), and Random Forests. An extensive ablation study was conducted to assess the contributions of key model components, confirming the critical role of attention mechanisms and bidirectional LSTM in improving performance and interpretability. Beyond its superior performance in sentiment classification, our model provides insights into the causal dynamics driving sentiment shifts on social media. This work significantly contributes to the fields of web data mining and causal inference by offering a robust framework capable of handling the complexities of noisy, high-dimensional data commonly encountered in online environments..

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Discovering Causal Relationships in Noisy Web Data for Sentiment Classification Using Attention Mechanisms

  • Miloud Mihoubi,
  • Meriem Zerkouk,
  • Belkacem Chikhaoui

摘要

In the era of big data, sentiment analysis on social media platforms presents unique challenges due to the noisy and unstructured nature of the data. This study introduces a novel deep learning approach for discovering causal relationships within such noisy web data using an attention-enhanced architecture. The proposed model leverages an embedding layer, bidirectional LSTM, attention mechanisms, global average pooling, and dropout techniques, followed by a dense layer with sigmoid activation for binary classification. We validated our model on the Sentiment140 dataset, achieving a test accuracy of 82.36% and an F1-score of 0.8246, outperforming traditional machine learning models such as Logistic Regression, Support Vector Machines (SVM), and Random Forests. An extensive ablation study was conducted to assess the contributions of key model components, confirming the critical role of attention mechanisms and bidirectional LSTM in improving performance and interpretability. Beyond its superior performance in sentiment classification, our model provides insights into the causal dynamics driving sentiment shifts on social media. This work significantly contributes to the fields of web data mining and causal inference by offering a robust framework capable of handling the complexities of noisy, high-dimensional data commonly encountered in online environments..