Digital media have grown at a remarkably fast pace, which in turn has increased the consumption of news through e-papers. In such a scenario, the ability to quickly estimate the sentiment of a news article becomes highly valuable. This research report encompasses a broad study on sentiment analysis related to e-paper news articles, by applying techniques in NLP, sentiment analysis models classify news articles into undertones that are either positive, negative, or neutral. The project begins with collecting a long and wide dataset of e-paper news articles, covering a wide range of categories, sources, and different timelines. Further sentiment analysis will hence be derived from the dataset. Preprocessing the dataset beforehand through NLP is done prior to the sentiment analysis. These steps, including text cleaning, stop word removal, and stemming, are involved in this cleaning process to further improve the quality of the dataset. We have chosen logistic regression over the other machine learning models due to its effectiveness in text classification tasks, its ability to model probabilities, and its strong performance with both simple and complex language structures. In addition, techniques from the group of Explainable AI such as LIME were applied to make the processes of sentiment classification more transparent and interpretable. These methods shall allow the users to comprehend which portion of any words and phrase in light of overall sentiment is contributing positive, thereby building trust in automated news classification. Consequently, this study adds to this emergent domain with a focus on e-paper news articles and proposes showcasing the part of Explainable AI in promoting transparency, trust, and fairness in automated news sentiment detection. It is highly relevant to understand the sentiment of news articles on e-paper for both readers and publishers.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leveraging Local Interpretable Model-Agnostic Explanations (LIME) for Sentiment Analysis of News Articles

  • Pragya Tewari,
  • Harshit Verma,
  • Rishabh Jaiswal,
  • Anurag Singh Baghel

摘要

Digital media have grown at a remarkably fast pace, which in turn has increased the consumption of news through e-papers. In such a scenario, the ability to quickly estimate the sentiment of a news article becomes highly valuable. This research report encompasses a broad study on sentiment analysis related to e-paper news articles, by applying techniques in NLP, sentiment analysis models classify news articles into undertones that are either positive, negative, or neutral. The project begins with collecting a long and wide dataset of e-paper news articles, covering a wide range of categories, sources, and different timelines. Further sentiment analysis will hence be derived from the dataset. Preprocessing the dataset beforehand through NLP is done prior to the sentiment analysis. These steps, including text cleaning, stop word removal, and stemming, are involved in this cleaning process to further improve the quality of the dataset. We have chosen logistic regression over the other machine learning models due to its effectiveness in text classification tasks, its ability to model probabilities, and its strong performance with both simple and complex language structures. In addition, techniques from the group of Explainable AI such as LIME were applied to make the processes of sentiment classification more transparent and interpretable. These methods shall allow the users to comprehend which portion of any words and phrase in light of overall sentiment is contributing positive, thereby building trust in automated news classification. Consequently, this study adds to this emergent domain with a focus on e-paper news articles and proposes showcasing the part of Explainable AI in promoting transparency, trust, and fairness in automated news sentiment detection. It is highly relevant to understand the sentiment of news articles on e-paper for both readers and publishers.