The spread of incorrect information is now a serious issue in this era of quick information distribution through various online platforms especially in the field of healthcare, necessitating the need of development of detecting fake news mechanisms. This research focuses on the critical issue of identifying fake news in the healthcare sector by leveraging a range of Machine Learning (ML) and Deep Learning (DL) models. Using Natural Language Processing (NLP) for text analysis, we conduct a thorough evaluation with various ML models such as XGBoost, Random Forest, Logistic Regression, Support Vector Machine, Decision Tree, and Multinomial Naive Bayes. Apart from the ML models, we also implemented Bi-directional Long Short-Term Memory (LSTM), a deep learning model. We trained and evaluated these models on a publicly available dataset of healthcare-related news articles, to understand the performance, accuracy, and robustness of each. Random forest, Support Vector Machine, and Logistic Regression have a similar F1-score of 0.93 with Logistic Regression showing the accuracy of 93.7%. Our results highlight the comparison of each approach, providing insightful information on how each can be used in actual scenarios. This study adds to the current efforts to identify fraudulent information, ultimately aiming to improve technology to provide better and reliable information in the healthcare sector.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fake News Detection in the Field of Healthcare

  • Yuvika Mishra,
  • Shlok Kohli,
  • Shiv Naresh Shivhare,
  • Thipendra P. Singh

摘要

The spread of incorrect information is now a serious issue in this era of quick information distribution through various online platforms especially in the field of healthcare, necessitating the need of development of detecting fake news mechanisms. This research focuses on the critical issue of identifying fake news in the healthcare sector by leveraging a range of Machine Learning (ML) and Deep Learning (DL) models. Using Natural Language Processing (NLP) for text analysis, we conduct a thorough evaluation with various ML models such as XGBoost, Random Forest, Logistic Regression, Support Vector Machine, Decision Tree, and Multinomial Naive Bayes. Apart from the ML models, we also implemented Bi-directional Long Short-Term Memory (LSTM), a deep learning model. We trained and evaluated these models on a publicly available dataset of healthcare-related news articles, to understand the performance, accuracy, and robustness of each. Random forest, Support Vector Machine, and Logistic Regression have a similar F1-score of 0.93 with Logistic Regression showing the accuracy of 93.7%. Our results highlight the comparison of each approach, providing insightful information on how each can be used in actual scenarios. This study adds to the current efforts to identify fraudulent information, ultimately aiming to improve technology to provide better and reliable information in the healthcare sector.