The rapid growth of digital platforms has revolutionized information dissemination, but it has also facilitated the spread of fake news, often amplified through biased journalism and social media. This paper addresses the challenge of fake news detection by proposing a robust methodology that categorizes news articles as genuine or fraudulent. Leveraging advanced natural language processing (NLP) techniques, such as TF-IDF and Bag of Words (BoW), the study employs six machine learning models—Logistic Regression, Support Vector Machines (SVM), Random Forests, Naïve Bayes, Decision Trees, and Neural Networks—to classify news articles. The study presents innovative feature extraction techniques, such as the Linguistic Inquiry and Word Count (LIWC2015) tool, to capture critical linguistic properties, including sentiment ratios and grammatical parts. Additionally, dimensionality reduction techniques like t-SNE and PCA enhance model performance by reducing noise and maintaining global patterns. The findings show that SVM, MLP, and logistic regression classifiers continuously produce high recall, accuracy, precision, and F1 scores, especially when adding speaker and party information. At the same time, Naïve Bayes performs poorly because it cannot handle the complexity of the dataset. The study places a strong emphasis on privacy and cultural diversity as ethical factors in the detection of fake news. The results show that the reliability of false news identification can be significantly increased by integrating complex feature sets with cutting-edge machine learning algorithms.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Techniques for Detecting False Information on Social Media to Strengthen Cybersecurity

  • Prabhat Kumar Sahu,
  • Smita Rath,
  • Alakananda Tripathy,
  • Rashmi Rani Patro,
  • Sangam Malla

摘要

The rapid growth of digital platforms has revolutionized information dissemination, but it has also facilitated the spread of fake news, often amplified through biased journalism and social media. This paper addresses the challenge of fake news detection by proposing a robust methodology that categorizes news articles as genuine or fraudulent. Leveraging advanced natural language processing (NLP) techniques, such as TF-IDF and Bag of Words (BoW), the study employs six machine learning models—Logistic Regression, Support Vector Machines (SVM), Random Forests, Naïve Bayes, Decision Trees, and Neural Networks—to classify news articles. The study presents innovative feature extraction techniques, such as the Linguistic Inquiry and Word Count (LIWC2015) tool, to capture critical linguistic properties, including sentiment ratios and grammatical parts. Additionally, dimensionality reduction techniques like t-SNE and PCA enhance model performance by reducing noise and maintaining global patterns. The findings show that SVM, MLP, and logistic regression classifiers continuously produce high recall, accuracy, precision, and F1 scores, especially when adding speaker and party information. At the same time, Naïve Bayes performs poorly because it cannot handle the complexity of the dataset. The study places a strong emphasis on privacy and cultural diversity as ethical factors in the detection of fake news. The results show that the reliability of false news identification can be significantly increased by integrating complex feature sets with cutting-edge machine learning algorithms.