The increasing concern about the spread of information and its negative impact on trust and democratic processes has led to the need for methods to identify fake news. Humans can only detect 54% of news accurately, with an additional 4% considered uncertain. This is especially crucial, during times such as the pandemic when inaccurate information can quickly spread. While machine learning algorithms have been considered for identifying news their widespread application is still limited. This study introduces a machine learning approach that combines techniques, including decision tree, support vector machines, random forest, logistic regression, Naive Bayes, and KNN. The model’s effectiveness is assessed by using an existing dataset aiming to achieve accuracy, without overfitting. The approach involves dealing with missing information columns and employing text preprocessing techniques. Important tactics such as K-fold cross-validation, TF-IDF vectorization, and principal component analysis (PCA) are utilized to maintain the model’s strength and practicality. The proposed decision tree model achieved a 99% accuracy rate on test data and 100% accuracy on the training set without signs of overfitting. This study contributes, toward automating news detection and offering a tool to combat misinformation in the age.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Advanced NLP Techniques for Detecting Online Fake News Using the Fake News Dataset

  • Roa’a Mohammedqasem,
  • Hayder Mohammedqasim,
  • Bilal A. Ozturk,
  • Isra Ben Brahim,
  • Heba Jadallah,
  • Najwa Alwesabi

摘要

The increasing concern about the spread of information and its negative impact on trust and democratic processes has led to the need for methods to identify fake news. Humans can only detect 54% of news accurately, with an additional 4% considered uncertain. This is especially crucial, during times such as the pandemic when inaccurate information can quickly spread. While machine learning algorithms have been considered for identifying news their widespread application is still limited. This study introduces a machine learning approach that combines techniques, including decision tree, support vector machines, random forest, logistic regression, Naive Bayes, and KNN. The model’s effectiveness is assessed by using an existing dataset aiming to achieve accuracy, without overfitting. The approach involves dealing with missing information columns and employing text preprocessing techniques. Important tactics such as K-fold cross-validation, TF-IDF vectorization, and principal component analysis (PCA) are utilized to maintain the model’s strength and practicality. The proposed decision tree model achieved a 99% accuracy rate on test data and 100% accuracy on the training set without signs of overfitting. This study contributes, toward automating news detection and offering a tool to combat misinformation in the age.