错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The Fake News Recognition Method Based on Naïve Bayes with Improved TF-IDF Algorithm

  • Liudmyla Mishchenkо,
  • Iryna Klymenkо,
  • Valentyna Tkachenko

摘要

The article introduces an innovative approach for recognition of fake news and misinformation on the internet through the utilization of Natural Language Processing (NLP). By integrating an enhanced TF-IDF (Term Frequency-Inverse Document Frequency) algorithm with the Naïve Bayes classifier. The proposed method aims to devise a method and tools for computer analysis of large volumes of dynamically changing textual data on the Internet and prompt recognition of deceptive information propagated through fake news. The study leverages the inherent features of NLP to analyze linguistic patterns and contextual cues, enhancing the accuracy of classification. The improved TF-IDF algorithm augments feature weighting, considering word frequency distribution and category distribution information to achieve a more precise assessment of feature significance. An analysis of the results of employing a fake news detection method on the Internet based on NLP and the integration of an intelligent Naїve Bayes classifier for predicting words in textual data has been conducted. As a result, the accuracy of binary classification of news texts has been increased from 85% to 93%. In comparison with a known NLP strategy based on the use of the TF-IDF statistical measure without integrating the Bayesian classifier, the classification accuracy ranges from 80% to 90%. Thus, the improved method has allowed for an average increase of 2.5% in the efficiency of fake news detection on the Internet. This research contributes to the evolving landscape of NLP-based techniques for combating the proliferation of fake news on the internet.