<p>Inaccurate information that is purposefully spread for a certain goal known as fake news related to COVID-19 tweets posing threats to public health, harm to the society. But creating a reliable method to identify news of this kind is difficult, particularly when the news is released combining factual and fictitious material. The identification of fake news with COVID-19 is challenging task in Natural Language Processing (NLP). This study uses Machine Learning (ML) models such as Support Vector Machine (SVM), Random Forest (RF), Gaussian Naıve Bayes (NB), Logistic Regression (LR) and Deep Learning (DL) models with several architectures such as Long-short Term memory (LSTM), Bidirectional LSTM with one-hot encoding representation and pre-trained word embeddings such as Word2Vec, GloVe and FastText, trained using 14,571 records on DataCOVID19 dataset obtained by combining datasets from three different sources available in Kaggle and GitHub repository. The tests on a real-world COVID-19 news dataset show that Bi-LSTM model improves performance when built over various embedding techniques such as Word2Vec, GloVe, FastText and One-hot encoding representation achieving highest accuracy score. The findings emphasise the potential of deep learning methods for identifying false news and have important ramifications for curbing the dissemination of false information about COVID-19 tweets known as Fake News Detection (FND).</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Detecting Misinformation in COVID-19 Content: A Machine Learning and Deep Learning Approach with Word Embeddings

  • Arati Chabukswar,
  • P. Deepa Shenoy,
  • K. R. Venugopal

摘要

Inaccurate information that is purposefully spread for a certain goal known as fake news related to COVID-19 tweets posing threats to public health, harm to the society. But creating a reliable method to identify news of this kind is difficult, particularly when the news is released combining factual and fictitious material. The identification of fake news with COVID-19 is challenging task in Natural Language Processing (NLP). This study uses Machine Learning (ML) models such as Support Vector Machine (SVM), Random Forest (RF), Gaussian Naıve Bayes (NB), Logistic Regression (LR) and Deep Learning (DL) models with several architectures such as Long-short Term memory (LSTM), Bidirectional LSTM with one-hot encoding representation and pre-trained word embeddings such as Word2Vec, GloVe and FastText, trained using 14,571 records on DataCOVID19 dataset obtained by combining datasets from three different sources available in Kaggle and GitHub repository. The tests on a real-world COVID-19 news dataset show that Bi-LSTM model improves performance when built over various embedding techniques such as Word2Vec, GloVe, FastText and One-hot encoding representation achieving highest accuracy score. The findings emphasise the potential of deep learning methods for identifying false news and have important ramifications for curbing the dissemination of false information about COVID-19 tweets known as Fake News Detection (FND).