错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fake news detection models using the largest social media ground-truth dataset (TruthSeeker)

  • Maysa Khalil,
  • Mohammad Azzeh

摘要

Twitter is a powerful platform for communication and information sharing but is also susceptible to spreading false information. This false information has adverse consequences for society and can significantly impact public perception, decision-making, and political outcomes. Therefore, there is an urgent need to build a fake news detection system that can accurately catch false information before it is disseminated. Building such a system requires the existence of good quality and trustworthy labeled datasets. The limitations of the existing datasets are undeniable. Most of them are not updated to reflect the advanced generation patterns of the new fake news creators. Thanks to Truth Seeker research team, who offered a large-scale fake news dataset that was labeled based on Amazon Mechanical Turk. The dataset was collected between 2009 and 2022 and then validated according to a robust procedure to ensure its quality and reliability. However, the credibility and trustability of this dataset is still questionable. In this paper, we study and analyze the feasibility of building a fake news detection model based on deep learning using Truth seeker dataset. Mainly we investigated the impact of different text representation techniques on the accuracy of deep learning models. Also, we investigated the importance of hand-crafted features associated with the dataset in the final results. The results have shown that using truth seeker dataset show potential to help social media platforms in detecting fake news. on the other hand, using deep contextualized text representation produced more accurate results compared to word2vec and TF-IDF techniques. The impact of hand-crafted features on the final performance of deep learning models is often negligible, and it is suggested to be excluded from the final models.