Social media are powerful platforms for sharing news and opinions, but their use may also facilitate the rapid spread of false information. Supervised algorithms for fake news detection, ranging from traditional machine learning to deep learning methods, rely heavily on the quality of training data. This work proposes a semi-automatic dataset creation technique to support the validation of fake news detection algorithms. The system aims to generate annotated datasets containing tweets with detailed information (e.g., text, user data, and relationships between users and data) and to determine the ground truth by assigning truth values to each tweet in the dataset. The tests conducted showed the effectiveness of the annotation process as well as the limitations and strengths of the approaches considered.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Annotated Dataset Creation for Fake News Detection on Online Social Networks

  • Farwa Batool,
  • Giuseppe Lo Re,
  • Marco Morana

摘要

Social media are powerful platforms for sharing news and opinions, but their use may also facilitate the rapid spread of false information. Supervised algorithms for fake news detection, ranging from traditional machine learning to deep learning methods, rely heavily on the quality of training data. This work proposes a semi-automatic dataset creation technique to support the validation of fake news detection algorithms. The system aims to generate annotated datasets containing tweets with detailed information (e.g., text, user data, and relationships between users and data) and to determine the ground truth by assigning truth values to each tweet in the dataset. The tests conducted showed the effectiveness of the annotation process as well as the limitations and strengths of the approaches considered.