错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fake news detection and corpus establishment from comment data for social network posts

  • Yean-Fu Wen,
  • Wen-Hsin Chang,
  • Chih-Chien Wang,
  • Kuo-Lin Yang

摘要

The objective of this study was to construct an association corpus of misinformation, disinformation, or fake news (MDFN) and categorize instances of MDFN in terms of severity to assess its effects. The hashtag function on social media platforms, such as Facebook, was used to search for mandarin posts, and the probability of these posts being MDFN was evaluated on the basis of user comments, which is a community-based approaches (CA). The posts were further categorized by MDFN severity. A corpus was established with Chinese keywords indicating MDFN probability and severity. Various MDFN probability and severity calculation schemes based on semantic analysis of words, sentences, and comments were developed. Two methods for evaluating the accuracy of the proposed schemes were adopted that are (1) the detected MDFN posts were sent to fact-checking organizations for validation and (2) a questionnaire was conducted, with respondents asked to read a post and indicate whether they believed it was MDFN. The contributions of this study are as follows: (1) it provides a mandarin corpus and detection candidate items of MDFN as an upstream resource for fact-checking organizations to more efficiently collect MDFN, (2) it provides a words-and-sentences database for semantic analysis, which is belong to word embeddings, and recursive detection of MDFN, and (3) it categorizes identified MDFN by severity to raise awareness among the public and protect them from the influence of MDFN. This study is the first to use comments on posts, which is a type of document embeddings, to simultaneously detect whether a post is MDFN and the severity of that MDFN; the method could enhance fact-checking efficiency.