Involving Society to Protect Society from Fake News and Disinformation: Crowdsourced Datasets and Text Reliability Assessment
摘要
Detecting fake information is a complex and multi-faceted challenge concerning data gathering, processing, analysis and detection, as well as countering fake content distribution. On the one hand, fake content is created intentionally and deliberately in order to mislead readers. On the other hand, it may take the form of unintentional disinformation resulting from careless selection and transmission of content. In this paper, we propose a novel approach to involve society in the fake news detection process. Moreover, we propose a methodology for text credibility assessment. In our research, we use natural language understanding (NLU) techniques to build an efficient ML pipeline that automates content evaluation. We report promising results on our new crowdsourced dataset in Polish language.