Weakly Supervised Learning for Textual Propaganda Detection on Twitter and a Multi-Label Tweet Propaganda Dataset
摘要
With the surge of the Internet and growing influence of social media, disinformation technique such as propaganda is increasingly getting used by both proponents and opponents of any modern-day social movement. Hence propaganda detection has emerged as an important task in modern disinformation research. However, the existing linguistic propaganda research is mostly limited to news articles, and to the best of our knowledge, there is no large labeled social media propaganda dataset available to date. The anti-CAA movement (2019) in India showcased the prominent role of social media in modern social movements and researchers have also conclusively proved the presence of inauthentic users on both sides of the discourse pursuing propagandistic goals. In this paper, we present a weakly supervised learning methodology to address the incomplete supervision challenge in supervised learning and present a large tweet-based dataset for textual propaganda detection. The data set consists of tweet ids, hashtags used, and corresponding propaganda techniques. The efficacy of the dataset is proved by building a transformer-based propaganda detection model that shows an accuracy of 87% on manually labeled test data. The dataset has been published to facilitate future research and automatic linguistic propaganda detection on Twitter.