Public perception towards deepfake through topic modelling and sentiment analysis of social media data
摘要
Since its inception in 2014, Deepfake technology has become prevalent across various sectors, provoking significant controversies and concerns. This study analyses 17,720 Deepfake-related posts and comments on the social media, Reddit, using topic modelling with Latent Dirichlet Allocation and sentiment analysis with TextBlob and VADER methods. Public discussions focus on eleven topics, categorised into two themes: Culture and Entertainment, Legal and Ethical Impacts. 47.0% of the public holds a positive attitude, while 36.8% are negative. The topic of Voice and Effects in Deepfakes has the highest proportion (59.3%) of positive sentiment, indicating public recognition of the creative allure of audio manipulation and voice synthesis by Deepfake. The topic of Abuse of Deepfakes in Adult Content draws the highest percentage of negative sentiment at 47.5%, reflecting social concern for the ethical and legal implications of non-consensual deepfake pornography and potential harm. Finally, it trains six machine learning models and three BERT-based models using the annotated negative data. Among these, the BERTweet model performs the best on the test data, achieving an accuracy of 87.03%. The finding suggests that public attitudes on the topics of Deepfake are divided, reflecting the complexity and contentiousness of the technology. While its innovative potential in entertainment is recognised, authenticity, legality and ethics should also be considered. The study reveals the differential impact of deepfakes on gender, especially when it comes to non-consensual pornography. This study underlines the balance of innovation and risks and provides valuable insights for policy-making, technological development, and future research.