A novel credibility dataset and annotation framework for COVID-19-related German tweets
摘要
Social media proved to be one of the most accessible and immediate sources for people to stay informed during the COVID-19 pandemic. However, it quickly became clear that it was important to distinguish between credible and non-credible information. For corresponding classifications, performed by non-domain experts as explicit annotations, the task of estimating the credibility of online content still remains a non-trivial one due to its subjective nature. We introduce a new approach to guiding the annotation of COVID-19-related tweets, written in the German language, where credibility is framed as informative and relevant content in social media regarding a predefined set of topics. We incorporate named entity annotations using an extended label set, targeted towards information extraction related to COVID-19. We present a curated dataset that consists of 643 tweets including 3591 entities and named entities. To the best of our knowledge, this dataset is the first to provide in-depth multi-layer annotations for COVID-19-related German texts. For the credibility annotation pipeline we achieve an inter-annotator agreement of 0.76 (Cohen’s Kappa) and for (named) entities 0.73 (Krippendorff’s Alpha). We conducted experiments with BERT and Llama models across all layers. The multilingual TwHIN models achieve the best results with an average 0.8 F1 score.