<p>Social media proved to be one of the most accessible and immediate sources for people to stay informed during the COVID-19 pandemic. However, it quickly became clear that it was important to distinguish between credible and non-credible information. For corresponding classifications, performed by non-domain experts as explicit annotations, the task of estimating the credibility of online content still remains a non-trivial one due to its subjective nature. We introduce a new approach to guiding the annotation of COVID-19-related tweets, written in the German language, where credibility is framed as informative and relevant content in social media regarding a predefined set of topics. We incorporate named entity annotations using an extended label set, targeted towards information extraction related to COVID-19. We present a curated dataset that consists of 643 tweets including 3591 entities and named entities. To the best of our knowledge, this dataset is the first to provide in-depth multi-layer annotations for COVID-19-related German texts. For the credibility annotation pipeline we achieve an inter-annotator agreement of 0.76 (Cohen’s Kappa) and for (named) entities 0.73 (Krippendorff’s Alpha). We conducted experiments with BERT and Llama models across all layers. The multilingual TwHIN models achieve the best results with an average 0.8 F<sub>1</sub> score.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A novel credibility dataset and annotation framework for COVID-19-related German tweets

  • Karolina Zaczynska,
  • Elena Leitner,
  • Georg Rehm

摘要

Social media proved to be one of the most accessible and immediate sources for people to stay informed during the COVID-19 pandemic. However, it quickly became clear that it was important to distinguish between credible and non-credible information. For corresponding classifications, performed by non-domain experts as explicit annotations, the task of estimating the credibility of online content still remains a non-trivial one due to its subjective nature. We introduce a new approach to guiding the annotation of COVID-19-related tweets, written in the German language, where credibility is framed as informative and relevant content in social media regarding a predefined set of topics. We incorporate named entity annotations using an extended label set, targeted towards information extraction related to COVID-19. We present a curated dataset that consists of 643 tweets including 3591 entities and named entities. To the best of our knowledge, this dataset is the first to provide in-depth multi-layer annotations for COVID-19-related German texts. For the credibility annotation pipeline we achieve an inter-annotator agreement of 0.76 (Cohen’s Kappa) and for (named) entities 0.73 (Krippendorff’s Alpha). We conducted experiments with BERT and Llama models across all layers. The multilingual TwHIN models achieve the best results with an average 0.8 F1 score.