<p>The success of the recent World Cup tournament has captured unprecedented global attention, leading to a substantial increase in online discourse that reflects the diverse sentiments and opinions of fans across the globe. Understanding this online conversation is essential, particularly in the context of significant sporting events such as the recent tournament held in Qatar from November 20 to December 18, 2022. However, access to social media data is often restricted, posing a significant obstacle to the comprehensive study of online sports discourse. To address this challenge and enhance the capabilities of the Computational Social Science research community, we present a large-scale dataset comprising tweets related to the World Cup. This multilingual dataset encompasses over 28 million posts from over 2.5 million unique users, documenting key events surrounding the tournament. Our dataset is meticulously curated and thoroughly documented, providing a valuable resource for researchers investigating critical social and scientific issues, including bot detection, sentiment analysis, hate speech and aggression identification, and the dissemination of misinformation. The dataset is publicly accessible at: <a href="https://data.mendeley.com/datasets/gw3mcnbkwr/2">https://data.mendeley.com/datasets/gw3mcnbkwr/2</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Tracking the global pulse: the first public twitter dataset from FIFA world cup

  • Kheir Eddine Daouadi,
  • Yaakoub Boualleg,
  • Oussama Guehairia,
  • Abdelmalik Taleb-Ahmed

摘要

The success of the recent World Cup tournament has captured unprecedented global attention, leading to a substantial increase in online discourse that reflects the diverse sentiments and opinions of fans across the globe. Understanding this online conversation is essential, particularly in the context of significant sporting events such as the recent tournament held in Qatar from November 20 to December 18, 2022. However, access to social media data is often restricted, posing a significant obstacle to the comprehensive study of online sports discourse. To address this challenge and enhance the capabilities of the Computational Social Science research community, we present a large-scale dataset comprising tweets related to the World Cup. This multilingual dataset encompasses over 28 million posts from over 2.5 million unique users, documenting key events surrounding the tournament. Our dataset is meticulously curated and thoroughly documented, providing a valuable resource for researchers investigating critical social and scientific issues, including bot detection, sentiment analysis, hate speech and aggression identification, and the dissemination of misinformation. The dataset is publicly accessible at: https://data.mendeley.com/datasets/gw3mcnbkwr/2.