Social media is a popular source of information for the public and therefore it is critical that health officials are able to engage with the public in an effective way on these platforms. We were interested in the sentiment reflected in tweets from the public responding to messaging from these agencies. This paper reports on a comparison of different machine learning models for use in the multi-classification of the sentiment of such tweets. Tweets were first collected and manually labeled into seven different sentiment classes. The labelled tweets were then processed to form a core dataset. Several machine learning models were compared using this dataset and augmentation of the dataset, including using up-scaling and the use of “artificial” tweets constructed from the core dataset. The paper reports on the techniques used during preprocessing, augmentation of the dataset, the machine learning models, and the results obtained.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparing Models for Sentiment Analysis of Tweets in Response to Public Health Announcements During the Pandemic

  • Kristina Kacmarova,
  • Heather McPhail,
  • Anita Kothari,
  • Lesley James,
  • Lyndsay Foisey,
  • Lorie Donelle,
  • Michael Bauer

摘要

Social media is a popular source of information for the public and therefore it is critical that health officials are able to engage with the public in an effective way on these platforms. We were interested in the sentiment reflected in tweets from the public responding to messaging from these agencies. This paper reports on a comparison of different machine learning models for use in the multi-classification of the sentiment of such tweets. Tweets were first collected and manually labeled into seven different sentiment classes. The labelled tweets were then processed to form a core dataset. Several machine learning models were compared using this dataset and augmentation of the dataset, including using up-scaling and the use of “artificial” tweets constructed from the core dataset. The paper reports on the techniques used during preprocessing, augmentation of the dataset, the machine learning models, and the results obtained.