Performance of Sentiment Analysis APIs on Political Opinion Polling
摘要
Social media, due to its deep use throughout the United States, has the potential to supply accurate opinion polling. This study aims to replicate job approval polls conducted by professional pollsters through the utilization of sentiment analysis. A Kaggle dataset of Twitter messages from the end of the 2020 United States Election was selected and prepared. The sentiment of each tweet within this dataset was classified by language models created by cloud providers and accessed through APIs. We used two evaluation methods: hypothesis testing and confusion matrix accuracy. Regarding the hypothesis testing, the sentiment classification proportions were evaluated against job approval data aggregated from professional pollsters and did not perform well enough to accept the null hypothesis of no independence. Regarding the confusion matrix accuracy, the classifiers were evaluated against two sets of manually labeled tweets: general sentiment and political sentiment. The classifiers performed reasonably with general sentiment but poorly for political sentiment.