Using Non-textual Content of Tweets in Sentiment Analysis: A Data Pre-processing Approach
摘要
Sentiment analysis of Twitter, now known as X, is becoming commonplace in consumer behaviour and decision research. The role of tweets’ non-textual content, namely hashtags, emoticons, and URLs in detecting the sentiment, has received little attention. This study proposes a data pre-processing approach to integrate non-textual content into sentiment analysis. Using 1.3 million tweets about the 2017 Ecuadorian Presidential election, sentiment analysis is conducted before and after pre-processing the datasets, and the results compared against the election results and official exit polls. Pre-processing involves splitting hashtags into single words, replacing emoticons and emoji with sentiment-conveying words, and replacing URLs with the sentiment positive, negative, or neutral. The results showed that the mean error using Twitter data outperformed results from traditional polling firms when measuring vote intention, with a mean error of 7.6% in the first round, and 0.25% in the second. The approach can be of help to researchers and practitioners involved in electoral campaigns. This study raises awareness of the need for novel and automated data pre-processing methods for sentiment analysis.