Comparative Evaluation of Sentiment Analysis Methods: Manual Annotation, Crowd-Coding, Dictionaries, and Machine Learning
摘要
This study investigates the validity of different sentiment analysis approaches by comparing manual annotation, crowd-coding, various dictionary methods, and machine learning algorithms using a Dutch language dataset of economic news headlines. Through a detailed analysis, we assess the performance of these methods against a gold-standard set of manually coded headlines. Our results show that human annotation, whether by trained coders or crowd coders, remains the most reliable method for sentiment analysis, achieving the highest accuracy. Machine learning, especially deep learning techniques, significantly outperforms dictionary-based approaches but still lags behind human performance. Our findings highlight the limitations of dictionary approaches and suggest that machine learning models require further validation before use in practical applications. We propose a step-by-step framework for implementing sentiment analysis to ensure both efficiency and validity, emphasizing the importance of human oversight in automated sentiment classification projects.