Can Machine Learning Models Recognise Emotions, Particularly Neutral, Better Than Humans?
摘要
Audio and visual data play vital roles in emotion recognition, with machine learning (ML) methods like SVMs and deep neural networks excelling in inferring human emotions. This study compared a state-of-the-art ML model’s performance in each modality to human performance. It also examined items frequently labelled as ‘neutral’, comparing ML and human results. Utilising the CREMA-D dataset, CNN-LSTM and SVM models were trained for visual-only and audio-only data. Evaluation included matching and non-matching test sets, and ML models consistently outperformed humans, especially on the latter. The CNN-LSTM achieved 80.8% accuracy on matching visual data compared to human accuracy of 75.9%, and 39.0% versus 19.4% on non-matching data. Similarly, the SVM model scored 81.0% and 34.3% accuracy on matching and non-matching audio data, surpassing human accuracy of 68.9% and 17.9%. These results underscore ML’s superiority in monomodal emotion recognition, particularly for challenging emotions. Implications for emotion recognition research and ethical concerns are inferred.