³Comparative Analysis of Audio–Video Multimodal Methods for Emotion Recognition
摘要
Research into human–computer interaction aims to provide a seamless interface by considering the humans emotional status. Human emotions, including facial expressions, signals from the body, and neuroimaging methods can be used to describe techniques. For better classification accuracy, multimodal affective computing systems and unimodal solutions are researched. This chapter presents a systematic and comparative study of existing multimodal methods for human emotion recognition through facial expressions and speech from the perspectives of multimodal datasets, detection methods, and multimodal fusion methods. In this paper, we wrap up the current influx of works on multimodal emotion detection and offer advice to researchers interested in understanding both cutting-edge and conventional multimodal emotion recognition techniques.