A Comprehensive Review of Emotion Recognition Using Multimedia Parameters: Challenges, Computational Intelligence and Applications
摘要
Advancement in human–computer interactions (HCI) demands optimized systems that depend on devising fast and reliable human emotion recognition algorithms. Images, audio, sensory physiological signals, and body gestures are some of the identified principal modalities that characterize human emotions. However, the fusion of different modalities (homogeneous and heterogeneous) towards emotion recognition is challenging research where the space of diversities in different modalities are explored. Apart from dimensionality reduction schemes, fitting feature extraction methods, and fusion through models are some challenges which impact the performance of HCI systems. With the capability of processing unstructured data and real-time system designs, various neural networks have been reported to be the preferable choice for development of a multimodal emotion recognition system. This paper presents a comprehensive review on emotion recognition works based on image, audio, and physiological modalities. It begins with a discussion on various emotions and their interpretation with the said modalities. Further, emotion recognition works have been explored in terms of established and crafted emotion features. Later features extraction techniques that range from general machine learning based schemes to the recent deep neural networks are presented. Consequently, the performance of various emotion recognition schemes has been reviewed for individual and multi modalities-based systems. This performance review goes higher for the models based on various modality fusions, i.e., audio–video, audio-text, and body gestures. The fusion schemes mentioned here have been analyzed keeping in mind the efficacy of modality combinations towards each human emotion recognition.