Random-coupled neural network with binary light spectrum optimization based context-aware multimodal emotion recognition using audio, video, and text feature fusion approach
摘要
Emotions play a vital role in shaping human behavior, making accurate emotion recognition crucial in areas like social robotics, mental health diagnostics, and human–computer interaction. Existing multimodal systems often suffer from cross-modal interference, poor feature extraction, and weak integration across text, audio, and video data. This paper proposes a novel framework Random-Coupled Neural Network with Binary Light Spectrum Optimization (BLSO) for Context-Aware Multimodal Emotion Recognition that addresses these challenges through tailored preprocessing and advanced feature fusion. Text is processed with tokenization and lemmatization, audio with noise reduction and voice activity detection, and video with frame extraction and face normalization. Features are extracted using an Elastic Decision Transformer for text, Short-Time Fourier and Continuous Wavelet Transforms for audio, and Local Directional Structural Patterns (LDSP) for video. A Context-Aware Attentive Multilevel Feature Fusion (CAMFF) network integrates these features, while a Random-Coupled Feedback Attention Neural Network (RCFANN) optimized via BLSO performs classification. Experiments on the IEMOCAP dataset demonstrate the proposed system achieves over 99% accuracy, precision, recall, and F1-score, outperforming existing multimodal fusion methods by 4–6%.