User Experience Optimization of Multimedia Art Interaction System Based on Affective Computing Algorithm
摘要
With the widespread application of human-computer interaction technology in digital art, immersive experience and intelligent display, traditional multimedia systems have problems such as low recognition accuracy, long feedback delay and incoherent experience in perceiving user emotions and dynamic feedback. To this end, this paper proposes a multimedia art interaction system that integrates emotional computing and multimodal deep learning to achieve high-precision collection and recognition of multimodal emotional information such as user visual expressions, voice intonation, behavioral trajectories and language texts. In the system structure, the visual channel uses the ResNet-based CNN (Convolutional Neural Network) model for facial expression recognition, the voice and behavior channels build a temporal emotion modeling module based on LSTM (Long Short-Term Memory), and the text channel uses the BERT (Bidirectional Encoder Representations from Transformers) semantic understanding model to extract user subjective emotions. In the emotion fusion stage, Transformer is used for modal alignment and cross-modal attention allocation to ensure that the fused emotional state is more in line with the actual user experience. The experimental results show that explicit emotions such as “happy” and “surprised” have a high recognition effect in CNN and BERT (for example, CNN recognizes happiness at 0.83/0.81, and BERT is 0.86/0.84), but the fusion model further improves to 0.92/0.91 and 0.90/0.88, reflecting the role of modal complementarity in model enhancement.