Enhancing Sentiment Analysis Accuracy Through Multimodal Data Fusion: A Deep Learning Approach
摘要
Sentiment analysis’s (SA) wide range of potential applications is making it more and more popular. Because prior study has focused on the SA of single modalities, it may be challenging to manage social media data with many modalities like words or images. However, after feature extraction, redundant data can be easily added to feature fusion, which has an impact on the final feature representation. Furthermore, rather than examining the intricate relationships between the two modalities, the majority of multimodal research has focused on simply integrating them, which has resulted in inadequate sentiment analysis findings. Motivated by this, we introduce multi-model data fusion (MMDF), a unique audio–video sentiment classification model that efficiently extracts pertinent information and the basic connection among the audio and video material by utilizing a mixed fusion framework for sentiment analysis (SA). This study compares the performance of models based on deep learning (DL) for sentiment analysis and multimodal data fusion. Our experimental results show that multi-task learning produces the best results across all modalities: for the classification of three emotional states using combinations of audio-visual, DEAP-visual, and DEAP-audio data, respectively; the results are 75.41, 68.33, and 78.75%.