Cross-Modal Fusion and Data Transformation Techniques: Adapting BERT-ResNet-50 Fusion to MELD Dataset and Extending to DistilBERT-MobileNetV2 Fusion
摘要
Multimodal sentiment analysis has become a prominent field of interest for researchers around the world. This growing interest is driven by the increased availability of data for research purposes. Additionally, focusing on multiple types of data like text, audio, and images has led to an increase in the accuracy of sentiment analysis models. In our paper, we build upon existing research that combines BERT and ResNet-50 models with sentence embedding and video encoding and test its validity on the MELD dataset. Moreover, we utilize these data transformation techniques along with two newer models, DistilBERT and MobileNetV2, to see the impact of their lightweight nature on accuracy parameters. Our findings are consistent with our expectations, as the accuracy of the lightweight models is slightly lower than that of their more robust counterparts, BERT, and ResNet-50. Nevertheless, our successful application of multimodal sentiment analysis techniques on the MELD dataset underscores the viability of these methods.