Analysis of Multimodal Social Media Data Utilizing VIT Base 16 and GPT-2 for Disaster Response
摘要
Multimedia systems, such as social media platforms, play a crucial role in disseminating vital information during calamities. This information is shared in various formats such as images, text, videos, audio, etc. Therefore, it becomes important to have a system that can identify multimodal data to classify relevant information. This paper proposes a new age classification method for multimodal data using advanced and improved transformer models, such as Vision Transformer and Generative Pre-trained Transformer 2, for image and text classification, respectively. These models were combined using an ensemble model (Random Forest Classifier), achieving an accuracy of 84.66% on the multimodal data. Furthermore, the proposed model demonstrates higher prediction accuracy compared to traditional Convolutional Neural Network (CNN) models which have an accuracy of 71.43%, exceeding it by 13.23%. A comparison with convolutional models is conducted to underscore the advantages of transformer models and to substantiate the necessity of the experiment. Our proposed classification model using Vision Transformer and GPT-2, along with an ensemble model, can be replicated by researchers in disaster management, humanitarian aid organizations, and social media platforms looking to filter and prioritize information during emergencies.