EmotiFusion-7: a multimodal framework for emotion recognition in comics using vision transformers and BERT
摘要
The proposed study presents an EmotiFusion-7 multimodal framework for emotion recognition using a self-created and compiled comic panel Custom-10K dataset comprised of 10,000 images. The proposed framework includes an attention-based fusion of vision transformers (ViTs) and bidirectional encoder representations from transformers (BERT) for extracting visual features and capturing textual semantics, respectively. Employing a Custom-10K dataset of annotated comic panels along with the fusion mechanism, the proposed work achieved a state-of-the-art performance with an exceptional overall accuracy of 92.3% and an area under the curve-receiver operating characteristic (AUC-ROC) score of 0.96 over seven emotions classes. In addition to this, the proposed study demonstrates superior performance in comparison with existing methods, conducting robustness testing and cross-domain evaluation. The proposed research establishes a new standard for multimodal emotion recognition along with a future scope for broader applications in visual storytelling and document image analysis.