Context <p>Comics, as a form of storytelling, integrate visual art and narrative text to convey emotions, presenting unique challenges for automated emotion detection. Existing methods tend to focus on either textual or visual data alone, limiting their effectiveness in fully understanding the emotions.</p> Objective <p>This study aims to develop a multimodal fusion model that combines textual and visual features for accurate emotion detection in comic panels. The model seeks to overcome the limitations of unimodal approaches by simultaneously leveraging dialogue and facial expressions to better capture the emotional depth of characters.</p> Method <p>The proposed model uses a&#xa0;variant of the Bidirectional Encoder Representations from Transformers (RoBERTa) for textual analysis and Vision Transformers (ViT) for visual feature extraction. A custom dataset of 10,000–12,000 comic panels was utilized, covering diverse genres and emotions. The model's cross-attention mechanism dynamically balances the contribution of text and visuals based on the emotional context.</p> Results <p>The proposed multimodal emotion detection model demonstrated superior performance, achieving an overall accuracy of 91.8%, significantly outperforming unimodal models that relied on either text or visual data. The model accurately classified five distinct emotion categories: Happy (94.8%), Sad (92.5%), Angry (90.2%), Surprise (91.8%), and Neutral (89.7%). ROC curve analysis further validated the model’s discriminative power, with the highest AUC value for Happy (0.97) and the lowest for Neutral (0.89), highlighting areas for improvement in subtle emotion detection. Additionally, the model demonstrated scalable inference times suitable for real-time applications, with a slight increase in computational cost compared to unimodal models.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leveraging multimodal fusion for emotion detection in comics: integrating text and visual cues with RoBERTa and vision transformers

  • Rishu,
  • Vinay Kukreja

摘要

Context

Comics, as a form of storytelling, integrate visual art and narrative text to convey emotions, presenting unique challenges for automated emotion detection. Existing methods tend to focus on either textual or visual data alone, limiting their effectiveness in fully understanding the emotions.

Objective

This study aims to develop a multimodal fusion model that combines textual and visual features for accurate emotion detection in comic panels. The model seeks to overcome the limitations of unimodal approaches by simultaneously leveraging dialogue and facial expressions to better capture the emotional depth of characters.

Method

The proposed model uses a variant of the Bidirectional Encoder Representations from Transformers (RoBERTa) for textual analysis and Vision Transformers (ViT) for visual feature extraction. A custom dataset of 10,000–12,000 comic panels was utilized, covering diverse genres and emotions. The model's cross-attention mechanism dynamically balances the contribution of text and visuals based on the emotional context.

Results

The proposed multimodal emotion detection model demonstrated superior performance, achieving an overall accuracy of 91.8%, significantly outperforming unimodal models that relied on either text or visual data. The model accurately classified five distinct emotion categories: Happy (94.8%), Sad (92.5%), Angry (90.2%), Surprise (91.8%), and Neutral (89.7%). ROC curve analysis further validated the model’s discriminative power, with the highest AUC value for Happy (0.97) and the lowest for Neutral (0.89), highlighting areas for improvement in subtle emotion detection. Additionally, the model demonstrated scalable inference times suitable for real-time applications, with a slight increase in computational cost compared to unimodal models.