Images can convey rich semantics and induce a wide range of emotions in viewers. However, predicting induced emotions from images can be challenging due to the subjective nature of emotions and the variability in how different viewers perceive them. Existing methods for image emotion prediction rely on neural networks to learn an image-to-emotion mapping. Such methods require large amounts of annotated training data to achieve good generalization. In this paper, we show that it is possible to train a model that generalizes better than state-of-the-art methods with significantly less data. Our method leverages the power of a pre-trained large multimodal model (LMM) with the addition of a shallow adapter module that transforms the LMM’s output embedding to a classification output. On three out of four benchmark datasets, our method outperforms the previous state of the art (SOTA) results by a significant margin, with one even showing around a 9% accuracy improvement. Additionally, our method achieves a new SOTA with only 20% of the data on these three datasets, and improves further using more data. On the fourth dataset, which is the smallest one, our method is on par with the SOTA. Moreover, our method can naturally provide human-readable intermediate results, which could serve as textual explanations of the classification outputs. The code is available at https://github.com/vimal-isi-edu/LMM_Emotion_Prediction .

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Large Multimodal Models Thrive with Little Data for Image Emotion Prediction

  • Peng He,
  • Mohamed Hussein,
  • Wael Abd Almageed

摘要

Images can convey rich semantics and induce a wide range of emotions in viewers. However, predicting induced emotions from images can be challenging due to the subjective nature of emotions and the variability in how different viewers perceive them. Existing methods for image emotion prediction rely on neural networks to learn an image-to-emotion mapping. Such methods require large amounts of annotated training data to achieve good generalization. In this paper, we show that it is possible to train a model that generalizes better than state-of-the-art methods with significantly less data. Our method leverages the power of a pre-trained large multimodal model (LMM) with the addition of a shallow adapter module that transforms the LMM’s output embedding to a classification output. On three out of four benchmark datasets, our method outperforms the previous state of the art (SOTA) results by a significant margin, with one even showing around a 9% accuracy improvement. Additionally, our method achieves a new SOTA with only 20% of the data on these three datasets, and improves further using more data. On the fourth dataset, which is the smallest one, our method is on par with the SOTA. Moreover, our method can naturally provide human-readable intermediate results, which could serve as textual explanations of the classification outputs. The code is available at https://github.com/vimal-isi-edu/LMM_Emotion_Prediction .