<p>The proliferation of social media has led to a surge in emotional image messages, necessitating advancements in image-affective computing. This field aims to recognize emotional information within images, with emotion classification being a pivotal area of research. However, due to the inherent uncertainty and ambiguity in emotion interpretation, conventional approaches relying on Convolutional Neural Networks (CNNs) often exhibit limited effectiveness. To address these challenges, this study introduces the EmoViTResNet architecture, a novel hybrid framework that synergistically integrates Vision Transformer (ViT) networks with Residual Networks (ResNet). The proposed model enhances deep feature representation by combining global attention mechanisms from ViT with local feature extraction and soft thresholding from ResNet, thereby improving emotion classification performance. Two classification networks, Res-ViT and the more advanced EmoViTResNet, were developed and evaluated across four distinct classification tasks, encompassing both multi-class and binary emotion classification on the FI and EmotionROI datasets. The EmoViTResNet model achieved outstanding accuracy scores of 94.58%, and 92.73%, respectively. These results demonstrate the model’s robustness, improved generalization, and the effectiveness of combining ViT with residual learning for image emotion classification in visual media.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Robust Transformer–Residual Hybrid Framework with Soft Thresholding for High-Performance Image Emotion Classification

  • Ala’a R. Al-Shamasneh,
  • Faten Khalid Karim,
  • Yu Wang

摘要

The proliferation of social media has led to a surge in emotional image messages, necessitating advancements in image-affective computing. This field aims to recognize emotional information within images, with emotion classification being a pivotal area of research. However, due to the inherent uncertainty and ambiguity in emotion interpretation, conventional approaches relying on Convolutional Neural Networks (CNNs) often exhibit limited effectiveness. To address these challenges, this study introduces the EmoViTResNet architecture, a novel hybrid framework that synergistically integrates Vision Transformer (ViT) networks with Residual Networks (ResNet). The proposed model enhances deep feature representation by combining global attention mechanisms from ViT with local feature extraction and soft thresholding from ResNet, thereby improving emotion classification performance. Two classification networks, Res-ViT and the more advanced EmoViTResNet, were developed and evaluated across four distinct classification tasks, encompassing both multi-class and binary emotion classification on the FI and EmotionROI datasets. The EmoViTResNet model achieved outstanding accuracy scores of 94.58%, and 92.73%, respectively. These results demonstrate the model’s robustness, improved generalization, and the effectiveness of combining ViT with residual learning for image emotion classification in visual media.