<p>This study proposes an efficient fine art image classification method integrating lightweight deep learning to address the limitations of low efficiency and poor generalization in art image classification tasks. The approach designs a lightweight hybrid network, MobileNet-Transformer Hybrid (MTH), that combines depthwise separable convolution with multi-head self-attention mechanisms to achieve efficient fusion of local details and global semantics. A dynamic channel-spatial attention module (DCSAM) adaptively enhances style-sensitive feature responses, while a cross-style feature transfer (CSFT) framework employs contrastive learning to align different style distributions and improve model robustness. Experiments conducted on the ArtBench-10 and WikiArt datasets validate the model’s performance in both classification accuracy and computational efficiency. The results demonstrate: (1) The proposed method enhances local style-discriminative features such as brushstrokes and colors through the DCSAM module, effectively alleviating the misjudgment problem of similar artistic styles. Among them, the CSFT framework constrains the cross-style feature distance through contrastive loss, improving the generalization of rare styles in long-tailed data. (2) The parameter count of the proposed model (1.2&#xa0;M) is 14.8% of that of EfficientNetV2-S (8.1&#xa0;million (M)) and 4.2% of that of Swin Tiny (28.3&#xa0;M). Under this lightweight condition, the proposed model still maintains an ArtBench-10 classification accuracy of 85.2%. Overall, this study provides an efficient solution for art design automation and cultural heritage digitization, offering theoretical innovation and practical application value.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fine art image classification and design methods integrating lightweight deep learning

  • Kexiang Ma,
  • SungWon Lee,
  • Xiaopeng Ma,
  • Hui Chen

摘要

This study proposes an efficient fine art image classification method integrating lightweight deep learning to address the limitations of low efficiency and poor generalization in art image classification tasks. The approach designs a lightweight hybrid network, MobileNet-Transformer Hybrid (MTH), that combines depthwise separable convolution with multi-head self-attention mechanisms to achieve efficient fusion of local details and global semantics. A dynamic channel-spatial attention module (DCSAM) adaptively enhances style-sensitive feature responses, while a cross-style feature transfer (CSFT) framework employs contrastive learning to align different style distributions and improve model robustness. Experiments conducted on the ArtBench-10 and WikiArt datasets validate the model’s performance in both classification accuracy and computational efficiency. The results demonstrate: (1) The proposed method enhances local style-discriminative features such as brushstrokes and colors through the DCSAM module, effectively alleviating the misjudgment problem of similar artistic styles. Among them, the CSFT framework constrains the cross-style feature distance through contrastive loss, improving the generalization of rare styles in long-tailed data. (2) The parameter count of the proposed model (1.2 M) is 14.8% of that of EfficientNetV2-S (8.1 million (M)) and 4.2% of that of Swin Tiny (28.3 M). Under this lightweight condition, the proposed model still maintains an ArtBench-10 classification accuracy of 85.2%. Overall, this study provides an efficient solution for art design automation and cultural heritage digitization, offering theoretical innovation and practical application value.