With the rapid development of artificial intelligence and computer vision technology, Image Aesthetic Quality Assessment (IAQA) has become an important topic in the field of computer vision. Traditional IAQA methods often rely on Convolutional Neural Networks (CNNs), but these methods have limitations when processing images with different resolutions and aspect ratios, and they typically require fixed input sizes. To address this issue, and considering computational energy consumption, this paper proposes an Image Aesthetic Quality Assessment method based on Multi-Scale Fusion Spiking Neural Network Vision Transformer (MUSVT). This method effectively handles full-size images by combining multi-scale input representations with a pulse-integrated Transformer architecture, overcoming the limitations of fixed input sizes and reducing computational energy consumption. The MUSVT model can capture detailed information from multiple scales, and by introducing a hash-based 2D spatial embedding module, it solves the spatial alignment problem of different scale inputs, further improving the accuracy and robustness of the image aesthetic quality evaluation. Experimental results show that the MUSVT model performs excellently on the AVA and KonIQ-10k image aesthetic quality evaluation datasets, demonstrating the potential of this method in multi-resolution image processing and IAQA.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Image Aesthetic Quality Assessment Method Based on Multi-scale Fusion Spiking Neural Network Vision Transformer

  • Yang Liu,
  • Shuo Zhang,
  • Tengyu Fan,
  • Minghu Fan

摘要

With the rapid development of artificial intelligence and computer vision technology, Image Aesthetic Quality Assessment (IAQA) has become an important topic in the field of computer vision. Traditional IAQA methods often rely on Convolutional Neural Networks (CNNs), but these methods have limitations when processing images with different resolutions and aspect ratios, and they typically require fixed input sizes. To address this issue, and considering computational energy consumption, this paper proposes an Image Aesthetic Quality Assessment method based on Multi-Scale Fusion Spiking Neural Network Vision Transformer (MUSVT). This method effectively handles full-size images by combining multi-scale input representations with a pulse-integrated Transformer architecture, overcoming the limitations of fixed input sizes and reducing computational energy consumption. The MUSVT model can capture detailed information from multiple scales, and by introducing a hash-based 2D spatial embedding module, it solves the spatial alignment problem of different scale inputs, further improving the accuracy and robustness of the image aesthetic quality evaluation. Experimental results show that the MUSVT model performs excellently on the AVA and KonIQ-10k image aesthetic quality evaluation datasets, demonstrating the potential of this method in multi-resolution image processing and IAQA.