Image Aesthetic Quality Assessment (IAQA) aims to simulate users’ subjective visual perception to predict the aesthetic quality of images. Due to the complex diversity in users’ aesthetics, existing IAQA methods mainly focus on multimodal-based models. However, these methods heavily rely on the aesthetic descriptions of images, making it challenging to effectively evaluate the aesthetic quality of images lacking textual information. To address this issue, this paper proposes an image aesthetic quality assessment method based on generative text prompts. The proposed method can generate aesthetics-related textual prompts for images that perform multimodal learning. Specifically, we first leverage the Multimodal Large Language Model (MLLM) to generate textual prompts describing the aesthetic aspects of images, which can effectively reflect the aesthetic experience that images bring to users. Secondly, we propose a novel cross-modal fusion strategy to facilitate the feature fusion between the generated text prompts and images, which can produce a multimodal-based image aesthetic quality assessment model even without aesthetic descriptions of images. Experimental results on several IAQA databases show that the proposed method achieves state-of-the-art performance in the IAQA task.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Generative Text Prompts for Image Aesthetic Quality Assessment

  • Ju Shi,
  • Hancheng Zhu,
  • Rui Yao,
  • Zhiwen Shao,
  • Leida Li

摘要

Image Aesthetic Quality Assessment (IAQA) aims to simulate users’ subjective visual perception to predict the aesthetic quality of images. Due to the complex diversity in users’ aesthetics, existing IAQA methods mainly focus on multimodal-based models. However, these methods heavily rely on the aesthetic descriptions of images, making it challenging to effectively evaluate the aesthetic quality of images lacking textual information. To address this issue, this paper proposes an image aesthetic quality assessment method based on generative text prompts. The proposed method can generate aesthetics-related textual prompts for images that perform multimodal learning. Specifically, we first leverage the Multimodal Large Language Model (MLLM) to generate textual prompts describing the aesthetic aspects of images, which can effectively reflect the aesthetic experience that images bring to users. Secondly, we propose a novel cross-modal fusion strategy to facilitate the feature fusion between the generated text prompts and images, which can produce a multimodal-based image aesthetic quality assessment model even without aesthetic descriptions of images. Experimental results on several IAQA databases show that the proposed method achieves state-of-the-art performance in the IAQA task.