The goal of arbitrary style transfer (AST) is to efficiently apply any artistic style to a content image. However, existing approaches are often hindered in pragmatic applications due to slow processing speeds or poor image quality. This paper proposes an innovative arbitrary image style transfer framework aimed at generating content-preserving images while transferring style features. Our model introduces multi-scale style projection and contrastive learning mechanisms to extract style information at multiple scales, project it into a latent style space, and optimize style representations through contrastive loss. This transforms them into discriminative style codes, which replace the style mean and standard deviation used in the traditional AdaIN method. Then, a style-guided attention mechanism is introduced when fusing style and content features, allowing content features to more precisely adapt to the target style under the direction of style, enhancing the visual quality of the synthesized image. Experimental results show that this model not only outperforms existing methods in visual quality but also improves inference speed, providing strong support for the practical application of AST.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

High-Fidelity Arbitrary Style Transfer Based on Latent Style Space Projection

  • Youzhi Xu,
  • Yuan Peng,
  • Chaochao Sun

摘要

The goal of arbitrary style transfer (AST) is to efficiently apply any artistic style to a content image. However, existing approaches are often hindered in pragmatic applications due to slow processing speeds or poor image quality. This paper proposes an innovative arbitrary image style transfer framework aimed at generating content-preserving images while transferring style features. Our model introduces multi-scale style projection and contrastive learning mechanisms to extract style information at multiple scales, project it into a latent style space, and optimize style representations through contrastive loss. This transforms them into discriminative style codes, which replace the style mean and standard deviation used in the traditional AdaIN method. Then, a style-guided attention mechanism is introduced when fusing style and content features, allowing content features to more precisely adapt to the target style under the direction of style, enhancing the visual quality of the synthesized image. Experimental results show that this model not only outperforms existing methods in visual quality but also improves inference speed, providing strong support for the practical application of AST.