The technique NST is used to convert one image into another image without changing of the content. The only change is the configurations of the image. The content describes the layout and style representing the paint or the colors. NST deals with Content and Style images. This research paper presents a novel integration of neural style transfer and image captioning, combining artistic expression with semantic understanding. Vision Transformers and GPT-2 are two transformer-based language models that have shown extraordinary proficiency in comprehending and producing genuine language. The “Fast Arbitrary Image Styling” model from TensorFlow Hub gives users to combine multiple styles and shows various artistic elements. To obtain the model’s capability to produce captions for stylized images fosters enhanced user interaction, facilitating deeper connections between artists and the model’s outputs. This fusion of artistic flair and semantic understanding offers promising potential for innovative and expressive artistic endeavours.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Image Captioning with Neural Style Transfer Using GPT-2 and Vision Transformer Architectures

  • Mamatha Mandava,
  • Surendra Reddy Vinta

摘要

The technique NST is used to convert one image into another image without changing of the content. The only change is the configurations of the image. The content describes the layout and style representing the paint or the colors. NST deals with Content and Style images. This research paper presents a novel integration of neural style transfer and image captioning, combining artistic expression with semantic understanding. Vision Transformers and GPT-2 are two transformer-based language models that have shown extraordinary proficiency in comprehending and producing genuine language. The “Fast Arbitrary Image Styling” model from TensorFlow Hub gives users to combine multiple styles and shows various artistic elements. To obtain the model’s capability to produce captions for stylized images fosters enhanced user interaction, facilitating deeper connections between artists and the model’s outputs. This fusion of artistic flair and semantic understanding offers promising potential for innovative and expressive artistic endeavours.