错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Recurrent fusion transformer for image captioning

  • Zhenping Mou,
  • Qiao Yuan,
  • Tianqi Song

摘要

Image captioning describes the visual content of a given image by using natural language sentences. However, in the existing image captioning model, the encoder only describes the image content from a specific pattern, and cannot fully understand the semantic sequence information of the input image. In this paper, we propose a multimodal recurrent fusion block (RF-Block), which uses a new recurrent attention and combines gated recurrent networks to capture feature correlation information. Also, a stack multiple feature fusion blocks is used can better enhance the relationship between higher-level features. Finally, the feature fusion block is inserted into the Transformer to form a recurrent fusion transformer (RCT), which can improve the performance of the image captioning model. Experimental results show that the proposed model is better than the traditional encoder-decoder image captioning model and provides comparable performance to the most advanced models.