<p>The Transformer based model has achieved significant performance improvement in the image captioning. However, at present, this field still faces two main problems: one is the lack of contextual visual information, and the other is how to effectively integrate two different types of visual features. Therefore, we propose a Dual-Visual Collaborative Enhanced Transformer (DVCET) model that makes full use of grid and region visual features to improve the performance of image captions. Specifically, we design a Grid Aggregation Encoding Layer (GAEL) to integrate adjacent context information into each feature and help the model capture local and global context information in the image. Then, we design the Region Graph Memory Encoding Layer (RGMEL) to read and write the visual region representation using the visual graph memory to achieve object level relational reasoning. Finally, we introduce Dual Gated Collaboration (DGC) to make better use of the multi-level characteristics of the two through the dual gate control operation and reduce the information redundancy of the grid structure. The experimental results on MSCOCO and Flickr30K verify that the proposed model can improve the description performance, achieve good results on multiple evaluation scores, and achieve the performance of competing with the most advanced methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dual-visual collaborative enhanced transformer for image captioning

  • Zhenping Mou,
  • Tianqi Song,
  • Hong Luo

摘要

The Transformer based model has achieved significant performance improvement in the image captioning. However, at present, this field still faces two main problems: one is the lack of contextual visual information, and the other is how to effectively integrate two different types of visual features. Therefore, we propose a Dual-Visual Collaborative Enhanced Transformer (DVCET) model that makes full use of grid and region visual features to improve the performance of image captions. Specifically, we design a Grid Aggregation Encoding Layer (GAEL) to integrate adjacent context information into each feature and help the model capture local and global context information in the image. Then, we design the Region Graph Memory Encoding Layer (RGMEL) to read and write the visual region representation using the visual graph memory to achieve object level relational reasoning. Finally, we introduce Dual Gated Collaboration (DGC) to make better use of the multi-level characteristics of the two through the dual gate control operation and reduce the information redundancy of the grid structure. The experimental results on MSCOCO and Flickr30K verify that the proposed model can improve the description performance, achieve good results on multiple evaluation scores, and achieve the performance of competing with the most advanced methods.