Image Captioning with Global Information Enhanced Image Representation
摘要
In recent years, significant progress has been made in image captioning using the encoder-decoder framework. However, most of these methods fail to exploit global information effectively. Global information provides a coarse understanding of the visual scene, which may contribute to the task of image captioning. In this paper, we propose a global information enhancement module which is adopted to enhance the extracted grid features of an image with global information. Besides, we introduce positional encoding of the grids in encoder to better model relationship among grid features. To validate our model, we conduct extensive experiments on the COCO image captioning dataset and achieve superior performance over many state-of-the-art methods.