Enhancing image captioning with spatial relational attention and grid decoder
摘要
This paper presents a novel image description model that integrates a spatial relational attention mechanism to enhance the model's ability to understand and generate accurate descriptions of images. By combining absolute position encoding (APE) and relative position encoding (RPE), the model improves its spatial awareness, enabling it to capture intricate spatial relationships within the image. The model also introduces the grid decoder, which utilizes a multi-path structure to effectively fuse information from multiple encoder layers, overcoming the limitations of traditional single-path decoders. This design promotes the interaction and integration of multi-level features, improving the model’s performance in generating more accurate and natural image de-scriptions. Experimental results demonstrated that the proposed model outperforms existing image captioning models, particularly in processing complex images with intricate spatial layouts. The model generated more contextually relevant and coherent de-scriptions, showing higher accuracy and naturalness compared to traditional methods.