Enhanced image captioning with advanced context-aware object relational model
摘要
Image captioning generates text in natural language to describe a given image. Recent advances in various object detection with attention mechanisms pushed for exploring multiple image captioning methodologies to create more meaningful and accurate captioning models. Although existing pipelines do well in describing the image, there has not been enough emphasis on relationship modeling between image features. Relationship modeling is crucial for establishing context between various objects in an image. We have proposed a novel Advanced Context-Aware Object Relational Model (ACAORM) that not only improves relationship modeling between image features but also better sentence generation due to Transformer architecture. ACAORM builds relation-aware visual representations for image description and builds captions with prior attention to relevant Regions of Interest (RoI). We tested the proposed methodology with three widely used datasets, Flickr8K, Flickr30K, and MS-COCO. The results show that it beats numerous cutting-edge techniques. ACAORM scores 0.3526, 0.4439, and 0.8813 in