A Survey on Automatic Image Captioning Approaches: Contemporary Trends and Future Perspectives
摘要
The automatic generation of image captions is one of the complex computer vision tasks that involve integration of object detection and natural language processing (NLP). In recent times, one of the significant aspects is to design image captioning approaches that accurately and efficiently generate appropriate image captions in a particular domain. With the emergence of deep learning paradigms, the task of image captioning becomes comparatively easier than traditional template-based approaches. In this article, we expound an in-depth examination of state of the art (SOTA) image captioning methods, along with the key conceptions. Besides, a comparative analysis of evaluation protocols is presented that are presently used to access the efficacy of the algorithms. Moreover, the study reveals open research issues in the existing methods that can be further investigated by the research community. One of the key challenges is to develop larger corpora of language specific dataset to design image captioning approaches in other regional languages such as Hindi, Marathi, Sanskrit, Telugu, and Gujarati etc. Furthermore, designing accurate and efficient image captioning approaches requisite the notion of attention mechanism in the images.