Learning to Generate: Text Guided Visual Feature Extraction for Image Captioning Using Joint Two-Phase Learning Model
摘要
Image Captioning can be defined as the ability of the machine to identify the features and environment to provide a caption that describes the situation in detail. Over the years, various authors have proposed unique solutions to either the problem of caption generation or an alteration in the caption generation mechanisms. The problem associated with these solutions is that, while they have provided a higher accuracy, they fail to properly correlate the features extracted to the caption, which could potentially decrease the descriptiveness of the caption. This may result in ambiguous captions that could be difficult to comprehend. In our project, we aim to address these issues by integrating a joint two- phase learning model, which will help in identification and classification of features present in the image, to generate a caption. From our results, we have seen an 89% accuracy in the generated captions. As a result, we will potentially be able to increase the implementation of this application in a wide range of domains such as security and education, without having to concern ourselves with limited accuracy.