Generation of Image Caption for Visually Challenged People
摘要
In recent years, advances in image interpretation and automatic image captioning have attracted lots of researchers to make use of and employ AI models. It integrates both computer vision and natural language processing (NLP) to generate descriptions in relation to the image observed. In our work, we present an assistive technology based on deep learning to better help visually challenged people to thoroughly understand images on the internet. The newly proposed automated image captioning (AIC) model consists of the following phases: data acquisition, non-captioned image selection, extraction of appearance and texture features, and generation of image captions. The model is trained to maximize the likelihood of the target description sentence to produce. Caption generation (CG) in computer vision is predicted to get a lot of interest owing to its numerous applications such as virtual assistants, image interpretation, image retrieval or indexing, and assisting visually challenged people hence improving their daily lives.