Interactive Attention AI to Translate Low-Light Photos to Captions for Night Scene Understanding in Women Safety
摘要
There is amazing progress in deep learning-based models for image captioning and low-light image enhancement. For the first time in literature, this paper develops a deep learning model that translates night scenes to sentences, opening new possibilities for AI applications in the safety of visually impaired women. Inspired by image captioning and visual question answering, a novel ‘Interactive Image Captioning’ is developed. A user can make the AI focus on any chosen person of interest by influencing the attention scoring. Attention context vectors are computed from CNN feature vectors and user-provided start words. The encoder–attention–decoder neural network learns to produce captions from low-brightness images. This paper demonstrates how women safety can be enabled by researching a novel AI capability in the interactive vision–language model for perception of the environment in the night.