Deep Learning–Based Surrounding Descriptor for the Visually Challenged
摘要
The proposed work presented in this chapter, “Deep learning–based surroundings descriptor for the visually challenged,” aims to make a system software that helps visually impaired or blind individuals in perceiving and understanding their surroundings. Visually challenged persons are not able to enjoy the surroundings and nature as we normally do, and this project aims at improving their quality of life. The system will capture images to detect environmental features and then provide real-world auditory feedback to the user. The goal of this project is to create a reliable, user-friendly, and assistive technology and an affordable device software that improves the visually challenged person’s life and makes them independent, allowing them to move around more safely and confidently. This project uses several deep learning and NLP technologies such as CNN and LSTM. The image inputs are taken and the features are extracted using the VGG16 model, which is an advanced version of CNN that is great at object identification and localization. In addition, we used LSTM for training the model with the extracted features and the corresponding captions. Finally, when a user gives an image input to the trained model, it predicts the caption, and then it converts the text-to-speech for the user. The Surroundings Descriptor technology aims to solve the difficulties that visually impaired people face when navigating their surroundings and to improve their experience of life. The project includes technology design and development, system testing and evaluation, and improving the model based on user feedback. Finally, the Surroundings Descriptor has the potential to significantly improve visually impaired individuals’ mobility and independence, allowing them to participate better in public life and live more fulfilling lives.