Writing expressive captions for visual content by automated procedures is a difficult task, but it has great potential to improve blind people’s perception of their environment. Due to the widespread availability of smartphones with cameras, people who are blind or visually challenged can take pictures of their surroundings. In this research, we use these photos to develop a ResNet-50-LSTM architecture that is specifically meant to produce succinct descriptions for images. In order to generate descriptive English captions, our method uses a ResNet-50 model to extract image characteristics, which are then smoothly fed into an LSTM network. Current methods use recurrent neural networks (RNNs) and convolutional neural networks (CNNs) or their variations to generate captions that are considered appropriate. Notable accomplishments include aligning our model with cutting-edge standards based on performance parameters like accuracy. With their high descriptiveness, the generated captions have the potential to greatly improve the quality of life for visually impaired people by providing them with comprehensive context-specific information. Furthermore, our model exhibits an impressive accuracy rate of 87%, demonstrating its effectiveness in providing visually impaired individuals with precise and informative image labels.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automatic Image Label Generator for Visually Impaired People

  • Tanu Singh,
  • Ashwani,
  • Dheeraj Maurya,
  • Sandhya Avasthi,
  • Kadambri Agarwal

摘要

Writing expressive captions for visual content by automated procedures is a difficult task, but it has great potential to improve blind people’s perception of their environment. Due to the widespread availability of smartphones with cameras, people who are blind or visually challenged can take pictures of their surroundings. In this research, we use these photos to develop a ResNet-50-LSTM architecture that is specifically meant to produce succinct descriptions for images. In order to generate descriptive English captions, our method uses a ResNet-50 model to extract image characteristics, which are then smoothly fed into an LSTM network. Current methods use recurrent neural networks (RNNs) and convolutional neural networks (CNNs) or their variations to generate captions that are considered appropriate. Notable accomplishments include aligning our model with cutting-edge standards based on performance parameters like accuracy. With their high descriptiveness, the generated captions have the potential to greatly improve the quality of life for visually impaired people by providing them with comprehensive context-specific information. Furthermore, our model exhibits an impressive accuracy rate of 87%, demonstrating its effectiveness in providing visually impaired individuals with precise and informative image labels.