Visually impaired people are frequently unaware of outside obstructions and require assistance to avoid dangerous collisions. The sense of sight is the most fundamental aspect of human perception and plays. We propose an object recognition algorithm and assistance system that is very useful for their safety and quality of life. This project aims to use an available mobile device as a talking guide dog to aid those with vision impairments in outdoor navigation and inform them about obstacles. The proposed technology will lower the risks of obstacle contact by allowing users to move outside without stumbling, potentially informing the user what the objects are. The approach proposed uses algorithms of deep learning. CNN recognizes salient functions and captions of photos and converts written text to speech by detecting features through the broadcasted picture alongside its relevant caption, while the Gated Recurrent Unit (GRU) network acts as a tool that captions and describes the text detected from photographs. The predicted caption is transformed into an audio message. The proposed Convolution Neural Networks [CNNs] Gated Recurrent Units [GRUs] model has investigated the usage of numerous network architectures: Inception Net, Mobile Net, Xception Net, ResNet, and VGG16.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Soft Computing for Visual Recognition Through Audio for the Visually Impaired

  • S. R. K. L. Amulya,
  • Vishakha Singh,
  • Tummalapalli Sandeep,
  • Surya Sasidhar,
  • Rita Roy

摘要

Visually impaired people are frequently unaware of outside obstructions and require assistance to avoid dangerous collisions. The sense of sight is the most fundamental aspect of human perception and plays. We propose an object recognition algorithm and assistance system that is very useful for their safety and quality of life. This project aims to use an available mobile device as a talking guide dog to aid those with vision impairments in outdoor navigation and inform them about obstacles. The proposed technology will lower the risks of obstacle contact by allowing users to move outside without stumbling, potentially informing the user what the objects are. The approach proposed uses algorithms of deep learning. CNN recognizes salient functions and captions of photos and converts written text to speech by detecting features through the broadcasted picture alongside its relevant caption, while the Gated Recurrent Unit (GRU) network acts as a tool that captions and describes the text detected from photographs. The predicted caption is transformed into an audio message. The proposed Convolution Neural Networks [CNNs] Gated Recurrent Units [GRUs] model has investigated the usage of numerous network architectures: Inception Net, Mobile Net, Xception Net, ResNet, and VGG16.