Drishti: Empowering the Visually Impaired Using Sustainable AI
摘要
Drishti, a gadget that uses deep learning models, visually impaired persons may be able to navigate their surroundings more accurately and conveniently. We suggest a new approach that combines the benefits of the Bootstrapping Language-Image Pre-Training (BLIP) model for image-to-text generation with the Bidirectional and Auto-Regressive Transformers for Acoustic Representations from Kinetics (BARK) model for text-to-speech synthesis. The vision language pre-trained BLIP model is used to reliably detect and classify things in real-world scenarios, going beyond the limitations of pre-defined libraries. After that, the BARK model turns these descriptions into speech that sounds natural, ensuring that the message is understood. This model uses transferred learning approach which lines with the sustainability goals as it reduces the training time and hence carbon footprints produced. This work opens the door for further developments in assistive devices driven by deep learning, which will help people with visual loss navigate the environment with more confidence and autonomy in the future.