A multitude of researchers have commenced employing AI models in Automatic Image Captioning (AIC), particularly owing to progress in image analysis. These models integrate computer vision with Natural Language Processing (NLP) to connect visual data and human language by producing accurate descriptions of visual content. AIC surpasses mere labelling, facilitating a profound comprehension of images, essential for individuals with visual impairments. This technology improves the engagement and experience of visually impaired users by automatically delivering text-based descriptions, aiding their understanding of images encountered online and in their surroundings. The proposed AIC model is organised into a sequence of meticulous steps to guarantee robustness. It primarily emphasises data collection, integrating a varied assortment of images to enhance the model’s efficacy. Subsequently, uncaptioned images are utilised to gather raw caption data. The model subsequently underscores the extraction of texture and appearance data from these images, which is essential for object identification and contextual understanding. This information is then employed for caption synthesis, wherein NLP techniques produce precise, contextually appropriate descriptions. The model is engineered to accommodate diverse scenarios and intricate visual contexts by improving its ability to generate pertinent target descriptions for each image, attaining an accuracy rate of 85%. This technology has potential applications in virtual assistants, image analysis, indexing and notably in assisting the visually impaired. AIC can greatly improve the daily experiences of individuals dependent on auditory descriptions by rendering visual content accessible. With the increasing interest in deep learning and natural language processing, automatic caption generation is anticipated to become more accurate and customised. The main aim of this project is to create a strong AIC model that successfully connects visual content with natural language, enhancing accessibility for individuals with visual impairments.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LSTM-Based Image Enrichment and Description Generation to Enhance Visual Perception for the Visually Challenged People

  • Riya Agarwal,
  • Yalavarthi Jaswanth,
  • P. Saranya

摘要

A multitude of researchers have commenced employing AI models in Automatic Image Captioning (AIC), particularly owing to progress in image analysis. These models integrate computer vision with Natural Language Processing (NLP) to connect visual data and human language by producing accurate descriptions of visual content. AIC surpasses mere labelling, facilitating a profound comprehension of images, essential for individuals with visual impairments. This technology improves the engagement and experience of visually impaired users by automatically delivering text-based descriptions, aiding their understanding of images encountered online and in their surroundings. The proposed AIC model is organised into a sequence of meticulous steps to guarantee robustness. It primarily emphasises data collection, integrating a varied assortment of images to enhance the model’s efficacy. Subsequently, uncaptioned images are utilised to gather raw caption data. The model subsequently underscores the extraction of texture and appearance data from these images, which is essential for object identification and contextual understanding. This information is then employed for caption synthesis, wherein NLP techniques produce precise, contextually appropriate descriptions. The model is engineered to accommodate diverse scenarios and intricate visual contexts by improving its ability to generate pertinent target descriptions for each image, attaining an accuracy rate of 85%. This technology has potential applications in virtual assistants, image analysis, indexing and notably in assisting the visually impaired. AIC can greatly improve the daily experiences of individuals dependent on auditory descriptions by rendering visual content accessible. With the increasing interest in deep learning and natural language processing, automatic caption generation is anticipated to become more accurate and customised. The main aim of this project is to create a strong AIC model that successfully connects visual content with natural language, enhancing accessibility for individuals with visual impairments.