<p>This research focuses on developing a Presentation Voice Descriptor (PVD) system using AI to transform visual presentation content into comprehensive audio descriptions for visually impaired students. The PVD system employs image recognition technologies to analyze and interpret Microsoft Power Point Presentation slide contents, generating accurate and contextually relevant voice narratives. The system converts power point presentations into images and reads text from these images using Optical Character Recognition. The extracted text is then converted to speech through text to speech conversion. This tool aims to provide real-time access to presentation content, enhancing educational accessibility and engagement for visually impaired learners. The generated voice descriptions are sent to the ear pods of the blind student as a live audio during the session and the same can be attached to the presentation for future reference. The development process involves creating a robust AI model trained on diverse data sets to ensure adaptability across various presentation styles and subjects. Initial testing in academic settings shows significant improvements in comprehension and student engagement. The evaluation of the results highlights how Generative AI-powered assistive technologies help in promoting educational equity and inclusion.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Presentation voice descriptor using generative AI as a learning aid for visually impaired learners

  • P. R. Renjith,
  • T. M. Prajesha

摘要

This research focuses on developing a Presentation Voice Descriptor (PVD) system using AI to transform visual presentation content into comprehensive audio descriptions for visually impaired students. The PVD system employs image recognition technologies to analyze and interpret Microsoft Power Point Presentation slide contents, generating accurate and contextually relevant voice narratives. The system converts power point presentations into images and reads text from these images using Optical Character Recognition. The extracted text is then converted to speech through text to speech conversion. This tool aims to provide real-time access to presentation content, enhancing educational accessibility and engagement for visually impaired learners. The generated voice descriptions are sent to the ear pods of the blind student as a live audio during the session and the same can be attached to the presentation for future reference. The development process involves creating a robust AI model trained on diverse data sets to ensure adaptability across various presentation styles and subjects. Initial testing in academic settings shows significant improvements in comprehension and student engagement. The evaluation of the results highlights how Generative AI-powered assistive technologies help in promoting educational equity and inclusion.