Transforming Images into Multilingual Captions for Accessibility and Interaction
摘要
In the digital era, the exponential growth of visual content necessitates innovative approaches for generating descriptive and meaningful captions that cater to a global and linguistically diverse audience. This research introduces a novel vision-language model designed to produce highly accurate and contextually relevant captions in English, which are subsequently translated into multiple languages. The proposed model combines advanced deep learning techniques, including an InceptionResNetV2 encoder and a transformer-based BLIP framework, to ensure precise feature extraction and seamless integration of visual and linguistic data. By leveraging probabilistic decoding strategies, the system also generates interactive and personalized stories from images, enhancing accessibility for visually impaired individuals and enabling diverse applications in education, content creation, and business.Evaluation on the COCO Captions dataset demonstrates the model’s effectiveness, achieving a BLEU-4 score of 34.38 and a METEOR score of 30.12, surpassing baseline methods. These results highlight the model’s superior ability to balance linguistic fluency with semantic accuracy, even in multilingual scenarios. This research contributes to the growing field of multimodal learning by advancing tools that foster inclusivity and enhance communication across linguistic and cultural barriers.