Advanced Image Captioning Using Deep Learning Techniques: A CNN-LSTM Approach
摘要
The rapid progress in AI technology has led to significant advancements in the field of image captioning, particularly through the application of deep learning techniques. Our research discusses the advancements in AI technology focusing on generating image captions through deep learning techniques. The process involves tasks such as object identification, establishing semantic connections, and translating scene information into relevant phrases. Our model, which incorporates computer vision and natural language processing, automatically generates information from images. The evaluation uses the Flickr8k dataset with 8000 photographs to assess the model’s fluency and accuracy, demonstrating its capability to generate appropriate captions.