错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AFSDCGN: Adaptive Feature Scaling and Dynamic Contextual Graph Networks for image captioning with unseen relationship detection

  • Yugandhara A. Thakare,
  • K. H. Walse,
  • Mohammad Atique

摘要

Automated image captioning systems play a crucial role in various applications such as assistive technologies, content indexing, and robotics. However, current frameworks face limitations in accurately identifying unseen relationships and contextual intricacies within images, hampering their interpretability. Most rely on standard Convolutional Neural Networks (CNNs) for feature extraction and traditional attention mechanisms for caption generation, resulting in sub-optimal precision, accuracy, and recall rates. This paper introduces an innovative image captioning framework addressing these challenges. It incorporates Adaptive Feature Scaling for nuanced object recognition, Dynamic Contextual Graph Networks for improved context understanding, Meta learning for Zero-Shot Detection of unseen relationships, and a Multi-modal Attention Mechanism for cohesive caption generation. Implementation of this model surpasses existing ones, exhibiting improvements in object classification precision (4.9%), accuracy (4.5%), recall (3.5%), and AUC (5.5%), while reducing classification delay by 2.9%. Notably, it enhances relationship representation metrics: Bleu Score (4.9%), METEOR Score (3.5%), CIDEr Score (8.5%), and ROUGE-L Score (3.9%). The proposed framework offers substantial advantages across diverse applications, including assistive technologies for the visually impaired, real-time robotics navigation, and automated content description on social media platforms. These advancements are pivotal for tasks demanding precise image understanding and natural language description.