Enhancing Self-attention Training in Transformers for Image Captioning Using ViT and BERT
摘要
The convergence of Vision Transformer (ViT) and Bidirectional Encoder Representations from Transformers (BERT) has emerged as a valuable fusion mechanism for several applications. This paper introduces an automated system that utilizes ViT's in comprehending complicate visual patterns and BERT's contextual understanding of language to generate captions seamlessly integrated with the visual context. By improving the self-attention mechanisms, the proposed model learns to dynamically weigh the importance of both visual and textual information, resulting in captions that are descriptive. This paper provides a comprehensive exploration of our system's architecture, training methodology, and fusion strategies. The experimentation and evaluation, reveal that the proposed ViT + BERT based model for image captioning is efficient than the existing methods.