This research proposes an efficient and fast visual description generation system that makes use of an encoder-decoder framework to maintain the trade-off between computational efficiency and captioning performance. In order to perform faster training and better parameter efficiency, the proposed model uses EfficientNetV2 in the encoder and a transformer model in the decoder to generate coherent, human-like textual descriptions. The inclusion of progressive learning and dynamic adjustment of regularization in the encoder from weak to strong for small to large-size images can speed up the extraction of rich visual features for the transformer decoder. Since the transformer model contextualizes information through multi-head self-attention mechanisms, it speeds up the learning process. The proposed visual description generation model undergoes evaluation on the extensively utilized MS COCO benchmark dataset, yielding compelling results in assessments. This underscores its capability to generate descriptive and contextually relevant descriptions across a diverse array of images.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Multi-modal Approach for Efficient and Contextually Rich Visual Description Generation

  • Biswajit Patra,
  • Dakshina Ranjan Kisku

摘要

This research proposes an efficient and fast visual description generation system that makes use of an encoder-decoder framework to maintain the trade-off between computational efficiency and captioning performance. In order to perform faster training and better parameter efficiency, the proposed model uses EfficientNetV2 in the encoder and a transformer model in the decoder to generate coherent, human-like textual descriptions. The inclusion of progressive learning and dynamic adjustment of regularization in the encoder from weak to strong for small to large-size images can speed up the extraction of rich visual features for the transformer decoder. Since the transformer model contextualizes information through multi-head self-attention mechanisms, it speeds up the learning process. The proposed visual description generation model undergoes evaluation on the extensively utilized MS COCO benchmark dataset, yielding compelling results in assessments. This underscores its capability to generate descriptive and contextually relevant descriptions across a diverse array of images.