<p>Image captioning aims to generate descriptive captions for visual content, thereby strengthening the connection between images and their semantic meanings. In this paper, we propose SCAP, a novel lightweight model that enhances image captioning through an innovative sifting attention mechanism. SCAP incorporates a summary module and a forget module within its encoder to refine visual information, selectively filtering visual information to retain relevant features and reduce redundancy. The hierarchical decoder then leverages sifting attention to align image features with text captions, generating accurate and contextually relevant descriptions. Extensive experiments conducted on multiple benchmark datasets, including COCO and Flickr30k, demonstrate SCAP’s effectiveness as a highly competitive model in the field. It achieves competitive performance while maintaining computational efficiency, making it particularly suitable for resource-constrained scenarios. This lightweight model represents a notable advancement in advancing image captioning.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SCAP: enhancing image captioning through lightweight feature sifting and hierarchical decoding

  • Yuhao Zhang,
  • Jiaqi Tong,
  • Honglin Liu

摘要

Image captioning aims to generate descriptive captions for visual content, thereby strengthening the connection between images and their semantic meanings. In this paper, we propose SCAP, a novel lightweight model that enhances image captioning through an innovative sifting attention mechanism. SCAP incorporates a summary module and a forget module within its encoder to refine visual information, selectively filtering visual information to retain relevant features and reduce redundancy. The hierarchical decoder then leverages sifting attention to align image features with text captions, generating accurate and contextually relevant descriptions. Extensive experiments conducted on multiple benchmark datasets, including COCO and Flickr30k, demonstrate SCAP’s effectiveness as a highly competitive model in the field. It achieves competitive performance while maintaining computational efficiency, making it particularly suitable for resource-constrained scenarios. This lightweight model represents a notable advancement in advancing image captioning.