错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CaLB: Caption and Language-Based with Simple Fusion Network

  • Yicong Shi,
  • Songyang Wu,
  • Zhiguo Ding,
  • Hang Shao,
  • Yingna Li

摘要

Multimodal Vision-Language models successfully integrate the advantages of the visual and linguistic domains. However, there exists an inherent semantic gap between images and language, which to some extent restricts the development of multimodal Vision-Language models. In current research, achieving semantic alignment between these two modalities has become a quite challenging task. The process of generating image captions inherently contains a complex pre-alignment mechanism, and the approach in this study leverages this characteristic, thereby enabling the comprehensive capture of detailed information related to multimodal alignment while effectively filtering out irrelevant noise. In summary, this paper validates the effectiveness of using image captions generated by a multimodal pre-training model as input to the visual part of the multimodal model, and through experiments on two datasets, clearly demonstrates the significant advantages of this feature.