CaLB: Caption and Language-Based with Simple Fusion Network
摘要
Multimodal Vision-Language models successfully integrate the advantages of the visual and linguistic domains. However, there exists an inherent semantic gap between images and language, which to some extent restricts the development of multimodal Vision-Language models. In current research, achieving semantic alignment between these two modalities has become a quite challenging task. The process of generating image captions inherently contains a complex pre-alignment mechanism, and the approach in this study leverages this characteristic, thereby enabling the comprehensive capture of detailed information related to multimodal alignment while effectively filtering out irrelevant noise. In summary, this paper validates the effectiveness of using image captions generated by a multimodal pre-training model as input to the visual part of the multimodal model, and through experiments on two datasets, clearly demonstrates the significant advantages of this feature.