A Comprehensive Review of Multimodal Visual Representation Learning: Tracing the Evolution from CNNs to Transformers and Beyond
摘要
The primary goal of multimodal visual representation learning is to generate implicit information that effectively represents multimodal information by exploring the commonalities and characteristics between different modalities. This research report will discuss currently widely used advanced methods in the field of multimodal visual representation learning. This article will discuss these methods in the following order, culminating in multimodal visual learning: (1) pre-trained visual representation learning, (2) generative visual representation learning, (3) contrastive multimodal visual representation learning, and (4) image-text multimodal visual representation learning methods. Each element provides useful clues that ultimately lead to multimodal visual learning. Pre-trained visual representation learning refers to the application of supervised pre-training models in visual representation learning, while generative visual representation learning uses generative models to learn feature representations that can integrate multimodal information. Contrastive multimodal visual representation learning uses contrastive learning methods to compare similar and dissimilar sample pairs, learning feature representations in a self-supervised manner. Image-text multimodal visual representation learning methods, on the other hand, attempt to enhance the capabilities of visual representation learning by fusing visual information (such as images) with textual information. This review report will explain the above research background, the classification of different research methods, commonly used evaluation methods , and future development trends.