Only Alignment and Reconstruction: Masked Multimodal Autoencoders for Vison-Language Pre-training
摘要
Vision and language pre-training (VLP) has demonstrated significant effectiveness across various vision-language (V+L) tasks. Traditional unsupervised VLP methods predominantly utilize a unimodal masking model, typically comprising a visual encoder and a textual en- coder. These models often restrict masking to textual data only. Consequently, such a unimodal approach limits the effective integration, or ’fusibility,’ of features extracted from each modality. Addressing this limitation, this paper introduces the Masked Multimodal Autoencoder (MME). MME innovatively enhances feature fusibility by optimizing mutual information across masked, cross-modal data. This is achieved through the integration of a dynamic masking strategy, applicable to both visual and textual encoders, coupled with a novel pre-training task.