Syntactic-guided optimization of image–text matching for intra-modal modeling
摘要
Image–text matching has become a research hotspot for multimodal matching tasks. Existing image–text matching models suffer from feature redundancy in intra-modal modeling of images. This redundant information not only dilutes the important region weights but also is introduced as irrelevant or distracting information. In addition, the syntactic structure of the text is made difficult to analyze by the model due to the lack of syntactic information, which affects the dependencies between words. To address the above problems, the syntactic-guided optimization of image–text matching for intra-modal modeling (SGIM) is proposed. Multi-view filtering of image features is used based on the differences in information richness in each region of the image. The importance weights of each region are adjusted. To obtain an enhanced representation of the image features, multi-view information is fused. A syntactic dependency enhancement method is proposed based on the contextual relationships between text words. To avoid the loss of long-range textual context, the attention distribution of the entire sentence is adjusted. The experimental results show that the SGIM model achieves a minimum improvement of 4.8%, 1.7%, and 2.1% in recall sum (Rsum) compared to MAG, MSR, MMCA, SGRAF, and ReSG on the publicly available datasets Flickr30K, MSCOCO 1K, and MSCOCO 5K.