Integrated Global Semantics and Local Details for Image-Text Retrieval
摘要
Image-text retrieval is a critical challenge for understanding the semantic relationship between vision and language domains. Previous studies have focused on analyzing either global or local features, neglecting the intrinsic connections between these two levels of granularity. In addition, current mainstream methods attempt to construct a unified semantic space by aggregating the weighted features of different segments, aiming to enhance the interaction with different granularities of information. However, these methods may be disrupted by irrelevant segments, leading to semantic misalignment. To address these challenges, we propose a Bidirectional Focused Attention and Global-Local Matching Fusion Network (BFAGL). The proposed method well integrates the similarity matching of global semantic contexts with the precision of local detail analysis, ensuring that global semantic alignment is achieved without compromising critical local information. Furthermore, it incorporates bi-directional focal attention into the local matching process to promote a nuanced understanding of the contextual semantic relationships of images and text. Experiments on the Flickr30K and MS-COCO datasets demonstrate the state-of-the-art performance of BFAGL in image-text retrieval tasks.