A dual-fusion strategy driven multimodal mobile visual search model for Dunhuang murals
摘要
Dunhuang murals, with their rich artistic, cultural, and scholarly significance, increasingly benefit from the integration of multimodal resources. However, current mobile visual search methods face two key challenges: (1) a reliance on single-modal retrieval that overlooks the complementary relationship between image and text, and (2) inconsistencies in the matching results across different modalities. To this end, we propose a novel dual-fusion strategy-driven multimodal mobile visual search (MVS) model tailored for Dunhuang murals. This model utilizes LERT and ViT to capture image and text features and introduces a horizontal weighted concatenation fusion (HWCF) strategy to integrate them. For effective result matching, we compute the similarity of L2 normalized features using the dot product. Furthermore, we devise a hybrid-modal cascade decision fusion (HCDF) strategy to re-rank results across different modalities, ensuring they are both visually similar and semantically relevant to the query. Extensive experiments indicate the proposed model is outperforming.