<p>Dunhuang murals, with their rich artistic, cultural, and scholarly significance, increasingly benefit from the integration of multimodal resources. However, current mobile visual search methods face two key challenges: (1) a reliance on single-modal retrieval that overlooks the complementary relationship between image and text, and (2) inconsistencies in the matching results across different modalities. To this end, we propose a novel dual-fusion strategy-driven multimodal mobile visual search (MVS) model tailored for Dunhuang murals. This model utilizes LERT and ViT to capture image and text features and introduces a horizontal weighted concatenation fusion (HWCF) strategy to integrate them. For effective result matching, we compute the similarity of L2 normalized features using the dot product. Furthermore, we devise a hybrid-modal cascade decision fusion (HCDF) strategy to re-rank results across different modalities, ensuring they are both visually similar and semantically relevant to the query. Extensive experiments indicate the proposed model is outperforming.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A dual-fusion strategy driven multimodal mobile visual search model for Dunhuang murals

  • Shouqiang Sun,
  • Ziming Zeng,
  • Jingjing Sun,
  • Tingting Li,
  • Qingqing Li,
  • Shuyue Xiao

摘要

Dunhuang murals, with their rich artistic, cultural, and scholarly significance, increasingly benefit from the integration of multimodal resources. However, current mobile visual search methods face two key challenges: (1) a reliance on single-modal retrieval that overlooks the complementary relationship between image and text, and (2) inconsistencies in the matching results across different modalities. To this end, we propose a novel dual-fusion strategy-driven multimodal mobile visual search (MVS) model tailored for Dunhuang murals. This model utilizes LERT and ViT to capture image and text features and introduces a horizontal weighted concatenation fusion (HWCF) strategy to integrate them. For effective result matching, we compute the similarity of L2 normalized features using the dot product. Furthermore, we devise a hybrid-modal cascade decision fusion (HCDF) strategy to re-rank results across different modalities, ensuring they are both visually similar and semantically relevant to the query. Extensive experiments indicate the proposed model is outperforming.