<p>This paper proposes a novel adaptive feature fusion strategy that combines a dual-layer attention mechanism and Multi-modal deep reinforcement learning (DRL) to optimize cross-modal information retrieval. The dual-layer attention mechanism enhances the model's ability to capture deep semantic relationships between different modalities, while DRL optimizes the feature extraction and fusion process, improving adaptability in complex environments. Experimental results demonstrate that this strategy outperforms traditional CNN and RNN methods in terms of accuracy, recall, and efficiency across a range of cross-modal retrieval tasks, particularly in multi-modal data environments such as text-image, text-video, and image-video. The proposed approach offers a promising solution for improving the accuracy and efficiency of cross-modal information retrieval.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An adaptive feature fusion strategy using dual-layer attention and multi-modal deep reinforcement learning for all-media similarity search

  • Jin Yue,
  • Jiayun Lang,
  • Rui Feng

摘要

This paper proposes a novel adaptive feature fusion strategy that combines a dual-layer attention mechanism and Multi-modal deep reinforcement learning (DRL) to optimize cross-modal information retrieval. The dual-layer attention mechanism enhances the model's ability to capture deep semantic relationships between different modalities, while DRL optimizes the feature extraction and fusion process, improving adaptability in complex environments. Experimental results demonstrate that this strategy outperforms traditional CNN and RNN methods in terms of accuracy, recall, and efficiency across a range of cross-modal retrieval tasks, particularly in multi-modal data environments such as text-image, text-video, and image-video. The proposed approach offers a promising solution for improving the accuracy and efficiency of cross-modal information retrieval.