Multi-view Correlation Learning Cross-Modal Retrieval Based on Multi-layer Attention
摘要
This paper proposes to combine multi-view relevance-based learning with deep adversarial learning model, aiming to resolve the cross-modal matching matter between text and images. Traditional cross-modal retrieval methods often overlook the rich correlation information between multimedia data, resulting in poor retrieval results. To handle this matter, we introduce a multi-layer attention mechanism to simultaneously consider the correlation between different views, effectively combining the semantic information of text and images, and accurately capturing their correlations. Firstly, we use deep neural networks to learn feature representations for text and images, transforming them into high-dimensional feature vectors. Then, we constructed a multi-layer attention model that captures the complex correlations between multi-view data by learning different levels of attention weights; Subsequently, these features were transformed into compact binary codes through hash methods, achieving efficient cross-modal retrieval. The experiments demonstrate that the proposed method has achieved certain performance improvements in cross-modal retrieval tasks, verifying its effectiveness and feasibility in cross-modal retrieval tasks.