<p>In layout image-text multimodal retrieval, fine-grained semantic alignment deviation and insufficient representation of geometric relationships seriously restrict the retrieval performance. This paper constructs a contrastive language-image pre-training-vision transformer (CLIP-ViT) framework that integrates spatial attention and position encoding enhancement. High-precision layout-text retrieval is achieved through layout-aware contrastive learning and structure-adaptive feature extraction. A dynamic spatial attention module is designed to establish a local regional association matrix in the CLIP image encoding stage, filter key visual elements through a learnable gating mechanism, insert a cross-head spatial attention sublayer in each layer of ViT, and compute the relative position sensitivity between feature patches. A hybrid position representation scheme is adopted to fuse the absolute coordinate encoding with the relative position bias vector by linear interpolation. According to the layout image characteristics, the normalized coordinate features of the element bounding box are additionally injected into the position embedding layer of ViT. A two-way contrastive learning objective is constructed. A spatial structural similarity constraint is added based on the traditional image-text matching loss. The Wasserstein distance between the predicted layout heat map and the actual element distribution is computed to enhance the model’s ability to model geometric arrangement rules. A multi-scale feature pyramid network is designed, and a deformable convolutional layer group is connected to the back end of the CLIP visual encoder. The feature fusion path is automatically selected according to the layout complexity. High-level semantic features are enabled for densely arranged areas, and the underlying detail features are retained in sparse areas. Experimental results show that the average mAP (mean average precision) of CLIP-ViT in this paper is 0.89 on five typical layout types, and the Recall@K (Recall at Top-K) is 0.76 when the K value is 20, which has a high cross-modal retrieval accuracy. In the fine-grained alignment quantification results, CLIP-ViT has significant advantages in the two core indicators of attention entropy (1.82) and KL (Kullback-Leibler) divergence (0.47). In the basic, inclusion, and overlapping relationships, CLIP-ViT maintains a stable average F1-score (88.8%), and its geometric relationship modeling ability has better generalization. Experimental data proves the effectiveness of this paper’s research on multimodal knowledge retrieval of layout image text.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal knowledge retrieval of layout image text based on CLIP and ViT

  • Bowen Zeng,
  • Rong Lu,
  • Guanghu Mao

摘要

In layout image-text multimodal retrieval, fine-grained semantic alignment deviation and insufficient representation of geometric relationships seriously restrict the retrieval performance. This paper constructs a contrastive language-image pre-training-vision transformer (CLIP-ViT) framework that integrates spatial attention and position encoding enhancement. High-precision layout-text retrieval is achieved through layout-aware contrastive learning and structure-adaptive feature extraction. A dynamic spatial attention module is designed to establish a local regional association matrix in the CLIP image encoding stage, filter key visual elements through a learnable gating mechanism, insert a cross-head spatial attention sublayer in each layer of ViT, and compute the relative position sensitivity between feature patches. A hybrid position representation scheme is adopted to fuse the absolute coordinate encoding with the relative position bias vector by linear interpolation. According to the layout image characteristics, the normalized coordinate features of the element bounding box are additionally injected into the position embedding layer of ViT. A two-way contrastive learning objective is constructed. A spatial structural similarity constraint is added based on the traditional image-text matching loss. The Wasserstein distance between the predicted layout heat map and the actual element distribution is computed to enhance the model’s ability to model geometric arrangement rules. A multi-scale feature pyramid network is designed, and a deformable convolutional layer group is connected to the back end of the CLIP visual encoder. The feature fusion path is automatically selected according to the layout complexity. High-level semantic features are enabled for densely arranged areas, and the underlying detail features are retained in sparse areas. Experimental results show that the average mAP (mean average precision) of CLIP-ViT in this paper is 0.89 on five typical layout types, and the Recall@K (Recall at Top-K) is 0.76 when the K value is 20, which has a high cross-modal retrieval accuracy. In the fine-grained alignment quantification results, CLIP-ViT has significant advantages in the two core indicators of attention entropy (1.82) and KL (Kullback-Leibler) divergence (0.47). In the basic, inclusion, and overlapping relationships, CLIP-ViT maintains a stable average F1-score (88.8%), and its geometric relationship modeling ability has better generalization. Experimental data proves the effectiveness of this paper’s research on multimodal knowledge retrieval of layout image text.