A Diffusion Scale-Enhanced CLIP Model for Cross-Lingual Cross-Modal Building Information Retrieval
摘要
Building information retrieval plays a crucial role in enhancing architectural design and management efficiency. Current building information retrieval faces challenges in dataset collection, language diversity, and feature extraction from complex architectural images. These multifaceted impediments significantly compromise the retrieval process’s precision and computational efficacy. This study proposes a cross-lingual cross-modal building information retrieval method. First, a Cross-lingual Building Scenes (CLBS) dataset is constructed, comprising rich building images and corresponding textual descriptions. Then, this paper devises a novel cross-lingual cross-modal building information retrieval model. The proposed model utilizes contrastive learning to capture the subtle relationship between textual and visual content in different languages. Finally, Diffusion Scale Networks (DSN) is introduced in retrieval models to enhance the capture of complex architectural features. DSN employs an entropy-driven adaptive temporal selection strategy to extract and fuse representative multi-scale building image features. The experimental results indicate that the average precision of Cross-lingual CLIP text-to-image retrieval is 73.6% (Chinese) and 71.6% (English), while the average accuracy of image-to-text retrieval is 83.3% (Chinese) and 82.4% (English).