Digital transformation has witnessed a boom in the preservation and dissemination of digital heritage, especially the classification and interpretation of artifacts. However, existing approaches to artfact classification and interpretation face challenges such as inefficiency, reliance on human expertise and multimodal information heterogeneity. To tackle these challenges, this paper proposes a model for classifying and interpreting cultural relics based on CLIP (Contrastive Language-Image Pre-training), which is fine-tuned to achieve accurate classification and contextualized description of cultural relic images. Firstly, all artifact images are standardised and enhanced, and labels are constructed based on artifact metadata to be associated with the corresponding images. Then, joint learning of images and text is achieved through contrast learning and the classification accuracy is optimised using a specially designed loss function. Next, textual features corresponding to the image are extracted and the image and text are compared, thus mapping the visual information of the image and the descriptive information of the text into a shared representation space. Finally, the generative model is used to generate detailed, artifact-specific interpretation. The method is applied to the Rijksmuseum dataset to improve the accuracy of artifact categorization and to be able to generate interpretations that are highly relevant to the artifacts.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Integrating CLIP for Contextualized Artifact Classification and Interpretation in Digital Heritage Archives

  • Xiangmin Liu,
  • Jian Hu,
  • Linghan Chen,
  • Yuqin Zhu

摘要

Digital transformation has witnessed a boom in the preservation and dissemination of digital heritage, especially the classification and interpretation of artifacts. However, existing approaches to artfact classification and interpretation face challenges such as inefficiency, reliance on human expertise and multimodal information heterogeneity. To tackle these challenges, this paper proposes a model for classifying and interpreting cultural relics based on CLIP (Contrastive Language-Image Pre-training), which is fine-tuned to achieve accurate classification and contextualized description of cultural relic images. Firstly, all artifact images are standardised and enhanced, and labels are constructed based on artifact metadata to be associated with the corresponding images. Then, joint learning of images and text is achieved through contrast learning and the classification accuracy is optimised using a specially designed loss function. Next, textual features corresponding to the image are extracted and the image and text are compared, thus mapping the visual information of the image and the descriptive information of the text into a shared representation space. Finally, the generative model is used to generate detailed, artifact-specific interpretation. The method is applied to the Rijksmuseum dataset to improve the accuracy of artifact categorization and to be able to generate interpretations that are highly relevant to the artifacts.