Foundation models in computer vision are trained in a special way on a very large datasets, allowing them to learn diverse and rich knowledge about the visual domain of our world. Therefore, they enable solving various complex tasks, including zero-shot learning. To build predictable AI-based solutions, high interpretability and explainability of the internal structure of knowledge representations in deep neural networks are necessary. How do Foundation models “see” knowledge in internal representations? Do features of the training dataset influence the properties of the learned knowledge representations? We systematically study these issues through the lens of geometry and consider the properties of representations in the embedding space inside Transformers from the perspective of manifold learning. We explore the deep neural networks embeddings dynamics on different layers through point cloud similarity analysis based on the Log-Euclidean Signature (LES) Distance method. Additionally, we introduce the Domain Geometry Divergence approach for better visual image domains space understanding. We are conducting extensive experiments with Foundation models architecture such as CLIP, DINOv2, BeiT, MAE and test our approach on various datasets including medical images.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Study of Foundation Models Knowledge Representations: Geometry Perspective

  • G. Magai,
  • A. Soroka

摘要

Foundation models in computer vision are trained in a special way on a very large datasets, allowing them to learn diverse and rich knowledge about the visual domain of our world. Therefore, they enable solving various complex tasks, including zero-shot learning. To build predictable AI-based solutions, high interpretability and explainability of the internal structure of knowledge representations in deep neural networks are necessary. How do Foundation models “see” knowledge in internal representations? Do features of the training dataset influence the properties of the learned knowledge representations? We systematically study these issues through the lens of geometry and consider the properties of representations in the embedding space inside Transformers from the perspective of manifold learning. We explore the deep neural networks embeddings dynamics on different layers through point cloud similarity analysis based on the Log-Euclidean Signature (LES) Distance method. Additionally, we introduce the Domain Geometry Divergence approach for better visual image domains space understanding. We are conducting extensive experiments with Foundation models architecture such as CLIP, DINOv2, BeiT, MAE and test our approach on various datasets including medical images.