Contrastive learning is a potent technique for self-supervised learning that maintains invariance between two positive views. Advancements such as the “core view” or multi-cropping have harnessed insights from multiple views, culminating in the latest state-of-the-art performance. However, the complexities of learning from multiple positive views remain partially unexplored. In this work, we provide a theoretical formulation on the use of multiple (two or more) positive views in contrastive learning, as well as a “plug-and-play” approach for learning from multiple positive views that seamlessly integrates with existing contrastive learning architectures and learn from just two views. Related work is often times limited to empirical analysis only, or considers the number of positive views as just one of many hyperparameters in their ablation studies. In contrast, we provide a holistic foundation for theorists and practitioners alike. We derive our theoretical formulation from the Information Bottleneck theorem, and provide an extensive empirical analysis through our experiments.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Theoretical Formulation on the Use of Multiple Positive Views in Contrastive Learning

  • Zhehao Liang,
  • Yanqing Luo,
  • Marvin Beese,
  • Patrik Reiske

摘要

Contrastive learning is a potent technique for self-supervised learning that maintains invariance between two positive views. Advancements such as the “core view” or multi-cropping have harnessed insights from multiple views, culminating in the latest state-of-the-art performance. However, the complexities of learning from multiple positive views remain partially unexplored. In this work, we provide a theoretical formulation on the use of multiple (two or more) positive views in contrastive learning, as well as a “plug-and-play” approach for learning from multiple positive views that seamlessly integrates with existing contrastive learning architectures and learn from just two views. Related work is often times limited to empirical analysis only, or considers the number of positive views as just one of many hyperparameters in their ablation studies. In contrast, we provide a holistic foundation for theorists and practitioners alike. We derive our theoretical formulation from the Information Bottleneck theorem, and provide an extensive empirical analysis through our experiments.