A Theoretical Formulation on the Use of Multiple Positive Views in Contrastive Learning
摘要
Contrastive learning is a potent technique for self-supervised learning that maintains invariance between two positive views. Advancements such as the “core view” or multi-cropping have harnessed insights from multiple views, culminating in the latest state-of-the-art performance. However, the complexities of learning from multiple positive views remain partially unexplored. In this work, we provide a theoretical formulation on the use of multiple (two or more) positive views in contrastive learning, as well as a “plug-and-play” approach for learning from multiple positive views that seamlessly integrates with existing contrastive learning architectures and learn from just two views. Related work is often times limited to empirical analysis only, or considers the number of positive views as just one of many hyperparameters in their ablation studies. In contrast, we provide a holistic foundation for theorists and practitioners alike. We derive our theoretical formulation from the Information Bottleneck theorem, and provide an extensive empirical analysis through our experiments.