Self-supervised contrastive learning of sentence representation with difficulty-based sampling
摘要
Contrastive learning is an effective method of self-supervised representation learning, which has made significant strides in the application of sentence representation learning in recent years. However, most previous studies typically focus on generating effective positive samples through various data augmentation methods, overlooking the fact that negative samples are often randomly sampled from the training data. This practice may introduce false negatives leading to negative sampling bias, which can consequently impair the performance of contrastive learning models. This issue has garnered considerable attention, yet current solutions have some limitations in dealing with negative sampling bias in a relatively straightforward way. To tackle these challenges, this paper proposes a