Video-text retrieval is a crucial and challenging task that involves searching for the most semantically relevant items in a given query. Existing works typically adopt symmetric contrastive learning loss, which encourage positive sample pairs to be pulled together and the negative to be pushed away in a shared embedding space. However, it solely treats one caption manually collected for a specific video as relevant. Indeed, numerous videos are semantically similar to the caption, yet they are disregarded as negatives. Therefore, the binary correlation-based learning strategy causes semantically similar samples to be irrationally separated, which results in confusing learning objectives. In this paper, we first explore the behavior of symmetric contrastive learning loss from the perspective of gradients, and reveal that the model may suffer from uncertain semantic biases, and the accumulation of the biases would arise to semantic collisions. Motivated by the discovery, an effective and efficient Uncertain SEmantic Consistency constraint (USEC) is proposed, which leverages the uncertain latent semantic correlation between all the two cross-matched negative sample pairs as a supervisory signal to minimize the distance between semantically similar negative sample pairs in a self-supervised manner, while ensuring semantic consistency of positive sample pairs. Based on the gradient analysis and visualization of USEC, we theoretically demonstrate the rationality of exploiting neglected negative sample semantics during training. Evaluation on a series of benchmarks shows that USEC improves retrieval performance and exhibits merits with state-of-the-art methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Uncertainty-Aware with Negative Samples for Video-Text Retrieval

  • Weitao Song,
  • Weiran Chen,
  • Jialiang Xu,
  • Yi Ji,
  • Ying Li,
  • Chunping Liu

摘要

Video-text retrieval is a crucial and challenging task that involves searching for the most semantically relevant items in a given query. Existing works typically adopt symmetric contrastive learning loss, which encourage positive sample pairs to be pulled together and the negative to be pushed away in a shared embedding space. However, it solely treats one caption manually collected for a specific video as relevant. Indeed, numerous videos are semantically similar to the caption, yet they are disregarded as negatives. Therefore, the binary correlation-based learning strategy causes semantically similar samples to be irrationally separated, which results in confusing learning objectives. In this paper, we first explore the behavior of symmetric contrastive learning loss from the perspective of gradients, and reveal that the model may suffer from uncertain semantic biases, and the accumulation of the biases would arise to semantic collisions. Motivated by the discovery, an effective and efficient Uncertain SEmantic Consistency constraint (USEC) is proposed, which leverages the uncertain latent semantic correlation between all the two cross-matched negative sample pairs as a supervisory signal to minimize the distance between semantically similar negative sample pairs in a self-supervised manner, while ensuring semantic consistency of positive sample pairs. Based on the gradient analysis and visualization of USEC, we theoretically demonstrate the rationality of exploiting neglected negative sample semantics during training. Evaluation on a series of benchmarks shows that USEC improves retrieval performance and exhibits merits with state-of-the-art methods.