The task of compositional action recognition holds significant importance in the field of video understanding; however, the issue of static bias severely limits the generalization capability of models. Existing models often overly rely on sensitive features in videos, such as object appearance and background morphology, for action recognition, without fully leveraging true temporal action features, leading to recognition errors when faced with novel object-action combinations. To address this issue, this paper proposes an innovative framework for compositional action recognition, utilizing Spatio-Temporal contrastive learning to construct a three-branch architecture that distinguishes appearance and spatiotemporal features at the feature extraction stage. The model is encouraged to contrast features that predict factual probabilities with those that predict biased probabilities through contrastive learning, thereby reducing the direct and indirect reliance on sensitive features and enhancing the accuracy and generalization of recognition. Experimental results show that this method achieves state-of-the-art performance on the Something-Else dataset, validating its effectiveness in composite action recognition tasks. Furthermore, it achieves comparable or superior results to state-of-the-art methods on standard action recognition datasets such as Something-Something-V2, UCF101, and HMDB51.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Spatio-Temporal Contrastive Learning for Compositional Action Recognition

  • Yezi Gong,
  • Mingtao Pei

摘要

The task of compositional action recognition holds significant importance in the field of video understanding; however, the issue of static bias severely limits the generalization capability of models. Existing models often overly rely on sensitive features in videos, such as object appearance and background morphology, for action recognition, without fully leveraging true temporal action features, leading to recognition errors when faced with novel object-action combinations. To address this issue, this paper proposes an innovative framework for compositional action recognition, utilizing Spatio-Temporal contrastive learning to construct a three-branch architecture that distinguishes appearance and spatiotemporal features at the feature extraction stage. The model is encouraged to contrast features that predict factual probabilities with those that predict biased probabilities through contrastive learning, thereby reducing the direct and indirect reliance on sensitive features and enhancing the accuracy and generalization of recognition. Experimental results show that this method achieves state-of-the-art performance on the Something-Else dataset, validating its effectiveness in composite action recognition tasks. Furthermore, it achieves comparable or superior results to state-of-the-art methods on standard action recognition datasets such as Something-Something-V2, UCF101, and HMDB51.