错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Prototypical Transformer for Weakly Supervised Action Segmentation

  • Tao Lin,
  • Xiaobin Chang,
  • Wei Sun,
  • Weishi Zheng

摘要

Weakly supervised action segmentation aims to recognize the sequential actions in a video, with only action orderings as supervision for model training. Existing methods either predict the action labels to construct discriminative losses or segment the video based on action prototypes. In this paper, we propose a novel Prototypical Transformer (ProtoTR) to alleviate the defects of existing methods. The motivation behind ProtoTR is to further enhance the prototype-based method with more discriminative power for superior segmentation results. Specifically, the Prediction Decoder of ProtoTR translates the visual input into action ordering while its Video Encoder segments the video with action prototypes. As a unified model, both the encoder and decoder are jointly optimized on the same set of action prototypes. The effectiveness of the proposed method is demonstrated by its state-of-the-art performance on different benchmark datasets.