Point-Level Enhancement and Background Suppression Network for Point-Supervised Temporal Action Localization
摘要
Point-Supervised Temporal Action Localization (PTAL) focuses on accurately pinpointing the temporal boundaries of actions in untrimmed videos, leveraging only a single point label for each action occurrence. Existing methods have primarily focused on snippet-level or proposal-level tasks in isolation. This paper introduces a novel framework that concurrently optimizes both aspects. For snippet-level learning, addressing the inadequate leverage of point-level label semantic information in existing methods for predicting class activation sequences, we propose the Point-Level Enhancement and Background Suppression Network (PEBS-Net). PEBS-Net effectively captures inter-segment dependencies using point-level labels through an enhanced attention block, emphasizing foreground information. Furthermore, an online-updated memory module is introduced to store background segment prototypes for each class, suppressing background information and significantly improving the differentiation between foreground and background segments. For proposal-level learning, existing work primarily concentrates on frame-level offsets, neglecting segment-level offsets. This paper not only considers proposal boundary information but also leverages internal features within the proposal for further refinement, generating high-confidence proposals. Our network exhibits state-of-the-art performance on challenging benchmarks such as THUMOS14, GTEA, according to experimental results. This performance significantly surpasses all prior point-supervised methods and even exceeds that of certain competitive fully-supervised approaches.