FDDet: Frequency-Decoupling for Boundary Refinement in Temporal Action Detection
摘要
Temporal action detection aims to locate and classify actions in untrimmed videos. While recent works exploit powerful pre-trained video encoders, they often suffer from background noise and semantic redundancy, which hinder precise localization. To tackle this issue, we propose FDDet, a frequency-aware decoupling network that enhances action representations by filtering irrelevant background semantics and preserving fine-grained motion patterns. Specifically, we design an adaptive temporal decoupling mechanism to suppress noisy signals while retaining atomic action details, and a category-aware relation module to capture both local transitions and long-range dependencies. These refined representations are then fed into a detection head for accurate action prediction. Extensive experiments on THUMOS14, HACS, and ActivityNet-1.3 demonstrate that FDDet, powered by InternVideo2-6B features, consistently outperforms state-of-the-art methods in temporal action detection.