Enhancing Repetitive Action Counting Through Hierarchical Transformer-Based Radar-Vision Fusion
摘要
In this paper, we explore the complex task of repetitive action counting in computer vision, where conventional sampling often misses complete action details, especially in videos with variable action cycles. Traditional vision-based methods use dynamic resolution, which leads to high computational costs and limited performance in poor lighting. To improve this, we introduce a radar-vision fusion network that combines millimeter-wave radar with video data, effectively handling varied action cycles. We develop a novel hierarchical Transformer network that integrates range-Doppler and time-Doppler radar spectrums with video, filling in gaps left by visual-only methods. This approach not only enhances local frame-level and global sequence-level motion sensing but also significantly improves learning action features from videos. Our experimental results confirm the effectiveness of our proposed method in enhancing repetitive action counting.