Catch before they fall: a pose-guided attention framework for indoor safety
摘要
Falls can pose a serious health threat, especially for older people, often leading to fractures, head injuries, or long-term disability. This highlights the need for reliable and non-invasive detection systems. Existing solutions often suffer from limitations such as user discomfort with wearables, constrained coverage of fixed sensors, or environmental challenges in vision-based methods. This paper proposes a dual-stream attention-guided robust indoor fall detection framework to solve the problem. The approach combines a pose-guided stream with a video stream, in which the former captures high-level skeletal features through MediaPipe and simultaneously processes spatiotemporal dynamics from raw RGB frames using a 3D Convolutional Neural Network (CNN) with a temporal attention module. In order to enhance classification precision, features from each stream are adaptively fused in a block based on attention mechanisms, which improves the model’s interpretation of posture and movement semantics. Testing on the KFall dataset reveals that the proposed method achieves 98.71% accuracy, surpassing pre-existing benchmarks. Triggering a visual alert by displaying a red rectangle on the screen upon the occurrence of this event is a subsequent outcome of this work. An ablation study highlights the effectiveness of each component. Finally, the proposed work advances fall detection by combining pose estimation and attention-based deep learning to deliver an accurate, interpretable, and deployable solution.