<p>Fall detection has played an important role in reducing the risk of disability among the elderly, alleviating pressure on public healthcare systems, and minimizing property damage. With the advancement of deep learning, vision-based fall detection has become an increasingly prominent research topic. However, many existing methods overlook the importance of short-term spatio-temporal motion information during fall events and fail to effectively integrate global and local features for comprehensive interaction. In this paper, we propose a dual-branch fall detection algorithm based on multi-scale temporal differences (MSTD-DB-FD). We build our method on the Vision Transformer (ViT) baseline. To improve local motion understanding, we design a dual-branch network that extracts spatio-temporal features at multiple scales and fuses multi-representation information. One branch uses ViT to capture global semantics and long-range dependencies. The other branch uses CNN to focus on local key features and short-term temporal information. This branch also complements the Transformer branch in the temporal dimension. We design two modules: the multi-scale spatio-temporal aggregation attention (MSAA) module and the multi-representation temporal difference feature fusion (MTDF) module. After experimental validation, the proposed method achieves an accuracy of 99.52% on the UR dataset and 99.21% on the Le2i dataset.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dual-branch fall detection algorithm based on multi-scale temporal difference

  • Xi Cai,
  • Yinuo Chen,
  • Mingchen Fan,
  • Guang Han

摘要

Fall detection has played an important role in reducing the risk of disability among the elderly, alleviating pressure on public healthcare systems, and minimizing property damage. With the advancement of deep learning, vision-based fall detection has become an increasingly prominent research topic. However, many existing methods overlook the importance of short-term spatio-temporal motion information during fall events and fail to effectively integrate global and local features for comprehensive interaction. In this paper, we propose a dual-branch fall detection algorithm based on multi-scale temporal differences (MSTD-DB-FD). We build our method on the Vision Transformer (ViT) baseline. To improve local motion understanding, we design a dual-branch network that extracts spatio-temporal features at multiple scales and fuses multi-representation information. One branch uses ViT to capture global semantics and long-range dependencies. The other branch uses CNN to focus on local key features and short-term temporal information. This branch also complements the Transformer branch in the temporal dimension. We design two modules: the multi-scale spatio-temporal aggregation attention (MSAA) module and the multi-representation temporal difference feature fusion (MTDF) module. After experimental validation, the proposed method achieves an accuracy of 99.52% on the UR dataset and 99.21% on the Le2i dataset.