<p>Currently, computer vision-based human activity recognition (HAR) technology is becoming increasingly mature, but its inherent drawbacks such as invasiveness and dependence on lighting conditions are also becoming more prominent; thus, HAR technology based on WiFi channel state information (CSI) has become a research hotspot due to its advantages of low cost, wide applicability, and non-invasiveness. However, most existing HAR models primarily focus on extracting temporal features from CSI data, often neglecting its spatial features. Even some hybrid models that attempt to simultaneously extract spatio-temporal features still suffer from problems such as insufficient spatio-temporal feature extraction, inadequate spatio-temporal feature fusion, and not being lightweight enough. Therefore, this paper proposes a lightweight HAR model called MCMTNet, which utilizes a dual-stream architecture combining MC-Mamba and Transformer; the model is divided into a temporal stream and a carrier-channel stream, where the temporal stream employs cascaded MC-Mamba modules, composed of a Mamba model and multi-scale convolutional neural network, to model long-range dependencies in time series and achieve local temporal feature extraction from CSI data; the carrier-channel stream is based on a Transformer structure, focusing on modeling inter-subcarrier dependencies and inter-antenna correlations to achieve spatial feature extraction from CSI data; finally, the temporal and spatial features are effectively reduced in dimensionality and fused. Overall, the model leverages the Mamba architecture’s advantage of processing long sequences with linear time complexity and combines it with an efficient dual-stream design, thereby significantly reducing the model’s parameter count and computational complexity while ensuring high accuracy. The model was evaluated on two public datasets, and the results show that it achieved an activity recognition accuracy of 97.54% for 16 human activities on the Wiar dataset and 99.5% activity recognition accuracy on the UT-HAR dataset. Compared to several current advanced methods, the MCMTNet model achieves highly competitive results by enabling high-precision HAR while simultaneously maintaining a low parameter count and computational complexity, which demonstrates its significant practical application value.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MCMTNet: A Lightweight Dual-Stream Model for High-Accuracy WiFi-Based Human Activity Recognition

  • Zhaofei Li,
  • Zekun Yang,
  • Yijie Zhang,
  • Ruiyu Zheng,
  • Linhong Li

摘要

Currently, computer vision-based human activity recognition (HAR) technology is becoming increasingly mature, but its inherent drawbacks such as invasiveness and dependence on lighting conditions are also becoming more prominent; thus, HAR technology based on WiFi channel state information (CSI) has become a research hotspot due to its advantages of low cost, wide applicability, and non-invasiveness. However, most existing HAR models primarily focus on extracting temporal features from CSI data, often neglecting its spatial features. Even some hybrid models that attempt to simultaneously extract spatio-temporal features still suffer from problems such as insufficient spatio-temporal feature extraction, inadequate spatio-temporal feature fusion, and not being lightweight enough. Therefore, this paper proposes a lightweight HAR model called MCMTNet, which utilizes a dual-stream architecture combining MC-Mamba and Transformer; the model is divided into a temporal stream and a carrier-channel stream, where the temporal stream employs cascaded MC-Mamba modules, composed of a Mamba model and multi-scale convolutional neural network, to model long-range dependencies in time series and achieve local temporal feature extraction from CSI data; the carrier-channel stream is based on a Transformer structure, focusing on modeling inter-subcarrier dependencies and inter-antenna correlations to achieve spatial feature extraction from CSI data; finally, the temporal and spatial features are effectively reduced in dimensionality and fused. Overall, the model leverages the Mamba architecture’s advantage of processing long sequences with linear time complexity and combines it with an efficient dual-stream design, thereby significantly reducing the model’s parameter count and computational complexity while ensuring high accuracy. The model was evaluated on two public datasets, and the results show that it achieved an activity recognition accuracy of 97.54% for 16 human activities on the Wiar dataset and 99.5% activity recognition accuracy on the UT-HAR dataset. Compared to several current advanced methods, the MCMTNet model achieves highly competitive results by enabling high-precision HAR while simultaneously maintaining a low parameter count and computational complexity, which demonstrates its significant practical application value.