Human Activity Recognition (HAR) using wearable sensors has gained significant attention due to its portability and unobtrusiveness. However, the data obtained from wearable sensors are limited to inertial data from predefined locations on the human body. In contrast, skeletal data from motion capture devices, such as the Kinect camera, offer richer information by capturing the whole body dynamics of a human action. Unfortunately, the use of skeletal data is impractical in wearable sensor-based HAR for real-world deployment. Currently, transformer neural networks, known for their self-attention mechanism, have shown effective handling of data from diverse modalities in wearable sensor-based HAR. However, the deployment of multimodal transformer on wearable devices is challenging due to their inherent large model size. We propose a Lightweight HAR Transformer (LightHART) framework that trains an unimodal Inertial Transformer (IT) network by transferring knowledge from a large multimodal transformer using a knowledge distillation approach. We evaluate the proposed framework on three public multimodal human activity datasets and compare the performance of the LightHART student model with various state-of-the-art approaches. Experimental results demonstrate that our LightHART model achieves competitive performance in terms of effectiveness and scalability with a model size of only 1.43 Mb. We are the first to deploy and validate the LightHART fall detection model on a SmartFall App running on a WearOS-compatible smartwatch showcasing its potential in advancing wearable sensor-based HAR research.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LightHART: Lightweight Human Activity Recognition Transformer

  • Syed Tousiful Haque,
  • Jianyuan Ni,
  • Jingcheng Li,
  • Yan Yan,
  • Anne Hee Hiong Ngu

摘要

Human Activity Recognition (HAR) using wearable sensors has gained significant attention due to its portability and unobtrusiveness. However, the data obtained from wearable sensors are limited to inertial data from predefined locations on the human body. In contrast, skeletal data from motion capture devices, such as the Kinect camera, offer richer information by capturing the whole body dynamics of a human action. Unfortunately, the use of skeletal data is impractical in wearable sensor-based HAR for real-world deployment. Currently, transformer neural networks, known for their self-attention mechanism, have shown effective handling of data from diverse modalities in wearable sensor-based HAR. However, the deployment of multimodal transformer on wearable devices is challenging due to their inherent large model size. We propose a Lightweight HAR Transformer (LightHART) framework that trains an unimodal Inertial Transformer (IT) network by transferring knowledge from a large multimodal transformer using a knowledge distillation approach. We evaluate the proposed framework on three public multimodal human activity datasets and compare the performance of the LightHART student model with various state-of-the-art approaches. Experimental results demonstrate that our LightHART model achieves competitive performance in terms of effectiveness and scalability with a model size of only 1.43 Mb. We are the first to deploy and validate the LightHART fall detection model on a SmartFall App running on a WearOS-compatible smartwatch showcasing its potential in advancing wearable sensor-based HAR research.