Hand activity recognition (HAR) based on a 3D hand skeleton is intuitive. It is widely applied in building systems that simulate hand activities and applications that guide blind people to grasp objects while saving memory because it does not require much input information. However, it also faces many challenges due to data noise and lack of information. In this study, we conduct a comparative study based on fine-tuning the activity recognition models (ISTA-Net, DDNet, MS-G3D, and PA-ResGCN) with 3D hand skeleton as input data captured from egocentric vision camera (EVC), the models are fine-tuned and evaluated on the HOI4D benchmark dataset. The recognition results show that PA-ResGCN has the best accuracy, with Acc being 99.44%. This result is close to 1.0, which can be applied to research models to build real applications. The recognition results are presented in detail and compared based on the available confusion matrix.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative Study of Hand Activity Recognition from Egocentric 3D Hand Pose

  • Nguyen Thi Loan,
  • Ninh Quang Tri,
  • Do Huu Son,
  • Pham Thi Thuy Linh,
  • Le Van Hung

摘要

Hand activity recognition (HAR) based on a 3D hand skeleton is intuitive. It is widely applied in building systems that simulate hand activities and applications that guide blind people to grasp objects while saving memory because it does not require much input information. However, it also faces many challenges due to data noise and lack of information. In this study, we conduct a comparative study based on fine-tuning the activity recognition models (ISTA-Net, DDNet, MS-G3D, and PA-ResGCN) with 3D hand skeleton as input data captured from egocentric vision camera (EVC), the models are fine-tuned and evaluated on the HOI4D benchmark dataset. The recognition results show that PA-ResGCN has the best accuracy, with Acc being 99.44%. This result is close to 1.0, which can be applied to research models to build real applications. The recognition results are presented in detail and compared based on the available confusion matrix.