Monitoring motion and physiological signals in virtual reality (VR) is essential to understand user interaction and system adaptability. Contact-based sensors, such as photoplethysmography (PPG) and inertial measurement units (IMUs), provide reference data but introduce motion artifacts and usability constraints. Camera-based methods offer a non-intrusive alternative, enabling richer spatial and temporal analysis by capturing full-body dynamics and semantic understanding. This study presents a methodology for extracting motion and physiological signals from video, demonstrating strong agreement with reference signals. The evaluation is conducted using a database of nursing students performing VR-based training under real-world conditions. Ground truth signals were recorded using a Shimmer3 GSR+ sensor. Pose estimation models, including ViTPose, YOLO11 and Meta Sapiens, were evaluated for motion tracking and compared to IMU-based acceleration. Remote photoplethysmography (rPPG) estimated blood volume pulse from visible skin regions, validated against earlobe PPG signals. Given facial occlusion by the VR headset, segmentation models such as Meta SAMv2, Sapiens and YOLO11 were used for skin region tracking and head pose estimation. These findings provide insights into the viability and reliability of camera-based sensing in VR scenarios as a standalone or complementary tool for real-time behavioral analysis and human-machine interaction.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating the Accuracy and Reliability of Camera-Based Physiological and Motion Signal Extraction Techniques in Virtual Reality Training Environments

  • Tharindu Ekanayake,
  • Constantino Álvarez Casado,
  • Nhi Nguyen,
  • Marta Sobocinski,
  • Sari Pramila-Savukoski,
  • Xiaoting Wu,
  • Kristina Mikkonen,
  • Miguel Bordallo López

摘要

Monitoring motion and physiological signals in virtual reality (VR) is essential to understand user interaction and system adaptability. Contact-based sensors, such as photoplethysmography (PPG) and inertial measurement units (IMUs), provide reference data but introduce motion artifacts and usability constraints. Camera-based methods offer a non-intrusive alternative, enabling richer spatial and temporal analysis by capturing full-body dynamics and semantic understanding. This study presents a methodology for extracting motion and physiological signals from video, demonstrating strong agreement with reference signals. The evaluation is conducted using a database of nursing students performing VR-based training under real-world conditions. Ground truth signals were recorded using a Shimmer3 GSR+ sensor. Pose estimation models, including ViTPose, YOLO11 and Meta Sapiens, were evaluated for motion tracking and compared to IMU-based acceleration. Remote photoplethysmography (rPPG) estimated blood volume pulse from visible skin regions, validated against earlobe PPG signals. Given facial occlusion by the VR headset, segmentation models such as Meta SAMv2, Sapiens and YOLO11 were used for skin region tracking and head pose estimation. These findings provide insights into the viability and reliability of camera-based sensing in VR scenarios as a standalone or complementary tool for real-time behavioral analysis and human-machine interaction.