错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Agile Translation of Human Motion to Virtual Scene Using Machine Vision for Industry 5.0

  • Songze Li,
  • Fadi Assad,
  • Xing Wu

摘要

Industry 5.0 promotes manufacturing systems in which human operators, intelligent machines, and digital twins are tightly integrated. However, many existing digital-twin workflows still lack a fast and inexpensive way to bring realistic human motions into virtual factory models. This paper develops a lightweight, camera-based pipeline that converts monocular RGB video into standard skeletal animation files, which can then be imported into an industrial simulation environment. The workflow combines state-of-the-art 2D pose estimation with temporal 3D pose lifting and uses Blender for motion retargeting before loading the animated avatar into Visual Components. To assess the practical utility of the pipeline, we record synchronized motion-capture and video data using a Vicon system and a single RGB camera. Six key joints (shoulder, elbow, wrist, hip, knee, ankle) are compared between the Visual Components scene and the Vicon ground truth. Across the evaluation sequence, the average positional RMSE is about 16.6 mm, and no joint exhibits a maximum frame-wise error greater than 50 mm. In addition, runtime measurements show that the end-to-end processing can reach approximately 12 frames per second on commodity hardware, providing near-real-time throughput. The results confirm that a single-camera, learning-based motion capture pipeline can achieve sufficient accuracy and efficiency for human-in-the-loop industrial simulations. The proposed approach offers a low-cost route to embed real human actions into digital twins for tasks such as ergonomic analysis, safety training, and cycle-time evaluation, thereby supporting the human-centric vision of Industry 5.0.