<p>3D hand-object interaction pose reconstruction is crucial for applications such as human–computer interaction and virtual reality. However, reconstructing poses from single images presents challenges due to the complexity and flexibility of hand joints, as well as mutual occlusion with objects. Existing methods often fail to capture the complex details of hand-object interactions fully. To address these issues, we propose the HOR2H framework, which synergistically extracts hand and object features from a single RGB image to reconstruct 3D hand-object interaction poses. Our framework comprises two key modules: HOPSE for keypoint location enhancement and HOPSR for pose and shape reconstruction. HOPSE leverages semantic graph convolution to enhance the recognition of keypoints locations, effectively inferring occluded keypoints and reducing ambiguity in pose estimation. HOPSR integrates global context and local details to reconstruct accurate hand-object interaction poses and shapes through two parallel processing branches. Experiments on the HO-3D and H2O-3D datasets demonstrate that our method achieves advanced results, outperforming a range of previous methods in both hand and object pose estimation tasks. Compared with the baseline method, on HO-3D, PHJE decreased by 11.5%, and PHME decreased by 3.7%. On H2O-3D, L-PHJE, L-PHME, R-PHJE, and R-PHME decreased by 6.5%, 8.8%, 9.4% and 4.3%, respectively. Our approach not only ensures complete information on hand-object interaction details but also enhances the model’s generalization and ability to model complex interactions. Source code and pretrained models are available at <a href="https://github.com/HOIMiao/HOR2H">https://github.com/HOIMiao/HOR2H</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing hand-object interaction pose reconstruction through semantic-enhanced and reconstruction modules

  • Wenji Yang,
  • Xingyang Miao

摘要

3D hand-object interaction pose reconstruction is crucial for applications such as human–computer interaction and virtual reality. However, reconstructing poses from single images presents challenges due to the complexity and flexibility of hand joints, as well as mutual occlusion with objects. Existing methods often fail to capture the complex details of hand-object interactions fully. To address these issues, we propose the HOR2H framework, which synergistically extracts hand and object features from a single RGB image to reconstruct 3D hand-object interaction poses. Our framework comprises two key modules: HOPSE for keypoint location enhancement and HOPSR for pose and shape reconstruction. HOPSE leverages semantic graph convolution to enhance the recognition of keypoints locations, effectively inferring occluded keypoints and reducing ambiguity in pose estimation. HOPSR integrates global context and local details to reconstruct accurate hand-object interaction poses and shapes through two parallel processing branches. Experiments on the HO-3D and H2O-3D datasets demonstrate that our method achieves advanced results, outperforming a range of previous methods in both hand and object pose estimation tasks. Compared with the baseline method, on HO-3D, PHJE decreased by 11.5%, and PHME decreased by 3.7%. On H2O-3D, L-PHJE, L-PHME, R-PHJE, and R-PHME decreased by 6.5%, 8.8%, 9.4% and 4.3%, respectively. Our approach not only ensures complete information on hand-object interaction details but also enhances the model’s generalization and ability to model complex interactions. Source code and pretrained models are available at https://github.com/HOIMiao/HOR2H.