<p>Egocentric manipulation systems, including humanoid robots, frequently suffer from severe self-occlusion caused by their own arms or grippers, resulting in partial observability and unstable visuomotor control. Imitation learning-based policies that rely on short-horizon visual inputs are especially vulnerable to such transient loss of task-relevant scene information. To address this challenge, we introduce a dual mechanism for improving workspace observability under self-occlusion. First, we employ Dynamic Multi-View Integration (DMVI) by mounting additional cameras on the wrist to secure task-critical local viewpoints near the end-effector. Second, we propose a Spatio-Temporal Point Cloud Memory (ST-PCM) that preserves previously observed but currently occluded regions as a confidence-aware 3D memory, gradually attenuating outdated information over time. The confidence modeling enables the policy to distinguish between fresh observations and residual memory. Since DMVI enhances spatial coverage while ST-PCM provides temporal continuity, their combination offers complementary strengths. We validate the proposed framework on real-world single-arm manipulation tasks designed to emulate self-occlusion-prone egocentric manipulation scenarios, demonstrating improved robustness to self-occlusion in both task success rate and execution efficiency.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Occlusion-Aware 3D Visuomotor Policy Via Dynamic Multi-View Integration and Spatio-Temporal Point Cloud Memory

  • Giseong Kwon,
  • Jungwook Lee,
  • Chibum Lee,
  • Ki-In Na

摘要

Egocentric manipulation systems, including humanoid robots, frequently suffer from severe self-occlusion caused by their own arms or grippers, resulting in partial observability and unstable visuomotor control. Imitation learning-based policies that rely on short-horizon visual inputs are especially vulnerable to such transient loss of task-relevant scene information. To address this challenge, we introduce a dual mechanism for improving workspace observability under self-occlusion. First, we employ Dynamic Multi-View Integration (DMVI) by mounting additional cameras on the wrist to secure task-critical local viewpoints near the end-effector. Second, we propose a Spatio-Temporal Point Cloud Memory (ST-PCM) that preserves previously observed but currently occluded regions as a confidence-aware 3D memory, gradually attenuating outdated information over time. The confidence modeling enables the policy to distinguish between fresh observations and residual memory. Since DMVI enhances spatial coverage while ST-PCM provides temporal continuity, their combination offers complementary strengths. We validate the proposed framework on real-world single-arm manipulation tasks designed to emulate self-occlusion-prone egocentric manipulation scenarios, demonstrating improved robustness to self-occlusion in both task success rate and execution efficiency.