<p>In this study, we propose a method for learning a latent space representing 6-DoF poses and performing 6-DoF control in the latent space using NewtonianVAE. NewtonianVAE, a type of world models based on Variational Autoencoder (VAE), can learn the dynamics of the environment as a latent space from observational data and perform proportional control based on the estimated position on the latent space. However, previous research has not demonstrated 6-DoF pose estimation and control using NewtonianVAE. Therefore, we propose 6D NewtonianVAE, which extends the latent space by incorporating the rotation vector to construct the latent space representing 6-DoF poses and perform 6-DoF control based on the estimated poses. Experimental results showed that our method achieves 6-DoF control with an accuracy within 7&#xa0;mm and 0.02 rad in a real-world. It was also shown that 6-DoF control is possible even in unseen environments. Our approach enables end-to-end 6-DoF pose estimation and control without annotated data. It also eliminates the need for RGB-D or point cloud data and relies solely on RGB images, reducing implementation and computational costs.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

6D NewtonianVAE: 6-DoF object pose estimation and control method for robotic tasks via learning from multi-view visual information

  • Mai Terashima,
  • Ryo Okumura,
  • Pedro Miguel Uriguen Eljuri,
  • Katsuyoshi Maeyama,
  • Yuanyuan Jia,
  • Tadahiro Taniguchi

摘要

In this study, we propose a method for learning a latent space representing 6-DoF poses and performing 6-DoF control in the latent space using NewtonianVAE. NewtonianVAE, a type of world models based on Variational Autoencoder (VAE), can learn the dynamics of the environment as a latent space from observational data and perform proportional control based on the estimated position on the latent space. However, previous research has not demonstrated 6-DoF pose estimation and control using NewtonianVAE. Therefore, we propose 6D NewtonianVAE, which extends the latent space by incorporating the rotation vector to construct the latent space representing 6-DoF poses and perform 6-DoF control based on the estimated poses. Experimental results showed that our method achieves 6-DoF control with an accuracy within 7 mm and 0.02 rad in a real-world. It was also shown that 6-DoF control is possible even in unseen environments. Our approach enables end-to-end 6-DoF pose estimation and control without annotated data. It also eliminates the need for RGB-D or point cloud data and relies solely on RGB images, reducing implementation and computational costs.