<p>Despite significant advancements in robotic grasp driven by visual perception, deploying robots in unstructured environments to perform user-specified tasks still poses considerable challenges. Natural language offers an intuitive means of specifying task objectives, reducing ambiguity. In this study, we introduced natural language into a vision-guided grasp system by employing visual attributes as a mediating bridge between language instructions and visual observations. We propose a command-driven semantic grasp architecture that integrates pixel attention within the visual attribute recognition module and includes a modified grasp pose estimation network to enhance prediction accuracy. Our experimental results show that our approach improves the performance of the submodules including visual attribute recognition and grasp pose estimation compared to baseline models. Furthermore, we demonstrate that our proposed model exhibits notable effectiveness in real-world user-specified grasping experiments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Command-driven semantic robotic grasping towards user-specified tasks

  • Qing Lyu,
  • Qingwen Ye,
  • Xiaoyan Chen,
  • Qiuju Zhang

摘要

Despite significant advancements in robotic grasp driven by visual perception, deploying robots in unstructured environments to perform user-specified tasks still poses considerable challenges. Natural language offers an intuitive means of specifying task objectives, reducing ambiguity. In this study, we introduced natural language into a vision-guided grasp system by employing visual attributes as a mediating bridge between language instructions and visual observations. We propose a command-driven semantic grasp architecture that integrates pixel attention within the visual attribute recognition module and includes a modified grasp pose estimation network to enhance prediction accuracy. Our experimental results show that our approach improves the performance of the submodules including visual attribute recognition and grasp pose estimation compared to baseline models. Furthermore, we demonstrate that our proposed model exhibits notable effectiveness in real-world user-specified grasping experiments.