Bcgn: BLIP-based cross-modal grasping network for language-conditioned robotic grasping
摘要
The performance of robots on the language-conditioned robotic grasping task reflects the intelligence level of robots. However, existing approaches lack the ability to handle implicit instructions and identify infeasible ones, which undermines the intelligence and operational safety of the robot. To overcome the above limitations, this paper introduces a novel Language-conditioned Robotic Grasping Dataset (LRGD), which covers a variety of instruction types. Correspondingly, an end-to-end BLIP-based Cross-modal Grasping Network (BCGN) for language-conditioned grasping is proposed. Specifically, BCGN integrates BLIP to jointly model cross-modal information, and introduces a learnable circuit breaker that enables the model to actively reject infeasible requests. Furthermore, through collaboration with LVLMs (Large Vision-Language Models), BCGN can easily achieve zero-shot recognition of implicit instructions. Experimental results the LRGD and in real-world scenarios demonstrate the effectiveness of BCGN in dealing with instructions of different complexity levels.