Towards Using Natural Language to Perform Robotic Tasks
摘要
This study explores automating robot programming using human-readable instructions, integrating textual and visual inputs. We present a framework combining a Visual Language Model (VLM), a vision processing model, and an adaptive skill library based on compliant control. The VLM converts textual instructions into executable commands linked to the skill library, while the vision model identifies and localizes referenced objects. This approach removes the need for additional training, enabling robots to execute tasks directly from natural language directives. Our method was evaluated using an internet-connected benchmarking device. It aims to streamline robot programming and enhance natural language communication in industrial and everyday settings.