Constrained Robotic Navigation on Preferred Terrains Using LLMs and Speech Instruction: Exploiting the Power of Adverbs
摘要
This paper explores leveraging large language models for mapless offroad navigation using generative AI, thereby eliminating the need for data collection and annotation. We propose a method where the robot receives speech instructions, converted to text through Whisper, and a large language model (LLM) like GPT model extracts landmarks, preferred terrains, and crucial adverbs translated into speed settings for constrained navigation. A language-driven semantic segmentation model generates text-based masks for identifying the required landmarks and terrains in images. By translating 2D image points to the vehicle’s motion plane using camera parameters, an MPC controller guides the vehicle towards the desired terrains. This approach enables adaptability to diverse environments and enhances instructions for navigating complex and challenging terrains.