LUCA: Look, Understand, Communicate, and Act – A Vision-Language System for Object-Oriented Navigation
摘要
We present LUCA: Look, Understand, Communicate, and Act – a novel vision-language system for object-oriented navigation in mobile robots. LUCA integrates state-of-the-art vision-language models (VLMs) with RGB-D perception and autonomous navigation, enabling robots to understand and execute natural language commands in real-world environments. Unlike traditional approaches that rely on pre-defined object classes or semantic maps, our system uses the open-vocabulary capabilities of GPT-4o to identify objects described in natural language. We implement a distributed architecture where computationally intensive vision and language processing occurs on a remote client while the resource-constrained mobile robot handles local sensing and actuation. Our system is validated in real-world tasks involving object-oriented navigation and active object search, demonstrating the feasibility of VLM-guided robot control on resource-constrained platforms. The results show promising performance without the need for task-specific training or onboard GPU computation, making LUCA a modular and extensible approach for natural language-based robot navigation.