错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adaptive Social Robot Navigation Using Vision-Language Models

  • Ketan Anand,
  • Daymond Chang,
  • Nathaniel John

摘要

Achieving socially aware navigation in robotics remains a challenge, particularly for integrating embodied agents into dynamic human environments. Building upon NaviSTAR and the MuSoHu dataset, we introduce a novel approach to social robot navigation using language-guided methods. Our methodology employs a multi-modal model leveraging Vision-Language Models (VLMs) to decode high-level text commands and social cues from human interactions. This approach embraces the complexity of human-environment interactions through a diverse array of sensors and datasets. A significant innovation is the dynamic adaptability to user preferences via instruction tuning mechanisms, enabling the system to encompass a wide range of human navigation styles and bridge the gap between robotic functionality and human expectations. Our system is rigorously tested using real-world datasets and simulated environments, employing qualitative assessments and quantitative metrics, including a 5-point Likert scale. The incorporation of VLMs enhances the robot’s understanding of social nuances and human intentions, grounding comprehension in a rich Vision-Language context. This fosters improved situational awareness and nuanced decision-making capabilities, leading to more intuitive and responsive human-robot interactions. Our contributions include the Instruct-MuSoHu dataset, combining language guidance with multi-modal sensory data in dynamic environments, advancing socially aware navigation in complex social settings.