<p>Autonomous Underwater Vehicle (AUV) docking is a critical technology for long-term autonomous operation, but its reliability is severely challenged by uncertain underwater environments, such as visual blur, varying illumination, and partial occlusions. Traditional vision-based methods often struggle with semantic understanding and exhibit limited robustness in these complex scenarios. To address this, we propose a novel framework, LLM-Augmented Semantic Reasoning for Robust AUV Docking (DockLLM), which pioneers the integration of multi-modal large models’ semantic generalization capabilities into the underwater docking task. Our approach fuses features from a lightweight visual encoder with a frozen Vision-Language Model (VLM) to achieve a high-level semantic understanding of the docking scene. A Large Language Model (LLM) then acts as a high-level policy coordinator, decomposing the docking mission into natural language sub-tasks and dynamically generating or refining low-level control policies. Furthermore, we introduce an uncertainty-aware prompting strategy that enables the LLM to output confidence scores and alternative actions, enhancing the system’s fault tolerance. This “perception-understanding-decision” closed-loop paradigm compensates for the deficiencies of purely data-driven models in few-shot and long-tail scenarios. Extensive simulated and real-world experiments demonstrate our method’s significant superiority over existing approaches. Specifically, in real-world physical deployments with natural turbidity, DockLLM achieves a 90.0% overall docking success rate (45/50), drastically outperforming the 70.0% (35/50) achieved by traditional deep learning baselines. Furthermore, under simulated heavy occlusion scenarios, our framework maintains a robust 89% success rate, representing a substantial improvement over the 58% success rate of advanced State-of-the-Art (ViT+RL) data-driven architectures.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LLM-augmented semantic reasoning for robust AUV docking under uncertain environments

  • Bolun Zhang,
  • Canjun Yang

摘要

Autonomous Underwater Vehicle (AUV) docking is a critical technology for long-term autonomous operation, but its reliability is severely challenged by uncertain underwater environments, such as visual blur, varying illumination, and partial occlusions. Traditional vision-based methods often struggle with semantic understanding and exhibit limited robustness in these complex scenarios. To address this, we propose a novel framework, LLM-Augmented Semantic Reasoning for Robust AUV Docking (DockLLM), which pioneers the integration of multi-modal large models’ semantic generalization capabilities into the underwater docking task. Our approach fuses features from a lightweight visual encoder with a frozen Vision-Language Model (VLM) to achieve a high-level semantic understanding of the docking scene. A Large Language Model (LLM) then acts as a high-level policy coordinator, decomposing the docking mission into natural language sub-tasks and dynamically generating or refining low-level control policies. Furthermore, we introduce an uncertainty-aware prompting strategy that enables the LLM to output confidence scores and alternative actions, enhancing the system’s fault tolerance. This “perception-understanding-decision” closed-loop paradigm compensates for the deficiencies of purely data-driven models in few-shot and long-tail scenarios. Extensive simulated and real-world experiments demonstrate our method’s significant superiority over existing approaches. Specifically, in real-world physical deployments with natural turbidity, DockLLM achieves a 90.0% overall docking success rate (45/50), drastically outperforming the 70.0% (35/50) achieved by traditional deep learning baselines. Furthermore, under simulated heavy occlusion scenarios, our framework maintains a robust 89% success rate, representing a substantial improvement over the 58% success rate of advanced State-of-the-Art (ViT+RL) data-driven architectures.