错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improved VLN-BERT with Reinforcing Endpoint Alignment for Vision-and-Language Navigation

  • Chuan Jin,
  • Boyuan Yang,
  • Ruonan Liu

摘要

Vision-and-Language Navigation (VLN) refers to an agent navigating a real-world environment by understanding natural language instructions and utilizing visual information from the surroundings. Currently, many pre-trained models and pre-training tasks have been proposed to assist agents in navigating unfamiliar environments using visual and linguistic information. However, ensuring that the agent stops near the endpoint is a challenging problem. Building on the existing VLN-BERT model, we propose an improved VLN-BERT model with a new pre-training task called Reinforcing Endpoint Alignment (REA-VLN-BERT). Through this pre-training task, the model can more effectively align the endpoint in the path with the corresponding instruction without requiring any additional data. Further experiments show that the reinforcing endpoint alignment task leads to improvements of 0.78% and 2.51% in Success Rate (SR) on the seen and unseen validation sets of the R2R dataset, respectively. Furthermore, inspired by Airbert, we combine shuffling loss with the reinforcing endpoint alignment task, resulting in a new model named SREA-VLN-BERT. SREA-VLN-BERT achieves improvements of 3.53% and 0.94% in SR on the seen and unseen validation sets of the R2R dataset, respectively, further enhancing the model’s average performance in VLN.