错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Sub-Instruction and Local Map Relationship Enhanced Model for Vision and Language Navigation

  • Yong Zhang,
  • Yinlin Li,
  • Jihe Bai,
  • Yi Feng,
  • Mo Tao

摘要

In this paper, different from most methods in vision and language navigation, which primarily relies on vision-language cross-modal attention modeling and the agent’s egocentric observations. We establish connections between sub-instructions and local maps to elaborately encode environment information and learn a path responsible for the whole instructions rather than the ultimate goal. We first obtain a local semantic map by ground projecting the RGB semantic segmentation map and the depth map. The segmented sub-instruction is passed through the sub-instruction attention module and then taken as input, together with the local map, to the cross-modal attention module. Finally, a set of waypoints are predicted by the navigation module until all sub-instructions of the long instruction are executed, which is the completion of an episode. Comparison experiments and ablation studies on the VLN-CE dataset show that our method outperforms most methods and has a good whole path predicting ability.