We present enhancements to the reinforcement learning (RL) approach used in the NeurIPS 2023 HomeRobot: Open Vocabulary Mobile Manipulation (OVMM) Challenge, focusing on augmenting the baseline model with advanced semantic segmentation and skill policy modifications. More specifically, we introduce refined semantic segmentation model (integrating the YOLOv8 and MobileSAM), improved place skill policy and a high-level heuristic strategy, which collectively advance the overall success rate from 0.8 to 5.2 (+550% relative) and the partial success rate from 9.7 to 25.8 (+165% relative) on the Test Standard split of the challenge dataset (ranked 2nd on the public leaderbord). These enhancements enabled our agent to achieve 3rd place in both the simulated and real-world stages of the competition. This paper details the strategies employed, discusses the insights gained, particularly in semantic segmentation and skill-specific training, and outlines potential avenues for future enhancements in embodied AI systems within open-vocabulary contexts.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Mobile Manipulation in Home Environments: A Case Study from the NeurIPS 2023 HomeRobot Challenge

  • Volodymyr Kuzma,
  • Vladyslav Humennyy,
  • Ruslan Partsey

摘要

We present enhancements to the reinforcement learning (RL) approach used in the NeurIPS 2023 HomeRobot: Open Vocabulary Mobile Manipulation (OVMM) Challenge, focusing on augmenting the baseline model with advanced semantic segmentation and skill policy modifications. More specifically, we introduce refined semantic segmentation model (integrating the YOLOv8 and MobileSAM), improved place skill policy and a high-level heuristic strategy, which collectively advance the overall success rate from 0.8 to 5.2 (+550% relative) and the partial success rate from 9.7 to 25.8 (+165% relative) on the Test Standard split of the challenge dataset (ranked 2nd on the public leaderbord). These enhancements enabled our agent to achieve 3rd place in both the simulated and real-world stages of the competition. This paper details the strategies employed, discusses the insights gained, particularly in semantic segmentation and skill-specific training, and outlines potential avenues for future enhancements in embodied AI systems within open-vocabulary contexts.