<p>Sample efficiency and systematic generalization remain persistent challenges in reinforcement learning. While previous studies demonstrate that incorporating natural language with other observation modalities enhances generalization and sample efficiency due to its compositional and open-ended nature, effectively leveraging these properties requires robust language grounding mechanisms. To address this, we propose <Emphasis Type="BoldItalic">I</Emphasis><i>nstruction</i> <Emphasis Type="BoldItalic">C</Emphasis><i>onditioned</i> <Emphasis Type="BoldItalic">MO</Emphasis><i>dular network</i> (<b>ICMO</b>) by introducing <i>language entrance</i> and <i>memory feedback</i> techniques on top of an existing modular and sparse architecture, NPS. The memory feedback mechanism aggregates high-level information, guides selective attention in NPS via attentional feedback, and strengthens the decision-making process in the presence of language guidance. ICMO achieves superior performance compared to previous methods during our rigorous experiments, demonstrating near-zero generalization gap that highlights its robustness. Additionally, an extensive ablation study confirms the contributions of these techniques to improving generalization, sample efficiency, and training stability.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Inductive biases for zero-shot systematic generalization in language-informed reinforcement learning

  • Negin Hashemi Dijujin,
  • Seyed Roozbeh Razavi Rohani,
  • Mohammad Mahdi Samiei,
  • Mahdieh Soleymani Baghshah

摘要

Sample efficiency and systematic generalization remain persistent challenges in reinforcement learning. While previous studies demonstrate that incorporating natural language with other observation modalities enhances generalization and sample efficiency due to its compositional and open-ended nature, effectively leveraging these properties requires robust language grounding mechanisms. To address this, we propose Instruction Conditioned MOdular network (ICMO) by introducing language entrance and memory feedback techniques on top of an existing modular and sparse architecture, NPS. The memory feedback mechanism aggregates high-level information, guides selective attention in NPS via attentional feedback, and strengthens the decision-making process in the presence of language guidance. ICMO achieves superior performance compared to previous methods during our rigorous experiments, demonstrating near-zero generalization gap that highlights its robustness. Additionally, an extensive ablation study confirms the contributions of these techniques to improving generalization, sample efficiency, and training stability.