The integration of Large Language Models (LLMs) into Reinforcement Learning (RL) presents a promising approach to enhancing agent performance by leveraging the world knowledge embedded in LLMs. However, existing approaches face significant challenges, including high computational costs, over-dependence on pre-trained knowledge, and limited adaptability to dynamic environments. In this paper, we propose Offline-to-Online: Case-Based Knowledge Distillation with Large Language Models for Reinforcement Learning (OCaL-RL). Our framework features an LLM-driven case-based reasoning system, which constructs a case repository of simulated experiences enriched with world knowledge. The agent first retrieves and leverages relevant cases for initial policy optimization and then refines its policies based on real environmental interactions, ensuring both knowledge efficiency and adaptability. This case-driven paradigm enables RL agents to dynamically retrieve, adapt, and refine distilled policies, balancing pre-trained knowledge and real-time learning. Experiments in challenging MiniGrid environments demonstrate that OCaL-RL outperforms strong baselines in both return rate and success rate.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Offline-to-Online: Case-Based Knowledge Distillation with Large Language Models for Reinforcement Learning

  • Hongzhe Liu,
  • Quan Liu,
  • Lan Wu,
  • Meilong Shi,
  • Zhiming Cui

摘要

The integration of Large Language Models (LLMs) into Reinforcement Learning (RL) presents a promising approach to enhancing agent performance by leveraging the world knowledge embedded in LLMs. However, existing approaches face significant challenges, including high computational costs, over-dependence on pre-trained knowledge, and limited adaptability to dynamic environments. In this paper, we propose Offline-to-Online: Case-Based Knowledge Distillation with Large Language Models for Reinforcement Learning (OCaL-RL). Our framework features an LLM-driven case-based reasoning system, which constructs a case repository of simulated experiences enriched with world knowledge. The agent first retrieves and leverages relevant cases for initial policy optimization and then refines its policies based on real environmental interactions, ensuring both knowledge efficiency and adaptability. This case-driven paradigm enables RL agents to dynamically retrieve, adapt, and refine distilled policies, balancing pre-trained knowledge and real-time learning. Experiments in challenging MiniGrid environments demonstrate that OCaL-RL outperforms strong baselines in both return rate and success rate.