错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Uncertainty-based bootstrapped optimization for offline reinforcement learning

  • Tianyi Li,
  • Genke Yang,
  • Jian Chu

摘要

Offline reinforcement learning (offline RL) promises to learn effective policies from previously-collected, static datasets without offering further possibility for exploration. However, offline RL encounters significant challenges primarily due to algorithmic difficulties arising from function approximation errors caused by extrapolating from out-of-distribution (OOD) data points. In this work, we propose uncertainty-based bootstrapped optimization (UBO), which aims to address the distributional shift induced by the fixed datasets. First, we take advantage of the bootstrapped architecture to implicitly approximate the epistemic uncertainty for the training instances. Then, we apply both the implicit and explicit penalties to the OOD data with high prediction uncertainties. Finally, we introduce a training paradigm based on the upper confidence bound (UCB) strategy for the bootstrapping updates, which enables the algorithm to thoroughly assess the varying performance of each bootstrapped head. We compare UBO with other prevailing offline RL algorithms on D4RL benchmarks. Experiments on various tasks demonstrate that the proposed algorithm can outperform or be competitive with the previous state-of-the-art on most of the tasks.