Deep reinforcement learning plays an important role in perception and decision-making of Internet of Things (IoT). In comparison with traditional reward-driven feature learning, it is more efficient to use contrastive learning as an auxiliary task to extract high-level discriminative features from images in Reinforcement Learning (RL). In this paper, we further promote the sample-efficiency of contrastive learning for RL. We present DCRL, Discrete Contrastive Representation Learning for Reinforcement Learning (DCRL). DCRL learns unsupervised representation via discrete representation learning under the framework of contrastive learning. Compared with previous pixel-based RL methods, the learned representation of DCRL is much more compact and representative. As a consequence, DCRL significantly improves the sample-efficiency of RL, as demonstrated by our experiments on the DeepMind Control suite. In addition, encouraging results demonstrate the superiority of DCRL over other strong baselines.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Efficient Visual Reinforcement Learning for IoT Control from Pixels

  • Haitao Wang,
  • Hejun Wu,
  • Qiang Tu

摘要

Deep reinforcement learning plays an important role in perception and decision-making of Internet of Things (IoT). In comparison with traditional reward-driven feature learning, it is more efficient to use contrastive learning as an auxiliary task to extract high-level discriminative features from images in Reinforcement Learning (RL). In this paper, we further promote the sample-efficiency of contrastive learning for RL. We present DCRL, Discrete Contrastive Representation Learning for Reinforcement Learning (DCRL). DCRL learns unsupervised representation via discrete representation learning under the framework of contrastive learning. Compared with previous pixel-based RL methods, the learned representation of DCRL is much more compact and representative. As a consequence, DCRL significantly improves the sample-efficiency of RL, as demonstrated by our experiments on the DeepMind Control suite. In addition, encouraging results demonstrate the superiority of DCRL over other strong baselines.