Efficient Visual Reinforcement Learning for IoT Control from Pixels
摘要
Deep reinforcement learning plays an important role in perception and decision-making of Internet of Things (IoT). In comparison with traditional reward-driven feature learning, it is more efficient to use contrastive learning as an auxiliary task to extract high-level discriminative features from images in Reinforcement Learning (RL). In this paper, we further promote the sample-efficiency of contrastive learning for RL. We present DCRL, Discrete Contrastive Representation Learning for Reinforcement Learning (DCRL). DCRL learns unsupervised representation via discrete representation learning under the framework of contrastive learning. Compared with previous pixel-based RL methods, the learned representation of DCRL is much more compact and representative. As a consequence, DCRL significantly improves the sample-efficiency of RL, as demonstrated by our experiments on the DeepMind Control suite. In addition, encouraging results demonstrate the superiority of DCRL over other strong baselines.