Intelligent Multi-zone Residential HVAC Control Strategy Based on Deep Reinforcement Learning
摘要
In this chapter, a novel data-driven method, which is called the deep deterministic policy gradient (DDPG), is applied for optimally controlling the multi-zone residential heating, ventilation, and air conditioning (HVAC) system. The DDPG method is a type of model-free deep reinforcement learning (deep RL) method that can generate HVAC control strategies without referring to any complex modeling formulation. The applied deep RL–based method can learn the optimal control strategy through continuous interaction with the simulated building environment. Simulation results of DDPG on real-world use cases and comparisons with the benchmark cases demonstrate the effectiveness and the generalization ability of DDPG in saving energy cost while maintaining occupant comfort, which proves its feasibility in solving real-world high-dimensional control problems with hidden information or vast solution spaces.