With the increasing severity of urban traffic congestion, adaptive traffic signal control based on deep reinforcement learning has been widely studied. Q-value based Deep Q Network reinforcement learning algorithms have been introduced for various types of expansion due to its effective efficacy. However, most of the studies design rewards and algorithms for specific application scenarios, without considering the impact of sampling method in experience replay during DRL-based traffic signal control model training. Moreover, Traffic control is often a multi-objective optimization problem, which contains traffic-participant pedestrians and vehicles. Most studies design comprehensive rewards to solve multi-objective problem without considering the sample size for each target in multi-objective. To address these, this paper proposes a soft-constraint reward for demands of pedestrian crossing, and a distributed multi-head policy DRL-based traffic signal control method based on cluster sampling for pedestrians and vehicles, Cluster Sampling-Multi Head Dueling Deep Q Network. For multi-intersection traffic signal control, each agent at intersection contains two policy networks, collects the local and neighboring intersection states from environments and controls phase duration according to two policies respectively. The experiences produced in environments exploring are stored separately. For each policy network, experience will be clustered by rewards, and then sampled for training. It can make experience replay in line with reward distribution, which improves model stability and generalization. In experiments, Cluster Sampling-Multi Head Dueling Deep Q Network is compared with existing traffic signal control methods in the real traffic environment of xuancheng to verify its performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Distributed Multi Head Policy Traffic Signal Control Method Based on Cluster Sampling for Efficient Training

  • Fuhao Yu,
  • Xiwei Mi,
  • Mei Han,
  • Yanchun Huang,
  • Yanhui Han,
  • Chao Chen

摘要

With the increasing severity of urban traffic congestion, adaptive traffic signal control based on deep reinforcement learning has been widely studied. Q-value based Deep Q Network reinforcement learning algorithms have been introduced for various types of expansion due to its effective efficacy. However, most of the studies design rewards and algorithms for specific application scenarios, without considering the impact of sampling method in experience replay during DRL-based traffic signal control model training. Moreover, Traffic control is often a multi-objective optimization problem, which contains traffic-participant pedestrians and vehicles. Most studies design comprehensive rewards to solve multi-objective problem without considering the sample size for each target in multi-objective. To address these, this paper proposes a soft-constraint reward for demands of pedestrian crossing, and a distributed multi-head policy DRL-based traffic signal control method based on cluster sampling for pedestrians and vehicles, Cluster Sampling-Multi Head Dueling Deep Q Network. For multi-intersection traffic signal control, each agent at intersection contains two policy networks, collects the local and neighboring intersection states from environments and controls phase duration according to two policies respectively. The experiences produced in environments exploring are stored separately. For each policy network, experience will be clustered by rewards, and then sampled for training. It can make experience replay in line with reward distribution, which improves model stability and generalization. In experiments, Cluster Sampling-Multi Head Dueling Deep Q Network is compared with existing traffic signal control methods in the real traffic environment of xuancheng to verify its performance.