<p>This paper studies the problem of Traffic Signal Control (TSC) for multiple intersections. A large-scale TSC algorithm based on multi-expert demonstrations and Multi-Agent Reinforcement Learning (MARL) called EXPs-XLight is proposed. In contrast to other human-in-the-loop expert demonstration approaches that rely on a single expert, a mechanism for multi-expert demonstrations is introduced to accelerate the training of large-scale multi-agents, reduce training difficulty, and improve the overall training effectiveness. Expert knowledge is derived from multiple sources, using Max Pressure (MP) and Max Queue-Length (M-QL) as expert policies. By combining diverse experiences from multiple experts, agents are able to learn from a broader range of traffic scenarios. During the process of learning from expert knowledge, a supervised large margin classification loss is introduced to encourage the learning of meaningful action values. EXPs-XLight incorporates mixed policy sampling, allowing dynamic adjustment of the balance between expert guidance and agents’ own experiences. Unlike previous methods that involve expert participation throughout entire episodes, EXPs-XLight enables partial expert participation during agent exploration. An <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\epsilon\)</EquationSource> </InlineEquation>-greedy algorithm based on multiple experts is introduced to encourage agents to explore novel state-action pairs while avoiding over-reliance on expert policies. To preserve the agents’ capacity for exploration and autonomous learning, their own strategies are consistently utilized in interactions with the environment. In addition, a replay buffer discrimination mechanism is introduced to ensure the accumulation of high-quality experience by storing interactions with higher rewards. EXPs-XLight has demonstrated excellent performance through experiments on real-world datasets, including a 16-intersection road network in Hangzhou and a 196-intersection road network in New York, as well as a large-scale simulated road network with 1000 intersections.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A deep reinforcement learning model for traffic signal control with multi-expert participating in exploration

  • Hui Zhang,
  • Zhicheng Zhou,
  • Shiyi Gu,
  • Ya Zhang

摘要

This paper studies the problem of Traffic Signal Control (TSC) for multiple intersections. A large-scale TSC algorithm based on multi-expert demonstrations and Multi-Agent Reinforcement Learning (MARL) called EXPs-XLight is proposed. In contrast to other human-in-the-loop expert demonstration approaches that rely on a single expert, a mechanism for multi-expert demonstrations is introduced to accelerate the training of large-scale multi-agents, reduce training difficulty, and improve the overall training effectiveness. Expert knowledge is derived from multiple sources, using Max Pressure (MP) and Max Queue-Length (M-QL) as expert policies. By combining diverse experiences from multiple experts, agents are able to learn from a broader range of traffic scenarios. During the process of learning from expert knowledge, a supervised large margin classification loss is introduced to encourage the learning of meaningful action values. EXPs-XLight incorporates mixed policy sampling, allowing dynamic adjustment of the balance between expert guidance and agents’ own experiences. Unlike previous methods that involve expert participation throughout entire episodes, EXPs-XLight enables partial expert participation during agent exploration. An \(\epsilon\) -greedy algorithm based on multiple experts is introduced to encourage agents to explore novel state-action pairs while avoiding over-reliance on expert policies. To preserve the agents’ capacity for exploration and autonomous learning, their own strategies are consistently utilized in interactions with the environment. In addition, a replay buffer discrimination mechanism is introduced to ensure the accumulation of high-quality experience by storing interactions with higher rewards. EXPs-XLight has demonstrated excellent performance through experiments on real-world datasets, including a 16-intersection road network in Hangzhou and a 196-intersection road network in New York, as well as a large-scale simulated road network with 1000 intersections.