<p>Existing electric vehicle (EV) charging management research aims to optimize charging power to reduce costs, considering the randomness of electricity prices and user behavior, but typically assumes fixed charging durations, ignoring the actual impact of traffic congestion on the dynamic changes in charging end times. This paper proposes a reinforcement learning-based approach that integrates the influence of real-time electricity prices and traffic congestion on charging durations. The problem is modeled as a discrete–continuous hybrid action space Markov decision process (MDP), where discrete actions determine the charging end time and continuous actions control the charging power; the INRIX index is used to quantify the level of traffic congestion, and the soft actor-critic (SAC) algorithm is improved (especially by restructuring the critic network) to address this hybrid action space issue. Simulations based on real data show that this method can effectively coordinate discrete and continuous actions and balance multiple reward objectives, significantly reducing congestion anxiety by about 18.36% and total costs by about 8.33% compared to benchmark solutions, highlighting its advantages.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deep Reinforcement Learning for Electric Vehicle Charging Management with Discrete–Continuous Hybrid Action Space

  • Qiang Zhao,
  • Yanan Jiang,
  • Hui Liu,
  • Yinghua Han

摘要

Existing electric vehicle (EV) charging management research aims to optimize charging power to reduce costs, considering the randomness of electricity prices and user behavior, but typically assumes fixed charging durations, ignoring the actual impact of traffic congestion on the dynamic changes in charging end times. This paper proposes a reinforcement learning-based approach that integrates the influence of real-time electricity prices and traffic congestion on charging durations. The problem is modeled as a discrete–continuous hybrid action space Markov decision process (MDP), where discrete actions determine the charging end time and continuous actions control the charging power; the INRIX index is used to quantify the level of traffic congestion, and the soft actor-critic (SAC) algorithm is improved (especially by restructuring the critic network) to address this hybrid action space issue. Simulations based on real data show that this method can effectively coordinate discrete and continuous actions and balance multiple reward objectives, significantly reducing congestion anxiety by about 18.36% and total costs by about 8.33% compared to benchmark solutions, highlighting its advantages.