<p>This research aims to find a schedule that adapts to the variations in the transport system that influence the traveled time by the buses. For this purpose, the stop-skipping control strategy is adopted with the goal to serve all passengers waiting at stations with a minimum total delay in a well known static environment. We introduce a novel measure to calculate the maximum total delay of passengers based on the notion of balancing the load inside the buses, denoted as <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\({\text { load-delay }}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mspace width="0.333333em" /> <mtext>load-delay</mtext> <mspace width="0.333333em" /> </mrow> </math></EquationSource> </InlineEquation>. To solve the stop-skipping decision problem with minimizing the <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\({\text { load-delay }}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mspace width="0.333333em" /> <mtext>load-delay</mtext> <mspace width="0.333333em" /> </mrow> </math></EquationSource> </InlineEquation>, we model it as a distributed game where the stations are the players. We explore three stateless reinforcement learning algorithms: Linear Reward Inaction (LRI), Upper Confidence Bound (UCB), and <InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(\epsilon \)</EquationSource> <EquationSource Format="MATHML"><math> <mi>ϵ</mi> </math></EquationSource> </InlineEquation>-Greedy, and compare their performance in terms of converge quality and speed. We propose an enhanced LRI denoted as LRI <InlineEquation ID="IEq4"> <EquationSource Format="TEX">\({\text { sub-strategy }}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mspace width="0.333333em" /> <mtext>sub-strategy</mtext> <mspace width="0.333333em" /> </mrow> </math></EquationSource> </InlineEquation> based on limiting the action set based on previous decisions. We validate the effectiveness of our proposed solution in comparison with the other RL methods through numerical experiments using real-world data.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Performance Investigation of Stateless RL Methods for Optimal Bus Scheduling

  • Perla Hajjar,
  • Leïla Kloul,
  • Dominique Barth

摘要

This research aims to find a schedule that adapts to the variations in the transport system that influence the traveled time by the buses. For this purpose, the stop-skipping control strategy is adopted with the goal to serve all passengers waiting at stations with a minimum total delay in a well known static environment. We introduce a novel measure to calculate the maximum total delay of passengers based on the notion of balancing the load inside the buses, denoted as \({\text { load-delay }}\) load-delay . To solve the stop-skipping decision problem with minimizing the \({\text { load-delay }}\) load-delay , we model it as a distributed game where the stations are the players. We explore three stateless reinforcement learning algorithms: Linear Reward Inaction (LRI), Upper Confidence Bound (UCB), and \(\epsilon \) ϵ -Greedy, and compare their performance in terms of converge quality and speed. We propose an enhanced LRI denoted as LRI \({\text { sub-strategy }}\) sub-strategy based on limiting the action set based on previous decisions. We validate the effectiveness of our proposed solution in comparison with the other RL methods through numerical experiments using real-world data.