<p>The rapid rise of industrial automation has accelerated the deployment of autonomous mobile robots for material handling, inspection, and collaborative operations. Effective performance in dynamic environments demands path-planning algorithms that ensure safety, energy efficiency, and adaptability—balancing global optimality with adaptive reactivity. No single algorithmic solution fully satisfies these requirements, necessitating hybrid frameworks that integrate the global optimization capability of metaheuristics with the adaptive learning of reinforcement learning. To address these challenges, we introduce RL-PFWOA, a novel hybrid hierarchical framework featuring bidirectional feedback between a metaheuristic core and a reinforcement learning agent. Global exploration is achieved through a hybrid Pufferfish optimization (PFO) and whale optimization algorithm (WOA) strategy, where the WOA coefficient <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\textbf{A}\)</EquationSource> <EquationSource Format="MATHML"><math> <mi mathvariant="bold">A</mi> </math></EquationSource> </InlineEquation> dynamically alternates between encircling and spiral exploitation phases. A DDPG agent performs path refinement and adaptively modulates exploration via hybrid weights <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(\textbf{w}_{\text {PFO}}\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi mathvariant="bold">w</mi> <mtext>PFO</mtext> </msub> </math></EquationSource> </InlineEquation> and <InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(\textbf{w}_{\text {WOA}}\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi mathvariant="bold">w</mi> <mtext>WOA</mtext> </msub> </math></EquationSource> </InlineEquation>. A Bayesian weighting mechanism further balances path length, energy consumption, and traversal time, ensuring adaptive multi-objective optimization and preventing premature convergence. Benchmarking on ten standard functions (Sphere, Rastrigin, Schwefel, etc.) demonstrates superior convergence and minimal cost values compared to metaheuristic (PSO, GA, ABC, WOA) and reinforcement learning (PPO, SAC) baselines. MATLAB/Simulink 2023a simulations validate practical performance: RL-PFWOA yields the shortest paths (55.0&#xa0;m static, 58.1&#xa0;m dynamic), lowest energy use (9.21&#xa0;J static, 9.87&#xa0;J dynamic), and minimal obstacle collisions (1.5–2.1%). Ablation results confirm the superiority of the hybrid PFO+WOA core. Overall, RL-PFWOA achieves high stability (<InlineEquation ID="IEq4"> <EquationSource Format="TEX">\(\hbox {SD} = \pm 0.35\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mtext>SD</mtext> <mo>=</mo> <mo>±</mo> <mn>0.35</mn> </mrow> </math></EquationSource> </InlineEquation>), a 56% reduction in collisions, and 3.7% energy savings, establishing it as a robust and efficient framework for adaptive navigation in complex industrial settings.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adaptive path planning in complex environments for mobile robots via fusion of reinforcement learning and hybrid swarm intelligence

  • Subhash Yadav,
  • Shubham Shukla,
  • Mahendra Pratap Yadav,
  • Ritu Tiwari

摘要

The rapid rise of industrial automation has accelerated the deployment of autonomous mobile robots for material handling, inspection, and collaborative operations. Effective performance in dynamic environments demands path-planning algorithms that ensure safety, energy efficiency, and adaptability—balancing global optimality with adaptive reactivity. No single algorithmic solution fully satisfies these requirements, necessitating hybrid frameworks that integrate the global optimization capability of metaheuristics with the adaptive learning of reinforcement learning. To address these challenges, we introduce RL-PFWOA, a novel hybrid hierarchical framework featuring bidirectional feedback between a metaheuristic core and a reinforcement learning agent. Global exploration is achieved through a hybrid Pufferfish optimization (PFO) and whale optimization algorithm (WOA) strategy, where the WOA coefficient \(\textbf{A}\) A dynamically alternates between encircling and spiral exploitation phases. A DDPG agent performs path refinement and adaptively modulates exploration via hybrid weights \(\textbf{w}_{\text {PFO}}\) w PFO and \(\textbf{w}_{\text {WOA}}\) w WOA . A Bayesian weighting mechanism further balances path length, energy consumption, and traversal time, ensuring adaptive multi-objective optimization and preventing premature convergence. Benchmarking on ten standard functions (Sphere, Rastrigin, Schwefel, etc.) demonstrates superior convergence and minimal cost values compared to metaheuristic (PSO, GA, ABC, WOA) and reinforcement learning (PPO, SAC) baselines. MATLAB/Simulink 2023a simulations validate practical performance: RL-PFWOA yields the shortest paths (55.0 m static, 58.1 m dynamic), lowest energy use (9.21 J static, 9.87 J dynamic), and minimal obstacle collisions (1.5–2.1%). Ablation results confirm the superiority of the hybrid PFO+WOA core. Overall, RL-PFWOA achieves high stability ( \(\hbox {SD} = \pm 0.35\) SD = ± 0.35 ), a 56% reduction in collisions, and 3.7% energy savings, establishing it as a robust and efficient framework for adaptive navigation in complex industrial settings.