<p>Parallel Disassembly Sequence Planning (DSP) deals with obtaining the order of disassembling the product with multiple parts disassembling simultaneously. Existing studies derive the product’s disassembly order based on the requirement for liaison data. However, the absence of stability data does not ensure that all the product parts are disassembled without damage. Since the order of retrieval of the product can also be obtained from stability data, this reduces the requirement of an initial precedence graph based on liaison data. The complexity of the problem increases if we adopt different representations for precedence and stability conditions of the product for constraint checking throughout the disassembly process. This paper constructs a stability-based precedence graph based on only the product’s stability data. The developed graph represents the precedence of the stable parts for disassembly, and the same is utilized to formulate the problem as a model-free deep Reinforcement Learning (RL) problem. A novel Exploration–Exploitation&#xa0;(EE) strategy-based Deep Q-Learning&#xa0;(DQL) algorithm is proposed to parallel DSP minimizing direction changes. Our effective reward mechanism accelerates the learning process for all the conditions and objectives of the problem considered. Based on past performance, the proposed algorithm dynamically decides the exploration and exploitation strategy to avoid local optima. We applied the proposed algorithm on five different-scale products and compared the results with uniform-sampling-based model-free deep RL algorithms. It showed significant improvements, particularly in cumulative rewards and disassembling all the parts for large-scale products with early convergence.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Novel Deep Reinforcement Learning Approach for Stability-Based Parallel Disassembly Sequence Planning Problem

  • M. R. Mahesh Kumar,
  • Chandrasekar Ravi

摘要

Parallel Disassembly Sequence Planning (DSP) deals with obtaining the order of disassembling the product with multiple parts disassembling simultaneously. Existing studies derive the product’s disassembly order based on the requirement for liaison data. However, the absence of stability data does not ensure that all the product parts are disassembled without damage. Since the order of retrieval of the product can also be obtained from stability data, this reduces the requirement of an initial precedence graph based on liaison data. The complexity of the problem increases if we adopt different representations for precedence and stability conditions of the product for constraint checking throughout the disassembly process. This paper constructs a stability-based precedence graph based on only the product’s stability data. The developed graph represents the precedence of the stable parts for disassembly, and the same is utilized to formulate the problem as a model-free deep Reinforcement Learning (RL) problem. A novel Exploration–Exploitation (EE) strategy-based Deep Q-Learning (DQL) algorithm is proposed to parallel DSP minimizing direction changes. Our effective reward mechanism accelerates the learning process for all the conditions and objectives of the problem considered. Based on past performance, the proposed algorithm dynamically decides the exploration and exploitation strategy to avoid local optima. We applied the proposed algorithm on five different-scale products and compared the results with uniform-sampling-based model-free deep RL algorithms. It showed significant improvements, particularly in cumulative rewards and disassembling all the parts for large-scale products with early convergence.