Reinforcement learning approach for decision making in disassembly processes
摘要
Remanufacturing represents a key strategy for the management of end-of-life products, with the potential to promote a circular economy. However, its implementation has been limited due to the labour-intensive and time-consuming nature of the disassembly processes required for component remanufacturing. The disassembly process has to cope with further uncertainties. The quality of the cores is often unknown, which can result in fluctuating processing times and random failures of cores or components during disassembly, due to damage. These uncertainties can have a significant impact on both on-time delivery and component service levels, both of which are costly and difficult to optimise simultaneously. In pursuit of multi-objective optimization, a novel reinforcement learning (RL) framework based on the Proximal Policy Optimization (PPO) algorithm has been formulated to inform decision-making processes during the disassembly process. Two distinct categories of RL agents have been developed to facilitate collaborative decision-making, namely one type for core decision and a second type for component decision. Furthermore, diverse configurations for merging these two types of RL agents have been explored. The approach aims to enable real-time decision-making during disassembly, potentially providing companies with an economic advantage. We compare our approach to heuristic control methods and the Deep-Q-Network (DQN) algorithm and prove the potential of the RL approach using the PPO algorithm through various test cases. Based on the reward function developed, the approach using the PPO algorithm received an approximately 12% higher reward than the DQN algorithm. In comparison to the Basic heuristic, it received an 22% higher reward and a 4% higher reward than the One Disassembly heuristic. This significantly increases adherence to on-time delivery and the service level. Simulations also demonstrate that the RL agent types can adapt to changing state values, resulting in high adaptivity and scalability.