Reward Shaping for Job Shop Scheduling
摘要
Effective production scheduling is an integral part of the success of many industrial enterprises. In particular, the job shop problem (JSP) is highly relevant for flexible production scheduling in the modern era. Recently, numerous approaches for the JSP using reinforcement learning (RL) have been formulated. Different approaches employ different reward functions, but the individual effects of these reward functions on the achieved solution quality have received insufficient attention in the literature. We examine various reward functions using a novel flexible RL environment for the JSP based on the disjunctive graph approach. Our experiments show that a formulation of the reward function based on machine utilization is most appropriate for minimizing the makespan of a JSP among the investigated reward functions.