Effective Reward Schemes for Tardiness Optimization
摘要
Over recent years, reinforcement learning has become a prominent method for the optimization of sequential decision-making problems. One group of sequential decision-making problems that has benefited significantly from reinforcement-learning-based optimization techniques is scheduling problems. However, most existing reinforcement learning works on scheduling optimization aim at optimizing a single, makespan-based objective. While the makespan—the overall time from the start of the first task to the end of the last task—is indeed important in some endeavors, other endeavors benefit more from the optimization of other types of objectives. In this work, we focus on Tardiness-based objectives and present a new reward scheme that aims at simultaneously optimizing multiple notions of Tardiness.