Investigating Tailored Retraining for Online Process Predictions Using Log Features
摘要
Predictive process monitoring leverages predictive models to forecast outcomes of ongoing processes. Such predictive models must be trained and often retrained to sustain high performance and adapt to the dynamic environments in which modern processes operate. While much research has focused on detecting drifts in process behavior, limited attention has been given to leveraging drift detection for tailored retraining. Consequently, this leaves a gap in understanding the impact of using log features for retraining the predictive models. This paper addresses this gap by examining two key log features for triggering tailored retraining: label distribution and variant coverage. Using 23 variations of ten event logs, we evaluate the impact of retraining methods based on these features against three baselines. Our findings show retraining using specific log features yields modest but consistent improvement in performance. These insights contribute to the development of more resilient predictive models, highlighting the potential of tailored retraining methods to mitigate performance degradation in dynamic process environments.