A Benchmark for Rule Induction in Automated Business Decisions
摘要
We consider rule induction as a valuable tool in developing automated business decision services. For this use case, we present first results towards a comprehensive benchmark of rule induction pipelines. In this paper, we focus on typical binary business decision classification problems. The chosen pipelines include classical algorithms as well as promising recent approaches. The data sets vary substantially in terms of their characteristics such as imbalance, noise, size and label complexity. Our results suggest that out-of-the-box rule induction often generates short rule sets within acceptable training time that have a predictive performance close to the non-interpretable reference xgboost. Our results also show some shortcomings and open problems. We also present several synthetic experiments that differentiate the selected pipelines further.