Using Mixed Exponentials for Unsupervised Discretization
摘要
In this study, discretization, fundamental to numerous data analysis and modeling activities, is reconceptualized through an innovative, unsupervised approach. This method enhances the interpretability and simplifies the management of continuous variables, especially those exhibiting right-skewed distributions. By initially fitting a variable to a mixture of exponential distributions using the Expectation-Maximization algorithm, and determining discretization cutpoints via intersections of the component distributions, our proposed approach outperforms traditional methods like Equal width and Equal Frequency. The method significantly enhances classification accuracy across various machine learning models, including Random Forest, Neural Network, and Gradient Boosting Machine, with improvements ranging from approximately 5% to 25%. Furthermore, it increases mutual information (between 1.5 to 600 times) and consistently maintains a stability measure of one. These results were observed across the three datasets evaluated. The paper highlights the potential of this proposed discretization method, discusses its promising applications, and indicates future research directions, emphasizing its superior performance especially for right-skewed distributions.