Optimizing balanced accuracy in medical data threshold models
摘要
We present a new approach to the problem of balancing data classes to optimize balanced accuracy in categorical models for which the classes are defined by applying a threshold to a real-valued dependent variable in the dataset. Many examples of such threshold problems are found in the context of medical conditions prefixed by “hypo” or “hyper”, such as hypothyroidism or hypertension. While our viewpoint and featured applications involve medical datasets, our methods apply to general threshold-defined classification problems. We offer a complete analytical solution for the special case in which the dataset is approximately bivariate normal, and apply it to a classical medical dataset of Galton relating parent and child heights. Our solution is expressed as a formula involving the Gaussian error function, which may be easily applied to real-world datasets computationally, as we demonstrate via the accompanying software. We also discuss approximations, estimations, and generalizations of this result involving deviations from normality in real-world datasets, and compare our class of threshold problems to other types of threshold problems appearing in the literature. Finally, we discuss how our solution relates to the goal of streamlining costly empirical data balancing procedures in machine learning models.