Learning from Imbalanced Data in Healthcare: State-of-the-Art and Research Challenges
摘要
Datasets associated with medical and healthcare domains are imbalanced in nature. An imbalanced dataset refers to a classification dataset where the number of instances of a given class is much lower than for other classes. Such imbalanced datasets require special attention because traditional classifiers tend to favor the classes with many instances. On the other hand, in healthcare, the class with fewer instances may correspond to a rare and uncommon event of potential interest. Ignoring the imbalance affects the performance of classifiers which hampers the detection of rare cases such as disease screening, severity analysis, detecting adverse drug reactions, cancer malignancy grading, and the identification of uncommon, chronic illnesses in the population. Therefore, the design of classifiers should aim at successfully classifying such classes as compared to classifying other classes. Many significant research works accept the intrinsic imbalance in healthcare data and propose methodologies to handle this. Considering the importance and identifying research challenges in imbalance, a compilation of such works is the aim of this chapter.