Machine learning based classification of catastrophic health expenditures: a cross-sectional study of Korean low-income households
摘要
Despite the National Health Insurance (NHI) system implemented in South Korea, concerns persist regarding access to health coverage for low-income households. To address this issue, this study aims to use machine learning-based data mining techniques to classify whether such households will face catastrophic health expenditures (CHEs).
MethodsA total of 4,031 low-income people were extracted using 2019 data from the Korea Health Panel Survey. The classification model was developed using four machine learning algorithms: Random Forest, Gradient boosting, Decision tree, Ridge regression, Neural network, and AdaBoost. Ten-fold cross validation was carried out to ensure the reliability of the analysis results. The model was evaluated based on the Area Under Receiver Operating Characteristics (AUROC) as well as accuracy, precision, recall, and F-1 score.
ResultsThe study’s findings revealed that the incidence of CHE was 26.2% in low-income households. The AdaBoost model had the highest classifiable power. It showed AUROC of 89.8%, accuracy of 83.1%, precision of 82.4%, recall of 83.1, and F1 score of 82.1%. The study found that economic activity, chronic disease, and age were significant factors that could lead to CHEs. Therefore, individuals over 65, with chronic conditions, and unemployed had the highest likelihood of developing CHE.
ConclusionIt is essential to identify low-income households that are at risk of CHEs in advance before facing the economic burden. This research is expected to provide fundamental data that can aid in developing an integrated support program to prevent and manage CHEs more effectively.