Machine-Learning Imbalanced Datasets for Predicting Stillbirth Risk
摘要
Stillbirth, the death of a baby before or during delivery, is a significant global public health issue impacting maternal and fetal health. Stillbirth rates vary widely by region, influenced by factors like inadequate fetal growth, maternal diseases, and healthcare infrastructure. This study aims to investigate the factors contributing to stillbirth and develop a predictive model using machine learning to identify high-risk pregnancies. We utilized a dataset of 31,149 birth cases from the Pattani Provincial Public Health Office in Thailand, collected between 2018 and 2021, including 217 stillbirths and 30,932 live births. Machine-learning algorithms, such as Naïve Bayes, logistic regression, and random forest, were applied to build and evaluate predictive models. Given the dataset’s imbalanced nature, undersampling, and oversampling techniques were employed to balance the data for effective model training and testing. Results indicate that the random forest algorithm, combined with oversampling, achieved the best performance, significantly improving accuracy, recall, F-measure, precision, and the area under the receiver operating characteristic curve (AUC). This predictive model can be use to healthcare professionals with a valuable tool for early risk assessment and intervention in future. Our research highlights the potential of machine learning to enhance stillbirth risk prediction, thereby improving obstetric care quality and reducing stillbirth rates. Implementing such predictive models can play a crucial role in safeguarding maternal and fetal health, representing a significant advancement in public health strategies to combat stillbirth.