Background: In recent decades, the rate of caesarean section (C-section) has been rapidly increasing in Bangladesh. This study explores different machine learning techniques to find out potential determinants of C-section in the country. Methods: Data on C-section from 4932 women were extracted from the clustered Bangladesh Demographic Health Survey (BDHS) 2017–18. Chi-square test was applied for feature selection. The machine learning techniques applied included decision tree, random forest, K-nearest neighbors, Gaussian Naïve-Bayes, logistic regression, and support vector machine. The performance of these techniques was evaluated using confusion matrices and receiver operating characteristics (ROC) curves. Results: The prevalence of C-section was found to be 33.0% (1628 out of 4932). Age of mother, wealth index, working status of mother, birth order, number of antenatal care visits, body mass index, birth weight, education, media access, birth interval, place of residence, and administrative division were significant features for the prediction of C-section. Among the ML models, the random forest (accuracy = 0.8143, precision = 0.7907, sensitivity = 0.7977, specificity = 0.8477, F1-score = 0.7939, and area under curve: AUC = 0.90) was found to be the best choice for predicting the C-section among Bangladeshi women. Conclusion: Study findings may help government and non-government organizations, health professionals, and policy-makers to take necessary actions in reducing unnecessary C-section childbirth.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Predicting C-Section Outcomes Among Bangladeshi Women: A Comparative Study of Machine Learning Techniques

  • Faisal Haque,
  • Sabikun Nahar,
  • Umme Marzia Haque,
  • Zakir Hossain,
  • Enamul Kabir

摘要

Background: In recent decades, the rate of caesarean section (C-section) has been rapidly increasing in Bangladesh. This study explores different machine learning techniques to find out potential determinants of C-section in the country. Methods: Data on C-section from 4932 women were extracted from the clustered Bangladesh Demographic Health Survey (BDHS) 2017–18. Chi-square test was applied for feature selection. The machine learning techniques applied included decision tree, random forest, K-nearest neighbors, Gaussian Naïve-Bayes, logistic regression, and support vector machine. The performance of these techniques was evaluated using confusion matrices and receiver operating characteristics (ROC) curves. Results: The prevalence of C-section was found to be 33.0% (1628 out of 4932). Age of mother, wealth index, working status of mother, birth order, number of antenatal care visits, body mass index, birth weight, education, media access, birth interval, place of residence, and administrative division were significant features for the prediction of C-section. Among the ML models, the random forest (accuracy = 0.8143, precision = 0.7907, sensitivity = 0.7977, specificity = 0.8477, F1-score = 0.7939, and area under curve: AUC = 0.90) was found to be the best choice for predicting the C-section among Bangladeshi women. Conclusion: Study findings may help government and non-government organizations, health professionals, and policy-makers to take necessary actions in reducing unnecessary C-section childbirth.