Explainable machine learning identifies diagnostic patterns for paediatric respiratory diseases in a high-dimensional underrepresented South African dataset
摘要
This study presents a clinically-derived, real-world paediatric pulmonology dataset across underrepresented areas in South Africa and benchmarks machine learning models for the diagnostic classification of three prevalent respiratory conditions: asthma, bronchiectasis, and bronchopulmonary dysplasia.
MethodsThe dataset comprises of 2,176 patient records with approximately 300 clinical features and 95 diagnostic labels, reflecting the multi-label nature of routine specialist care. The records used were from the period 2010 to 2024. Using a stratified 70–30 train-test split with controlled cross-validation and hyperparameter tuning, we evaluated logistic regression, linear support vector machines, decision trees, random forests, gradient-boosted trees, LightGBM, XGBoost, CatBoost, a multilayer perceptron, and an ensemble CatBoost neural network model.
ResultsAcross all diseases, modern boosted tree models and the ensemble achieved consistently strong performance, with AUC values up to 0.975 for asthma, 0.947 for bronchiectasis, and 0.961 for BPD, and high precision–recall performance under class imbalance. Probabilistic scores and calibration analyses further highlighted differences in probability reliability between models, even when discrimination was similar. This is relevant for clinical screening and decision-support where reliable risk estimates are necessary. We further evaluated alternative decision thresholds and used decision-curve analysis to assess clinical usefulness, showing that the operating point can be aligned with an intended clinical objective and that the strongest models provided positive net benefit over default strategies. Lastly, model interpretability using SHAP-based explanations identified internally consistent predictor patterns for each disease target.
ConclusionsThese results demonstrate that a robust and internally consistent classification signal is present in a real-world, multi-label paediatric respiratory dataset and that explainable machine learning evaluation can support transparent comparison of modelling approaches in complex clinical environments.