Predicting Successful Treatment Completion Using Baseline Case Characteristics through Machine Learning and Ensemble Modeling: A Two-Step Approach
摘要
This study aims to enhance treatment completion prediction beyond the limitations of traditional logistic regression.
MethodsOur approach utilizes a two-step machine learning method, integrating cross-validation and ensemble modeling to achieve this goal. First, we employ a feature selection process using random forest to identify the most pivotal variables for our prediction model. Various models are then created using common machine learning algorithms, and a stacking approach combines the set into an ensemble model. Model selection is guided by a comprehensive assessment of performance and practical considerations. Our predictive model not only prioritizes accuracy but also provides insights into the impact of individual attributes on treatment success. By forecasting success, solely using baseline characteristics, researchers can assess participants’ likelihood of completing treatment before it starts, aiding cost reduction, especially in resource-intensive programs, by selecting individuals who are more likely to complete treatment. Additionally, emphasis on relevant variables helps identify areas for improving adherence and commitment to the treatment regimen.
ResultsOur approach’s validation involved a thorough assessment using a dataset comprising over 800 real child welfare cases, showcasing the practicality and resilience of our predictive model. The ensemble model adeptly strikes a balance between machine learning models and statistical logistic regression, rendering notable improvements in sensitivity and specificity, particularly in scenarios marked by imbalanced data.
ConclusionThis contribution marks a substantial stride forward in the area of enhancing decision-making and optimizing resource allocation within treatment program management.