Feature Selection and Comparison of Logistic Regression and Random Forest for Stability Assessment of Landslide Dams
摘要
Reliable stability assessment of landslide dams (LDs) remains challenging due to the scarcity and heterogeneity of available data, which hinder the generalization of both empirical formulas and single-model approaches. This study addresses this issue by proposing a reproducible and interpretable machine learning framework that integrates physically meaningful composite features with ensemble and statistical modeling. A global dataset of 135 LD cases was compiled, from which four dimensionless composite features and one material indicator were constructed to represent fundamental geometric and hydrological properties. Logistic regression (LR) and random forest (RF) were adopted as complementary models that combine interpretability and nonlinear generalization. Under stratified cross-validation and independent testing, both models consistently outperformed empirical formulas, confirming the effectiveness of the composite features. Additional classifiers including support vector machine (SVM) and neural network (NN) validated the robustness and transferability of the proposed feature design. Application to the Tangjiashan LD demonstrated that LR can capture gradual probability transitions, whereas RF identifies threshold-like behaviors during instability onset. The results indicate that for small and heterogeneous LD datasets, model reliability depends not only on algorithm selection but also on data engineering and feature construction. The proposed hybrid framework establishes a reproducible foundation for interpretable and data-driven LD stability assessment and provides a scalable basis for future integration of spatial and temporal information in dynamic risk evaluation.