A model for predicting pseudospectral acceleration and peak ground acceleration utilizing supervised machine learning algorithms for seismically hazardous areas in India
摘要
Machine learning (ML) techniques offer major improvements for ground motion prediction in India's high seismic hazard zones—specifically Seismic Zones IV and V, encompassing the Himalayas, Indo-Gangetic Plain, and Kachchh. This study harnesses a dataset of 564 three-component acceleration records from 145 earthquakes (Mw 2.3–7.9) and 95 strong-motion stations to develop and benchmark XGBoost (eXtreme Gradient Boosting), LightGBM (Light Gradient Boosting Machine), and artificial neural network (ANN) models. The XGBoost model, trained with rigorous cross-validation strategies and explicit regularization, achieves excellent generalization (test R2 = 0.96, Pearson’s correlation coefficient ρ = 0.998), outperforming established ground motion prediction equations (GMPEs) and ANNs while capturing regional and site-specific variability. Model robustness and uncertainties are analyzed using RMSE, MAE, F1-Score, Bayesian Information Criterion (BIC), and comprehensive residual checks. The Bayesian Information Criterion (BIC) values obtained for the training and full datasets are −1710.24 and -2251.49, respectively. The substantial negative BIC values demonstrate that our XGBoost regression model achieves excellent predictive performance by balancing fit and simplicity effectively. The XGBoost approach demonstrates robust physical consistency but reveals elevated uncertainties for long-period/distant events, highlighting data-driven limitations and motivating further research. This ML-based framework offers substantial advances for seismic hazard assessment and resilient structural design tailored to India's most hazardous regions.