<p>Road safety studies benefit from analyzing crash injury severity. Developing reliable models predicting crash severity is crucial for law enforcement and policymakers. While statistical models dominate research on crash injury severity, they rely on assumptions prone to estimation errors. Machine learning techniques are rapidly emerging to overcome limitations. This study presents a methodology using tree-based machine learning algorithms to assess crash severity accurately. The process involves cleaning and preprocessing raw data, converting it for machine learning, performing feature engineering, data standardization, and handling missing values. Six ensemble models—bagging, random forest, gradient boosting, light boosting, and&#xa0;extreme gradient boosting—are tested, with hyperparameters optimized using Bayesian optimization. Model performance metrics include accuracy, precision, recall, f1-score, and model training time. Results show that categorical boosting outperformed other models with an f1-score of 0.839, requiring 6.27&#xa0;s to train. Random Forest had a lower f1-score of 0.806, and eXtreme gradient boosting trained in 0.246&#xa0;s with an f1-score of 0.807. Carriageway hazards are the prominent variable that has affected all models except LightGBM, whereas road type and day of the week have also shown significance in some models.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Traffic Crash-Severity Prediction using Interpretable Optimized Ensemble Models

  • Md Shafiul Alam,
  • Syed Masiur Rahman,
  • Karim Sattar,
  • Md Nurul Islam,
  • Md Shamimul Haque Choudhury,
  • Khaled Jamal Assi,
  • Nedal Taisir Al-Ratrout

摘要

Road safety studies benefit from analyzing crash injury severity. Developing reliable models predicting crash severity is crucial for law enforcement and policymakers. While statistical models dominate research on crash injury severity, they rely on assumptions prone to estimation errors. Machine learning techniques are rapidly emerging to overcome limitations. This study presents a methodology using tree-based machine learning algorithms to assess crash severity accurately. The process involves cleaning and preprocessing raw data, converting it for machine learning, performing feature engineering, data standardization, and handling missing values. Six ensemble models—bagging, random forest, gradient boosting, light boosting, and extreme gradient boosting—are tested, with hyperparameters optimized using Bayesian optimization. Model performance metrics include accuracy, precision, recall, f1-score, and model training time. Results show that categorical boosting outperformed other models with an f1-score of 0.839, requiring 6.27 s to train. Random Forest had a lower f1-score of 0.806, and eXtreme gradient boosting trained in 0.246 s with an f1-score of 0.807. Carriageway hazards are the prominent variable that has affected all models except LightGBM, whereas road type and day of the week have also shown significance in some models.