<p>Occupational accidents represent a serious social problem in Brazil, often causing deaths of workers or permanent incapacity to work. Therefore, it is necessary to study the statistical distribution of such accidents in the country and evaluate the possibility of predicting the occurrence of these phenomena. In the present work, we obtain an integrated dataset containing socioeconomic variables from multiple sources, such as statistical surveys and government data, performing an exploratory analysis of these variables and the occupational accidents in the country. These variables are then used as features to predict the number of accidents in Brazilian municipalities. In the proposed regression problem, we use linear regression as a baseline predictor and decision tree, artificial neural networks (ANN), and gradient boosting algorithms (GradientBoosting, LightGBM, XGBoost, and CatBoost), with the adoption of Bayesian optimization to tune the hyperparameters of the algorithms. In addition, a feature selection technique based on the importance of the variables is adopted to evaluate which variables present the greatest significance in the predictions. The models show high performance and explainability, with the <i>R</i>-squared metric reaching values greater than 0.90. The results may help the government and enterprises adopt preventive actions to avoid occupational accidents, reducing the country’s human and social security costs.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Predicting occupational accidents in Brazilian cities: a machine learning-based approach

  • J. M. Toledo,
  • T. J. M. Moura

摘要

Occupational accidents represent a serious social problem in Brazil, often causing deaths of workers or permanent incapacity to work. Therefore, it is necessary to study the statistical distribution of such accidents in the country and evaluate the possibility of predicting the occurrence of these phenomena. In the present work, we obtain an integrated dataset containing socioeconomic variables from multiple sources, such as statistical surveys and government data, performing an exploratory analysis of these variables and the occupational accidents in the country. These variables are then used as features to predict the number of accidents in Brazilian municipalities. In the proposed regression problem, we use linear regression as a baseline predictor and decision tree, artificial neural networks (ANN), and gradient boosting algorithms (GradientBoosting, LightGBM, XGBoost, and CatBoost), with the adoption of Bayesian optimization to tune the hyperparameters of the algorithms. In addition, a feature selection technique based on the importance of the variables is adopted to evaluate which variables present the greatest significance in the predictions. The models show high performance and explainability, with the R-squared metric reaching values greater than 0.90. The results may help the government and enterprises adopt preventive actions to avoid occupational accidents, reducing the country’s human and social security costs.