Abstract <p>Air pollutants particularly PM2.5 have been shown to largely affect environments and human health across the world including Thailand. Air pollutants prediction therefore plays an important role in air pollution alerts and in management to keep pollutant under control. This study presents a comparative analysis of six supervised machine learning models—artificial neural networks, random forest, support vector machines, naïve Bayes, <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(k\)</EquationSource> <!--LobJMat2561115Tosasukul-m1--> </InlineEquation>-nearest neighbors, and extreme gradient boosting—for classifying PM2.5 concentration levels in Chiang Mai, northern part of Thailand. The models were trained and evaluated using weekly environmental data collected over 524 weeks from 2015 to 2024. Among them, the random forest model achieved the highest accuracy (72.82<InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(\%\)</EquationSource> <!--LobJMat2561115Tosasukul-m2--> </InlineEquation>). The ANN and NB models also demonstrated moderate to substantial agreement levels based on Cohen’s kappa statistics. Model interpretability was further enhanced using SHAP (Shapley additive explanations) analysis. The results indicate that the number of hotspot occurrences in Chiang Mai, Tak, and Mae Hong Son are the most influential predictors. These results underscore the value of ensemble tree-based methods for air quality classification tasks and highlight the utility of explainable AI in environmental health research.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Machine Learning-Based Classification with Shapley Additive Explanations: A Case Study on PM2.5 Concentration Levels

  • Jiraroj Tosasukul,
  • Ratchada Viriyapong

摘要

Abstract

Air pollutants particularly PM2.5 have been shown to largely affect environments and human health across the world including Thailand. Air pollutants prediction therefore plays an important role in air pollution alerts and in management to keep pollutant under control. This study presents a comparative analysis of six supervised machine learning models—artificial neural networks, random forest, support vector machines, naïve Bayes, \(k\) -nearest neighbors, and extreme gradient boosting—for classifying PM2.5 concentration levels in Chiang Mai, northern part of Thailand. The models were trained and evaluated using weekly environmental data collected over 524 weeks from 2015 to 2024. Among them, the random forest model achieved the highest accuracy (72.82 \(\%\) ). The ANN and NB models also demonstrated moderate to substantial agreement levels based on Cohen’s kappa statistics. Model interpretability was further enhanced using SHAP (Shapley additive explanations) analysis. The results indicate that the number of hotspot occurrences in Chiang Mai, Tak, and Mae Hong Son are the most influential predictors. These results underscore the value of ensemble tree-based methods for air quality classification tasks and highlight the utility of explainable AI in environmental health research.