Development of Machine Learning Based 2-Step Regression Model for Mass Closure of PM2.5
摘要
Mass closure is a technique to establish a balance between measured PM2.5 concentrations and their corresponding chemical species, which include elements, cation, anion, and carbonaceous species. It highlights the role of chemical composition in PM2.5, providing new insights into the emission source. This study intends to develop a 2-step regression model to automate the mass closure computation. Random forest (RF) and multiple linear regression (MLR) algorithms work admirably in delivering a reconstructed mass (RM) with composite species concentrations that are similar to the actual one. Initially, mass closure of PM2.5 chemically characterized data of two sites (Mine top and CEPI) at Singrauli region was done manually using empirical formulae in the published literature. The obtained RM lies within the 80–120 range of percentage change, and their regression coefficient value should be physically reasonable. Subsequently, it is used to train with 1-step (RF) regression model to produce multiple output (PM components); further these outcomes become input for 2-step (MLR) regression model to produce RM. The model performance is evaluated by accuracy (fivefold cross validation), root mean square error (RMSE), and R2 value. RF algorithm predicted composite species are found highly correlated (R2 > 0.83) and less accurate (<62%), whereas MLR predicted RM with high correlation (R2 > 0.99) and accuracy (>98%). Furthermore, the predicted composite species holds strong correlation (R2 > 0.85) with actual ones. Correspondingly, the predicted RM also represents strong correlation (R2 > 0.98) with measured PM2.5 concentration. Less accuracy in the 1-step regression model needs to be improved by enhancing data size, data preprocessing, or by the application of other advanced ensemble approaches. This composite model helps in automation of mass closure computation in a real-time manner.