Abstract
This study aimed to determine predictors of particulate matter (PM \({}_{2.5}\) ) concentrations at different heights (30 m, 75 m, and 110 m) at a monitoring tower in Bangkok, Thailand, a megacity representative of various Asian megacities suffering from air pollution problems. Ground level PM \({}_{2.5}\) air pollution data were used as controls. In a first step, feature selection methods were applied to narrow down the number of relevant predictors. Subsequently, a machine learning algorithm, namely, a deep long short-term memory model, was applied to identify the most relevant predictors. With this approach, we identified the following key PM \({}_{2.5}\) predictors: carbon monoxide for low levels (ground level and 30 m), ozone for medium and high (75 and 110 m) levels, and pressure for high levels (110 m). We also found that using all available features as predictors led to higher prediction errors and, consequently, in general, seems to be a modeling strategy that should not be pursued. Moreover, varying the input data length of the model was a key step to improve prediction accuracy. PM \({}_{2.5}\) signals observed at different vertical levels exhibited height-specific optimal input lengths ranging from 1 to 7 days. Finally, predicting low level (30 m) PM \({}_{2.5}\) concentrations as observed in our vertical PM \({}_{2.5}\) data set and predicting PM \({}_{2.5}\) concentrations from one of the ground level control monitoring sites located nearby a heavy-traffic road could be achieved best with two models that both used carbon monoxide as predictor and a data input of 5 days.