Topological Approach and Kernel Principal Component Analysis for Air Pollution Source Apportionment
摘要
Air pollution is driven by human and natural activities, with complex nonlinear interactions shaping its dynamics and sources. For air pollution source apportionment, Principal Component Analysis (PCA) is commonly used; however, it is limited to capturing only linear variance in a dataset. This study addresses this limitation using Kernel Principal Component Analysis (KPCA), an extension of PCA that captures nonlinear variance, to apportion pollution sources across air quality monitoring stations in Malaysia. As a pre-processing step, Ball Mapper, a topological data analysis technique, categorizes stations based on the Air Pollution Index (API), which reflects air quality status. Multiple linear regression (MLR) is used to quantify the contribution of individual pollutants to API, identifying Particulate matter (PM2.5) as the dominant contributor. Finally, the PM2.5 predictions using MLR are compared with those from random forest regression (RFR), long short-term memory (LSTM), and artificial neural networks (ANN). LSTM and ANN outperform the other models, underscoring the importance of nonlinear methods for understanding pollutant dynamics and offering a robust framework for effective environmental management.