<p>Air pollution is driven by human and natural activities, with complex nonlinear interactions shaping its dynamics and sources. For air pollution source apportionment, Principal Component Analysis (PCA) is commonly used; however, it is limited to capturing only linear variance in a dataset. This study addresses this limitation using Kernel Principal Component Analysis (KPCA), an extension of PCA that captures nonlinear variance, to apportion pollution sources across air quality monitoring stations in Malaysia. As a pre-processing step, Ball Mapper, a topological data analysis technique, categorizes stations based on the Air Pollution Index (API), which reflects air quality status. Multiple linear regression (MLR) is used to quantify the contribution of individual pollutants to API, identifying Particulate matter (PM<sub>2.5</sub>) as the dominant contributor. Finally, the PM<sub>2.5</sub> predictions using MLR are compared with those from random forest regression (RFR), long short-term memory (LSTM), and artificial neural networks (ANN). LSTM and ANN outperform the other models, underscoring the importance of nonlinear methods for understanding pollutant dynamics and offering a robust framework for effective environmental management.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Topological Approach and Kernel Principal Component Analysis for Air Pollution Source Apportionment

  • Vine Nwabuisi Madukpe,
  • Bright Chukwuma Ugoala,
  • Nur Fariha Syaqina Zulkepli

摘要

Air pollution is driven by human and natural activities, with complex nonlinear interactions shaping its dynamics and sources. For air pollution source apportionment, Principal Component Analysis (PCA) is commonly used; however, it is limited to capturing only linear variance in a dataset. This study addresses this limitation using Kernel Principal Component Analysis (KPCA), an extension of PCA that captures nonlinear variance, to apportion pollution sources across air quality monitoring stations in Malaysia. As a pre-processing step, Ball Mapper, a topological data analysis technique, categorizes stations based on the Air Pollution Index (API), which reflects air quality status. Multiple linear regression (MLR) is used to quantify the contribution of individual pollutants to API, identifying Particulate matter (PM2.5) as the dominant contributor. Finally, the PM2.5 predictions using MLR are compared with those from random forest regression (RFR), long short-term memory (LSTM), and artificial neural networks (ANN). LSTM and ANN outperform the other models, underscoring the importance of nonlinear methods for understanding pollutant dynamics and offering a robust framework for effective environmental management.