Background <p>Accurate personal exposure assessment to air pollution is crucial in environmental epidemiological studies, but it remains challenging due to substantial differences in pollutant concentrations between indoor and outdoor environments. Indirect assessment methods estimate personal exposure by incorporating microenvironment-specific pollutant concentrations and personal time-activity patterns, while requiring participants to record detailed time-location diaries.</p> Objective <p>This study aims to develop microenvironment classification models using global positioning system (GPS) tracking data.</p> Methods <p>We used data from the Korean Air pollutant EXposure (KAPEX) model project, which reflects seasonal and daily activity patterns of the Korean urban population. Three microenvironment categorizations were defined: Two-level (indoor and outdoor), Three-level (indoor, walk, and transit), and Four-level (home and workplace, other indoor, walk, and transit), where the outdoor category is subdivided into walk and transit in Three- and Four-level. Classification models were developed for each categorization using various machine learning (ML) and deep learning (DL) methods, incorporating both individual mobility patterns and GPS signal quality information. Important variables contributing to classification accuracy were identified using interpretable ML or DL methods.</p> Results <p>Random forest achieved the highest test AUROC of 0.963 and 0.958 for the Two-level and Three-level categorization, respectively, while boosting achieved the highest test AUROC of 0.918 for the Four-level categorization. Variables related to both mobility patterns and GPS signal quality were found to enhance classification accuracy.</p> Significance <p>These findings highlight that incorporating both mobility patterns and signal quality into ML models significantly improves classification accuracy.</p> Impact statement <p><UnorderedList Mark="Bullet"> <ItemContent> <p>Microenvironment classification models were developed using GPS tracking data to improve personal exposure assessment. Both mobility patterns and GPS signal quality were used to enhance model accuracy. Random forest and boosting outperformed the deep neural network across various outcome categorizations. SHAP values were used to identify important variables contributing to model performance.</p> </ItemContent> </UnorderedList></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Developing microenvironment classification models for personal exposure assessment based on global positioning system tracking data

  • Jiwoong Yu,
  • Hyunwoo Jeon,
  • Kiyoung Lee,
  • Woojoo Lee

摘要

Background

Accurate personal exposure assessment to air pollution is crucial in environmental epidemiological studies, but it remains challenging due to substantial differences in pollutant concentrations between indoor and outdoor environments. Indirect assessment methods estimate personal exposure by incorporating microenvironment-specific pollutant concentrations and personal time-activity patterns, while requiring participants to record detailed time-location diaries.

Objective

This study aims to develop microenvironment classification models using global positioning system (GPS) tracking data.

Methods

We used data from the Korean Air pollutant EXposure (KAPEX) model project, which reflects seasonal and daily activity patterns of the Korean urban population. Three microenvironment categorizations were defined: Two-level (indoor and outdoor), Three-level (indoor, walk, and transit), and Four-level (home and workplace, other indoor, walk, and transit), where the outdoor category is subdivided into walk and transit in Three- and Four-level. Classification models were developed for each categorization using various machine learning (ML) and deep learning (DL) methods, incorporating both individual mobility patterns and GPS signal quality information. Important variables contributing to classification accuracy were identified using interpretable ML or DL methods.

Results

Random forest achieved the highest test AUROC of 0.963 and 0.958 for the Two-level and Three-level categorization, respectively, while boosting achieved the highest test AUROC of 0.918 for the Four-level categorization. Variables related to both mobility patterns and GPS signal quality were found to enhance classification accuracy.

Significance

These findings highlight that incorporating both mobility patterns and signal quality into ML models significantly improves classification accuracy.

Impact statement

Microenvironment classification models were developed using GPS tracking data to improve personal exposure assessment. Both mobility patterns and GPS signal quality were used to enhance model accuracy. Random forest and boosting outperformed the deep neural network across various outcome categorizations. SHAP values were used to identify important variables contributing to model performance.