Background <p>Low-cost sensors (LCS) are widely used for air quality monitoring, but their accuracy depends on proper calibration. This study compares linear regression (LR) and machine learning (ML) techniques, particularly random forest (RF), to determine optimal calibration strategies.</p> Objectives <p>This study aims to compare the effectiveness of LR and RF models in calibrating the Plantower PMS 3003 sensor under different environmental conditions. It also explores ways to streamline calibration efforts whilemaintaining accuracy.</p> Methods <p>Sensor data were collected in a controlled laboratory setting, with measurements compared against a reference monitor. LR and RF models were developed to calibrate the sensor, and their performance was evaluated based on RMSE, R<sup>2</sup>, and bias. Additionally, the study examined whether using fewer sensors for training could still produce reliable calibration models.</p> Results <p>Both LR and RF models demonstrated strong calibration performance. LR models were effective for low to moderate PM2.5 concentrations and required fewer computational resources, making them suitable for large-scale monitoring with limited resources. RF models captured nonlinear relationships, showing superior accuracy at high PMconcentrations and in conditions with high relative humidity. The findings suggest that LR models trained on smallerdatasets can achieve practical accuracy, reducing the need for extensive individual sensor calibration.</p> Conclusions <p>The selection of a calibration model should be guided by study-specific requirements, including environmental conditions and resource availability. LR models are recommended for large-scale studies with constrained resources, while RF models may offer advantages in high-exposure environments due to their ability to model complex interactions. This study is the first to explore reducing sensor calibration efforts while maintaining accuracy, highlighting the potential for optimized strategies in resource-limited settings. Future research should validate these findings in real-world deployments to further refine calibration models for LCS applications.</p> Graphical Abstract <p></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimizing Air Quality Monitoring: Comparative Analysis of Linear Regression and Machine Learning in Low-Cost Sensor Calibration

  • Runcheng Fang,
  • Scott Collingwood,
  • Yue Zhang,
  • Joseph B. Stanford,
  • Christina Porucznik,
  • Darrah Sleeth

摘要

Background

Low-cost sensors (LCS) are widely used for air quality monitoring, but their accuracy depends on proper calibration. This study compares linear regression (LR) and machine learning (ML) techniques, particularly random forest (RF), to determine optimal calibration strategies.

Objectives

This study aims to compare the effectiveness of LR and RF models in calibrating the Plantower PMS 3003 sensor under different environmental conditions. It also explores ways to streamline calibration efforts whilemaintaining accuracy.

Methods

Sensor data were collected in a controlled laboratory setting, with measurements compared against a reference monitor. LR and RF models were developed to calibrate the sensor, and their performance was evaluated based on RMSE, R2, and bias. Additionally, the study examined whether using fewer sensors for training could still produce reliable calibration models.

Results

Both LR and RF models demonstrated strong calibration performance. LR models were effective for low to moderate PM2.5 concentrations and required fewer computational resources, making them suitable for large-scale monitoring with limited resources. RF models captured nonlinear relationships, showing superior accuracy at high PMconcentrations and in conditions with high relative humidity. The findings suggest that LR models trained on smallerdatasets can achieve practical accuracy, reducing the need for extensive individual sensor calibration.

Conclusions

The selection of a calibration model should be guided by study-specific requirements, including environmental conditions and resource availability. LR models are recommended for large-scale studies with constrained resources, while RF models may offer advantages in high-exposure environments due to their ability to model complex interactions. This study is the first to explore reducing sensor calibration efforts while maintaining accuracy, highlighting the potential for optimized strategies in resource-limited settings. Future research should validate these findings in real-world deployments to further refine calibration models for LCS applications.

Graphical Abstract