Background <p>Communities along Ghana’s Pra, Ankobra, and Tano rivers depend heavily on surface water for drinking, irrigation, and fisheries but face increasing arsenic (As) exposure risks due to intensified small-scale (galamsey) mining. Uneven monitoring density complicates hotspot detection and burden estimation. The present study aimed to (i) identify a high-performing, interpretable machine learning model for forecasting arsenic (As) concentrations and exceedance of the WHO 10 µg/L guideline and (ii) produce bias-aware spatial risk evidence that distinguishes true hotspots from artifacts of unequal monitoring density.</p> Methods <p>We audited data quality (0% missingness; predictor-only winsorization of longitude at 2.5/97·5th to limit leverage; raw As preserved). Spatial burden was quantified with (i) an effort-adjusted exposure index (site As × inverse local station density), (ii) a density-weighted, row-standardized Getis–Ord Gi* (25 km), and (iii) an effort-offset relative-risk KDE (exceedance intensity ÷ effort). For prediction, random forest (RF), SVM-RBF, and a neural network were compared via 5× repeated 5-fold nested CV and multiple train–validation–test splits; the selected RF (70–20–10) was interpreted with SHAP. STL decomposition characterized the 2020–2026 temporal structure; OOB curves and residuals were used to assess model adequacy.</p> Results <p>The overall exceedance of the WHO 10 µg/L guideline was high (63%). After effort adjustment, Pra had the largest exposure index (median 321.45 [47.50]), whereas Ankobra presented the highest exceedance share (76.5% [8.8%]). Tano was lower in both. Bias-aware Gi* yielded sparse, localized hotspots, with NA flags where the density was too low for stable z scores. The RFs consistently matched or outperformed the alternatives; the OOB error plateaued at ~250–350 trees, and the test residuals showed no pathological patterns. SHAP ranked temperature, pH, and latitude as the dominant drivers. STL indicated weak seasonality and a slight downward trend.</p> Conclusion <p>Arsenic risk is substantial and spatially heterogeneous. After effort imbalance is corrected, Pra should be prioritized for mitigation (largest aggregate burden), whereas Ankobra warrants intensified compliance surveillance (highest exceedance probability). The selected interpretable RF model provides robust, policy-ready forecasts to guide targeted enforcement and resilient water-service investments. Together, the interpretable ML predictions and effort-adjusted hotspot maps provide actionable, uncertainty-aware evidence to prioritize mitigation in Pra and high-frequency compliance surveillance in Ankobra while acknowledging data gaps and model limits that future monitoring (broader covariates, more balanced station coverage) should address.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative performance of machine learning models in predicting arsenic exposure risk in Ghana’s riverine ecosystems

  • Senyefia Bosson-Amedenu,
  • Ebenezer Boakye,
  • Emmanuel Ayitey,
  • John Awuah Addor

摘要

Background

Communities along Ghana’s Pra, Ankobra, and Tano rivers depend heavily on surface water for drinking, irrigation, and fisheries but face increasing arsenic (As) exposure risks due to intensified small-scale (galamsey) mining. Uneven monitoring density complicates hotspot detection and burden estimation. The present study aimed to (i) identify a high-performing, interpretable machine learning model for forecasting arsenic (As) concentrations and exceedance of the WHO 10 µg/L guideline and (ii) produce bias-aware spatial risk evidence that distinguishes true hotspots from artifacts of unequal monitoring density.

Methods

We audited data quality (0% missingness; predictor-only winsorization of longitude at 2.5/97·5th to limit leverage; raw As preserved). Spatial burden was quantified with (i) an effort-adjusted exposure index (site As × inverse local station density), (ii) a density-weighted, row-standardized Getis–Ord Gi* (25 km), and (iii) an effort-offset relative-risk KDE (exceedance intensity ÷ effort). For prediction, random forest (RF), SVM-RBF, and a neural network were compared via 5× repeated 5-fold nested CV and multiple train–validation–test splits; the selected RF (70–20–10) was interpreted with SHAP. STL decomposition characterized the 2020–2026 temporal structure; OOB curves and residuals were used to assess model adequacy.

Results

The overall exceedance of the WHO 10 µg/L guideline was high (63%). After effort adjustment, Pra had the largest exposure index (median 321.45 [47.50]), whereas Ankobra presented the highest exceedance share (76.5% [8.8%]). Tano was lower in both. Bias-aware Gi* yielded sparse, localized hotspots, with NA flags where the density was too low for stable z scores. The RFs consistently matched or outperformed the alternatives; the OOB error plateaued at ~250–350 trees, and the test residuals showed no pathological patterns. SHAP ranked temperature, pH, and latitude as the dominant drivers. STL indicated weak seasonality and a slight downward trend.

Conclusion

Arsenic risk is substantial and spatially heterogeneous. After effort imbalance is corrected, Pra should be prioritized for mitigation (largest aggregate burden), whereas Ankobra warrants intensified compliance surveillance (highest exceedance probability). The selected interpretable RF model provides robust, policy-ready forecasts to guide targeted enforcement and resilient water-service investments. Together, the interpretable ML predictions and effort-adjusted hotspot maps provide actionable, uncertainty-aware evidence to prioritize mitigation in Pra and high-frequency compliance surveillance in Ankobra while acknowledging data gaps and model limits that future monitoring (broader covariates, more balanced station coverage) should address.