Background <p>Ensuring sustainable water management (SDG 6) requires reliable IoT-based sensor networks, yet low-cost deployments are often hindered by environmental noise and the opaque nature of complex AI models. This research aims to develop a robust contaminant classification framework that prioritizes actionable transparency and physical consistency in environmental monitoring.</p> Methods <p>Time-series water quality data were analyzed using Random Forest (RF), LSTM, and Transformer models. Beyond standard performance metrics (Weighted F1-Score), this study introduces a Dual-Layer Validation (DLV) framework powered by SHAP-based Explainable AI (XAI) to verify the models’ alignment with hydrogeochemical principles.</p> Results <p>The RF model achieved the highest performance (F1 = 0.872), outperforming LSTM (0.852) and Transformer (0.817) (<i>p</i> &lt; 0.05). Crucially, the XAI analysis revealed a strong cross-architecture consensus: both ensemble and attention-based models independently identified Turbidity and pH as the principal eco-chemical drivers of contamination.</p> Conclusion <p>The DLV framework demonstrates that transparent models can effectively learn genuine physical and hydrogeochemical fingerprints rather than statistical artifacts. This cross-validated resilience to sensor noise offers a critical pathway toward sustainable, trustworthy, and cost-effective IoT water quality management systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Explainable AI with dual-layer validation framework for ensuring noise resilience and physical consistency in low-cost sustainable water quality monitoring

  • Wibowo Harry Sugiharto,
  • Rizkysari Meimaharani,
  • Ahmad Abdul Chamid,
  • Muhammad Imam Ghozali,
  • Alif Catur Murti

摘要

Background

Ensuring sustainable water management (SDG 6) requires reliable IoT-based sensor networks, yet low-cost deployments are often hindered by environmental noise and the opaque nature of complex AI models. This research aims to develop a robust contaminant classification framework that prioritizes actionable transparency and physical consistency in environmental monitoring.

Methods

Time-series water quality data were analyzed using Random Forest (RF), LSTM, and Transformer models. Beyond standard performance metrics (Weighted F1-Score), this study introduces a Dual-Layer Validation (DLV) framework powered by SHAP-based Explainable AI (XAI) to verify the models’ alignment with hydrogeochemical principles.

Results

The RF model achieved the highest performance (F1 = 0.872), outperforming LSTM (0.852) and Transformer (0.817) (p < 0.05). Crucially, the XAI analysis revealed a strong cross-architecture consensus: both ensemble and attention-based models independently identified Turbidity and pH as the principal eco-chemical drivers of contamination.

Conclusion

The DLV framework demonstrates that transparent models can effectively learn genuine physical and hydrogeochemical fingerprints rather than statistical artifacts. This cross-validated resilience to sensor noise offers a critical pathway toward sustainable, trustworthy, and cost-effective IoT water quality management systems.