Explainable AI with dual-layer validation framework for ensuring noise resilience and physical consistency in low-cost sustainable water quality monitoring
摘要
Ensuring sustainable water management (SDG 6) requires reliable IoT-based sensor networks, yet low-cost deployments are often hindered by environmental noise and the opaque nature of complex AI models. This research aims to develop a robust contaminant classification framework that prioritizes actionable transparency and physical consistency in environmental monitoring.
MethodsTime-series water quality data were analyzed using Random Forest (RF), LSTM, and Transformer models. Beyond standard performance metrics (Weighted F1-Score), this study introduces a Dual-Layer Validation (DLV) framework powered by SHAP-based Explainable AI (XAI) to verify the models’ alignment with hydrogeochemical principles.
ResultsThe RF model achieved the highest performance (F1 = 0.872), outperforming LSTM (0.852) and Transformer (0.817) (p < 0.05). Crucially, the XAI analysis revealed a strong cross-architecture consensus: both ensemble and attention-based models independently identified Turbidity and pH as the principal eco-chemical drivers of contamination.
ConclusionThe DLV framework demonstrates that transparent models can effectively learn genuine physical and hydrogeochemical fingerprints rather than statistical artifacts. This cross-validated resilience to sensor noise offers a critical pathway toward sustainable, trustworthy, and cost-effective IoT water quality management systems.