<p>Access to clean and safe water has been a crucial issue of global concern in line with Sustainable Development Goal (SDG) 6. Nevertheless, traditional Water Quality Index (WQI) models usually possess problems of class bias, inadequate flexibility, and interpretability, which decrease their efficiency in practice. To overcome these constraints, this research suggests an Explainability-Based Multi-Stage Framework to effectively and meaningfully assess river water quality by both classification and regression methods. The framework integrates multiple advanced techniques into a unified pipeline. First, five resampling algorithms, SMOTE, Borderline-SMOTE, SMOTE-ENN, SMOTE-Tomek Links, and ADASYN, are carefully compared to address the issue of class imbalance. This is then followed by a model selection process based on reinforcement learning (RL) that considers Q-Learning, Deep Q-Network (DQN), Proximal Policy Optimization (PPO) and Multi-Armed Bandit (MAB) to select the most effective predictive model. To promote transparency and trust, LIME, SHAP, and Multi-Head Attention (MHA) are integrated throughout pre-ad hoc, ad hoc and post-ad hoc steps, and a systematic explainability-vast feature selection is achieved. The framework is evaluated using 1096 river water samples collected from the Kaveri River basin during the period 2008–2025, with each sample comprising nineteen water-quality parameters. The experimental findings have shown that SMOTE has the highest performance among resampling methods with a Macro F1-score of 0.90, which successfully combats the issue of class imbalance. The RL models show that DQN has a higher validation accuracy of 0.932, test accuracy of 0.941, precision of 0.948, recall of 0.941, and F1-score of 0.942. Additionally, the explainability-guided feature selection establishes ten important features, which considerably improve the performance of regression in regards to Biochemical Oxygen Demand (BOD), with R<sup>2</sup> of 0.9839 (training), 0.9801 (validation), and 0.9725 (testing). Altogether, the suggested Explainability-Driven Multi-Stage Framework provides a powerful, interpretable, and efficient model to assess the water quality in rivers, promoting efficient environmental monitoring and developing sustainable water management.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Explainability-driven multi-stage framework for river water quality assessment using reinforcement learning

  • S. Ramya,
  • S. Srinath,
  • Pushpa Tuppad,
  • V. Chandan

摘要

Access to clean and safe water has been a crucial issue of global concern in line with Sustainable Development Goal (SDG) 6. Nevertheless, traditional Water Quality Index (WQI) models usually possess problems of class bias, inadequate flexibility, and interpretability, which decrease their efficiency in practice. To overcome these constraints, this research suggests an Explainability-Based Multi-Stage Framework to effectively and meaningfully assess river water quality by both classification and regression methods. The framework integrates multiple advanced techniques into a unified pipeline. First, five resampling algorithms, SMOTE, Borderline-SMOTE, SMOTE-ENN, SMOTE-Tomek Links, and ADASYN, are carefully compared to address the issue of class imbalance. This is then followed by a model selection process based on reinforcement learning (RL) that considers Q-Learning, Deep Q-Network (DQN), Proximal Policy Optimization (PPO) and Multi-Armed Bandit (MAB) to select the most effective predictive model. To promote transparency and trust, LIME, SHAP, and Multi-Head Attention (MHA) are integrated throughout pre-ad hoc, ad hoc and post-ad hoc steps, and a systematic explainability-vast feature selection is achieved. The framework is evaluated using 1096 river water samples collected from the Kaveri River basin during the period 2008–2025, with each sample comprising nineteen water-quality parameters. The experimental findings have shown that SMOTE has the highest performance among resampling methods with a Macro F1-score of 0.90, which successfully combats the issue of class imbalance. The RL models show that DQN has a higher validation accuracy of 0.932, test accuracy of 0.941, precision of 0.948, recall of 0.941, and F1-score of 0.942. Additionally, the explainability-guided feature selection establishes ten important features, which considerably improve the performance of regression in regards to Biochemical Oxygen Demand (BOD), with R2 of 0.9839 (training), 0.9801 (validation), and 0.9725 (testing). Altogether, the suggested Explainability-Driven Multi-Stage Framework provides a powerful, interpretable, and efficient model to assess the water quality in rivers, promoting efficient environmental monitoring and developing sustainable water management.