Water quality is crucial for every segment of life and with the increasing importance of reliable and prompt water quality information for environmental monitoring and public health, advanced computational methods are vital. Bayesian approaches, known for their probabilistic foundation and capability to handle uncertainties, provide a promising approach for addressing this need. This study undertakes an in-depth examination of the performance of various Bayesian models in classifying water quality as safe or not safe, using a comprehensive dataset of chemical water parameters. The methodology involves a complete workflow that begins with data preparation through quantile-based discretization and dataset splitting. Next, four distinct Bayesian classifiers—Naive Bayes (NB), Tree Augmented Naive Bayes (TAN), Forest Augmented Naive Bayes (FAN), and a General Bayesian Network (BN)—were implemented and evaluated. Each model’s performance was rigorously assessed using a variety of metrics, including Accuracy, Precision, Recall, F1 score and particularly Matthews Correlation Coefficient (MCC) to consider the dataset’s imbalance. The results indicate that while NB provides a fundamental approach, its performance, along with TAN and FAN, may be compromised by overfitting. In contrast, the BN model, with its focus on significant features—cadmium and aluminium—exhibited consistent performance across training and testing datasets, indicating effective learning and generalization. Specifically, the BN model achieved an accuracy of 91.731%, a precision of 95.121%, a recall of 95.583%, an F1-score of 95.351%, and an MCC of 0.58. These results underscore the model’s balanced and robust performance, especially in the context of an unbalanced dataset. Additionally, the Conditional Probability Distributions (CPDs) extracted from the BN model explained the relationships between specific chemical concentrations and water safety. Taking all factors and results into account this study underscores the integration of Bayesian models into water quality assessment frameworks, particularly advocating for the BN model due to its balanced precision, recall, and reduced bias towards majority class.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating Bayesian Approaches for Water Quality Classification: A Comparative Study

  • Ana Dodig,
  • Tatjana Lutovac

摘要

Water quality is crucial for every segment of life and with the increasing importance of reliable and prompt water quality information for environmental monitoring and public health, advanced computational methods are vital. Bayesian approaches, known for their probabilistic foundation and capability to handle uncertainties, provide a promising approach for addressing this need. This study undertakes an in-depth examination of the performance of various Bayesian models in classifying water quality as safe or not safe, using a comprehensive dataset of chemical water parameters. The methodology involves a complete workflow that begins with data preparation through quantile-based discretization and dataset splitting. Next, four distinct Bayesian classifiers—Naive Bayes (NB), Tree Augmented Naive Bayes (TAN), Forest Augmented Naive Bayes (FAN), and a General Bayesian Network (BN)—were implemented and evaluated. Each model’s performance was rigorously assessed using a variety of metrics, including Accuracy, Precision, Recall, F1 score and particularly Matthews Correlation Coefficient (MCC) to consider the dataset’s imbalance. The results indicate that while NB provides a fundamental approach, its performance, along with TAN and FAN, may be compromised by overfitting. In contrast, the BN model, with its focus on significant features—cadmium and aluminium—exhibited consistent performance across training and testing datasets, indicating effective learning and generalization. Specifically, the BN model achieved an accuracy of 91.731%, a precision of 95.121%, a recall of 95.583%, an F1-score of 95.351%, and an MCC of 0.58. These results underscore the model’s balanced and robust performance, especially in the context of an unbalanced dataset. Additionally, the Conditional Probability Distributions (CPDs) extracted from the BN model explained the relationships between specific chemical concentrations and water safety. Taking all factors and results into account this study underscores the integration of Bayesian models into water quality assessment frameworks, particularly advocating for the BN model due to its balanced precision, recall, and reduced bias towards majority class.