Forecasting the availability of clean water is becoming increasingly difficult. This study aims to identify the numerous contaminants present in water using various criteria and then use a machine learning classification algorithm to determine whether the water is clean. Each of the 1043 samples contains a unique set of 28 columns of data, including biochemical oxygen demand, temperature, PH, conductivity, dissolved oxygen, feces-associated coli form, total coli form, and nitrate. This study used Adaboost, RF, KNN, logistic regression, DT, SVM, and other machine learning techniques to construct a robust indicator and classification system for water quality. More accurate prediction models can be developed by examining more reliable data sources, highlighting the importance of ensuring that drinking water is safe. The correlation between predicted and actual values indicates the model or system’s accuracy when predicting water quality. Based on the obtained accuracy, the model predicts which method (RF, DT, KNN, or LR) would be the most effective at purifying a specific range of water and, consequently, The KNN algorithm is most effective at detecting whether water is safe to consume (94%). In contrast, the SVM algorithm is the least effective (66%).

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Early Detection and Prediction of Water Quality Using Machine Learning Techniques

  • N. Pooja kumara,
  • H. S. Vandana,
  • Bobby Lukose

摘要

Forecasting the availability of clean water is becoming increasingly difficult. This study aims to identify the numerous contaminants present in water using various criteria and then use a machine learning classification algorithm to determine whether the water is clean. Each of the 1043 samples contains a unique set of 28 columns of data, including biochemical oxygen demand, temperature, PH, conductivity, dissolved oxygen, feces-associated coli form, total coli form, and nitrate. This study used Adaboost, RF, KNN, logistic regression, DT, SVM, and other machine learning techniques to construct a robust indicator and classification system for water quality. More accurate prediction models can be developed by examining more reliable data sources, highlighting the importance of ensuring that drinking water is safe. The correlation between predicted and actual values indicates the model or system’s accuracy when predicting water quality. Based on the obtained accuracy, the model predicts which method (RF, DT, KNN, or LR) would be the most effective at purifying a specific range of water and, consequently, The KNN algorithm is most effective at detecting whether water is safe to consume (94%). In contrast, the SVM algorithm is the least effective (66%).