Groundwater is a crucial resource, serving as a primary source of drinking water in various regions worldwide, such as India. Consequently, maintaining high groundwater quality is essential. Because of the inevitable advantages of data-driven models, Machine Learning (ML) techniques can contribute to predicting and analyzing water quality. This study employed four ML methods to assess water quality in Haryana state, India: Multiple Linear Regression (MLR), Support Vector Regression (SVR), Gaussian Process Regression (GPR), and Random Forest Regression (RFR). The Pollution Index of Groundwater (PIG) results indicated that 56.72% of groundwater samples exhibited insignificant pollution levels. Additionally, 14.92% of samples showed low pollution, 9.88% moderate pollution, 7.56% high pollution, and 10.92% very high pollution. Among the models tested, Random Forest Regression (RFR) demonstrated superior performance across all datasets compared to the other three models. These findings could assist decision-makers in formulating effective water management and quality initiatives in the future.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Prediction of Pollution Index of Groundwater Using Machine Learning Approach: A Case Study of Haryana State, Northern India

  • Hemant Raheja,
  • Arun Goel,
  • Mahesh Pal

摘要

Groundwater is a crucial resource, serving as a primary source of drinking water in various regions worldwide, such as India. Consequently, maintaining high groundwater quality is essential. Because of the inevitable advantages of data-driven models, Machine Learning (ML) techniques can contribute to predicting and analyzing water quality. This study employed four ML methods to assess water quality in Haryana state, India: Multiple Linear Regression (MLR), Support Vector Regression (SVR), Gaussian Process Regression (GPR), and Random Forest Regression (RFR). The Pollution Index of Groundwater (PIG) results indicated that 56.72% of groundwater samples exhibited insignificant pollution levels. Additionally, 14.92% of samples showed low pollution, 9.88% moderate pollution, 7.56% high pollution, and 10.92% very high pollution. Among the models tested, Random Forest Regression (RFR) demonstrated superior performance across all datasets compared to the other three models. These findings could assist decision-makers in formulating effective water management and quality initiatives in the future.