Improving the Efficiency of Water Quality Prediction Using the SuperTML Approach in Machine Learning
摘要
The quality of water is important to the growth of a sustainable environment. Contaminated water causes serious waterborne diseases to human life. Predicting the potability of water is more practical for water management, pollution control, and cutting off the threat of water-related diseases. Machine learning (Al-Adhaileh and Alsaade in Journal on Sustainability 13(8), 2021, [1])–(Zhang et al. in Expo Health 12:487–500, 2020, [4]) has widely developed its scope in analyzing, classifying, and predicting data models. Potability is a measure used to identify the quality of water to be consumable. Potability is taken categorically to perform the classification of water samples based on their characteristics. The features used for this analysis are pH, solids, hardness, chloramines, sulphates, organic carbons, conductivity, turbidity, and trihalomethanes. This paper's water quality prediction used SuperTML (tabular machine learning) and CNN models to predict water potability efficiently. Data are collected and pre-processing is done using SMOTE analysis. It involves projection of features into two-dimensional embeddings that resemble images for every input of processed data, and the resultant image is then given into CNN models for classification. The Super Characters approach has recently produced two-dimensional word embeddings that have advanced in text categorization challenges. The resultant images from SuperTML are projected into the convolutional neural network model. The CNN model is typically more accurate and produces reliable results. The proposed model using SuperTML and CNN produces training results of 98% and testing accuracy above 90% which is better than the existing models.