Everywhere, there is an increase in chronic diseases. These diseases affect the immune system and quality of life. Chronic illness treatment can be expensive and difficult to manage. People should focus on healthy lifestyle habits like exercising regularly, eating a well-balanced diet, and getting enough sleep because prevention is the best option. These illnesses can make it hard for doctors to treat. Dimensionality reduction methods allow for the processing of large data sets, allowing for decision support in the area of chronic diseases. A large number of symptoms are included in the records that are important to predict disease. By using feature selection methods, there is a need to remove unnecessary and meaningless symptoms from data sets so as to deliver better classification accuracy. By using feature selection methods, prediction accuracy and model complexity can be reduced. Furthermore, it could improve the interpretation of models and make them easier to understand and provide more insight into the data that lies beneath them. In addition, the option of selecting features may speed up model creation and reduce training time. Moreover, it may reduce the likelihood of overfitting which can have an adverse effect on model performance. The primary contribution of this paper is a comparative analysis of the effects of the analyzed filtering methods, which include the Wrapper methods, which include the Best First Search (BFS) method, the Linear Forward Selection (LFS) method, and the Greedy Step Wise Search (GSS) method, as well as the Information Gain (IG) method, Chi-Square (CS) method, and the Cor-relation Feature Selection (CFS) method. The results show that the Wrapper approach, based on the Best First Search method, has been shown to be more accurate and effective in selecting features. Furthermore, it has been noted that the data gain and correlation element choice strategies have become more persuasive in terms of preciseness. A Random Forest algorithm was used in Python Anaconda as a classifier for this analysis. Attribute significance analysis was performed on the breast cancer, diabetes, kidney disease, and heart disease study datasets. In terms of accuracy and time taken to complete, the CFS approach was better than other filtering techniques. The CFS method, as well as the datasets for heart disease, diabetes, kidney disease, and breast cancer, all had an ac-curacy rate of 96.5 percent. Likewise, similar idle periods were recorded for each of the datasets: 1.07 s, 1.04 s, 1.06 s and 1.01 s. As demonstrated by the accuracy rate and the time taken to complete the analysis, the CFS method was a reliable and efficient method of medical analysis. Furthermore, the CFS method has a substantially higher accuracy rate compared to other filter methods. That indicates that the CFS approach can be considered as a viable option for diagnosis. BFS performed exceptionally well in terms of covering strategies compared to other methods. The breast cancer, diabetes, and heart disease datasets achieved maximum accuracy of 94.7%, 89.9%, 92.3%, and 92.8%, respectively. For specific datasets, the same method has been used to detect a latency delay of 1.42 s, 1.44 s, 1.32 s and 1.37 s. An improved random forest classifier has also been used to evaluate the hybrid approach. The classification and clustering are combined by an improved random forest classifier. Its performance has been recorded and verified in 12 separate chronic disease datasets. The results of this classification were very consistent and effective. A mean of 96.7%, 96.5%, 95.6% and 96.2% was achieved for accuracy, specificity, sensitivity and F1-score metrics, respectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Enhanced Learning Model Based on an Improved Random Forest Classifier and an Integrated Attribute Selector for Healthcare Datasets

  • S. Rajeashwari,
  • K. Arunesh

摘要

Everywhere, there is an increase in chronic diseases. These diseases affect the immune system and quality of life. Chronic illness treatment can be expensive and difficult to manage. People should focus on healthy lifestyle habits like exercising regularly, eating a well-balanced diet, and getting enough sleep because prevention is the best option. These illnesses can make it hard for doctors to treat. Dimensionality reduction methods allow for the processing of large data sets, allowing for decision support in the area of chronic diseases. A large number of symptoms are included in the records that are important to predict disease. By using feature selection methods, there is a need to remove unnecessary and meaningless symptoms from data sets so as to deliver better classification accuracy. By using feature selection methods, prediction accuracy and model complexity can be reduced. Furthermore, it could improve the interpretation of models and make them easier to understand and provide more insight into the data that lies beneath them. In addition, the option of selecting features may speed up model creation and reduce training time. Moreover, it may reduce the likelihood of overfitting which can have an adverse effect on model performance. The primary contribution of this paper is a comparative analysis of the effects of the analyzed filtering methods, which include the Wrapper methods, which include the Best First Search (BFS) method, the Linear Forward Selection (LFS) method, and the Greedy Step Wise Search (GSS) method, as well as the Information Gain (IG) method, Chi-Square (CS) method, and the Cor-relation Feature Selection (CFS) method. The results show that the Wrapper approach, based on the Best First Search method, has been shown to be more accurate and effective in selecting features. Furthermore, it has been noted that the data gain and correlation element choice strategies have become more persuasive in terms of preciseness. A Random Forest algorithm was used in Python Anaconda as a classifier for this analysis. Attribute significance analysis was performed on the breast cancer, diabetes, kidney disease, and heart disease study datasets. In terms of accuracy and time taken to complete, the CFS approach was better than other filtering techniques. The CFS method, as well as the datasets for heart disease, diabetes, kidney disease, and breast cancer, all had an ac-curacy rate of 96.5 percent. Likewise, similar idle periods were recorded for each of the datasets: 1.07 s, 1.04 s, 1.06 s and 1.01 s. As demonstrated by the accuracy rate and the time taken to complete the analysis, the CFS method was a reliable and efficient method of medical analysis. Furthermore, the CFS method has a substantially higher accuracy rate compared to other filter methods. That indicates that the CFS approach can be considered as a viable option for diagnosis. BFS performed exceptionally well in terms of covering strategies compared to other methods. The breast cancer, diabetes, and heart disease datasets achieved maximum accuracy of 94.7%, 89.9%, 92.3%, and 92.8%, respectively. For specific datasets, the same method has been used to detect a latency delay of 1.42 s, 1.44 s, 1.32 s and 1.37 s. An improved random forest classifier has also been used to evaluate the hybrid approach. The classification and clustering are combined by an improved random forest classifier. Its performance has been recorded and verified in 12 separate chronic disease datasets. The results of this classification were very consistent and effective. A mean of 96.7%, 96.5%, 95.6% and 96.2% was achieved for accuracy, specificity, sensitivity and F1-score metrics, respectively.