Predictive Modeling of Genetic Diseases: A Comparison of Various Machine Learning Approaches
摘要
The aim of this research is to improve the prediction of genetic disorder types by employing large patient data and comparing several machine learning algorithms like XGBoost, Random Forest, AdaBoost, and k-Nearest Neighbors. The study's goal is to enhance our understanding of the underlying patterns and factors contributing to various genetic disorders by analyzing a large dataset. To improve prediction accuracy, the study uses feature selection and dimensionality reduction techniques like Recursive Feature Elimination, Pearson Correlation, and Kendall feature selection. Principal Component Analysis (PCA) and t-distributed Stochastic Neighbor Embedding (t-SNE) simplify the data without losing vital details. The primary goal is to develop a robust and interpretable predictive model for genetic disorders. When the XGBoost Classifier was used with Recursive Feature Elimination, the highest accuracy of 82% was achieved. This research aims to determine the best methods for early detection and personalized treatment planning in clinical settings by thoroughly analyzing several machine learning approaches. Enhancing our capacity to anticipate genetic condition types provides healthcare professionals with useful insights, potentially leading to more tailored and successful treatment options for patients. This approach combines modern machine learning techniques with domain-specific expertise to address the complex issue of predicting the type of genetic disorders.