Assessment of AI models in healthcare: performance metrics and applications across critical diseases
摘要
By identifying critical disease patterns and severity, Artificial Intelligence (AI) has the potential to enable timely diagnosis and personalized treatment planning. Effective classification of diseases reduces diagnostic errors, expedites treatments, and improves patient outcomes. Medical data is complex, characterized by high noise, strong feature correlations, and high dimensionality, which presents challenges for traditional classification techniques. Artificial intelligence (AI), particularly machine learning (ML) and deep learning (DL), can be explored to address these challenges by analyzing diverse datasets of several diseases. ML algorithms like support vector machines (SVMs), logistic regression, and random forests excel with structured datasets. Meanwhile, DL models, including convolutional neural networks (CNNs), and Generative Adversarial Networks (GANs), handle high-dimensional and unstructured data, including medical images and text. GANs are particularly valuable for augmenting datasets by generating synthetic medical images, which enhances model training and performance, especially for rare diseases. The proposed research work provides a scientific and objective assessment of AI models for performance metrics and applications in critical diseases, including breast cancer, osteoarthritis, diabetes, liver, kidney, thyroid and heart disease, showing strong results in evaluation metrics such as accuracy, precision, recall, and F1 score. The novelty of this research lies in its practical integration of multiple advanced AI methods such as ensemble learning, GAN-based data augmentation, and deep learning models applied across seven real-world healthcare datasets, including breast cancer, diabetes, heart disease, and liver disease. By applying more than nine deep learning and machine learning models and comparing their performance using key metrics, the study provides a detailed and data-driven evaluation of the strengths of each model. The CNN models achieved up to 93.1% accuracy on knee X-ray data, while ensemble methods like Random Forest and XGBoost consistently delivered above 90% accuracy across structured datasets like Indian Liver and Heart Disease. The research also used techniques such as SMOTE and GAN to address data imbalance, improving recall for minority classes in conditions such as liver disease and osteoarthritis. This end-to-end framework, from pre-processing and model tuning to performance benchmarking, offers a scalable and reproducible blueprint for using AI in clinical decision-making, taking a meaningful step forward in the pursuit of precision medicine.