Prediction of lung cancer outcome is a vital aspect in the medical diagnostics. This study works on building predictive models with the help of different machine learning methodologies. The research comprised 309 patients’ data, which obtained demographic details, smoking histories, medical and family histories, and other health indicators such as age, BMI, lung functions, and blood pressures. Data preprocessing involved cleaning, imputation of missing values by the means of mean/mode, normalization by min–max scaling, and encoding categorical variables. Descriptive statistics and Pearson correlation analysis were performed to explore distributions of data and important relationships between features and the target, whether the patient has lung cancer. The study assesses several algorithms, ranging from logistic regression through support vector machines, decision trees, random forests, and gradient boosting methods. Performance evaluation of all algorithms was done using accuracy, precision, recall, F1 score, and AUC-ROC score. Hyperparameter tuning via grid search was done on all models, and class imbalance handling was done using SMOTE. The feature importance analysis is done using SHAP values, which highlighted the essential contributors such as smoking history and metrics around lung function. The findings indicate that the gradient boosting method performed better than the other models, achieving the highest accuracy and AUC-ROC score. The study presents the opportunities machine learning courses lend to enhancing the accuracy of lung cancer prediction with perspectives on implementing AI-based diagnosis in outpatient settings. The findings provide substantial insights for early diagnosis and therapeutic tactics to support oncologists during decision-making and electronic health record integration. Future work should investigate computational efficiency and scalability concerning wider and scattered datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Comprehensive Analysis of Lung Cancer Prediction Using Machine Learning Models

  • Aditya Patil,
  • Sanket Lodha,
  • Arun Patil,
  • Atharv Ekavire,
  • Sudhir Chitnis,
  • Vaishali Patil

摘要

Prediction of lung cancer outcome is a vital aspect in the medical diagnostics. This study works on building predictive models with the help of different machine learning methodologies. The research comprised 309 patients’ data, which obtained demographic details, smoking histories, medical and family histories, and other health indicators such as age, BMI, lung functions, and blood pressures. Data preprocessing involved cleaning, imputation of missing values by the means of mean/mode, normalization by min–max scaling, and encoding categorical variables. Descriptive statistics and Pearson correlation analysis were performed to explore distributions of data and important relationships between features and the target, whether the patient has lung cancer. The study assesses several algorithms, ranging from logistic regression through support vector machines, decision trees, random forests, and gradient boosting methods. Performance evaluation of all algorithms was done using accuracy, precision, recall, F1 score, and AUC-ROC score. Hyperparameter tuning via grid search was done on all models, and class imbalance handling was done using SMOTE. The feature importance analysis is done using SHAP values, which highlighted the essential contributors such as smoking history and metrics around lung function. The findings indicate that the gradient boosting method performed better than the other models, achieving the highest accuracy and AUC-ROC score. The study presents the opportunities machine learning courses lend to enhancing the accuracy of lung cancer prediction with perspectives on implementing AI-based diagnosis in outpatient settings. The findings provide substantial insights for early diagnosis and therapeutic tactics to support oncologists during decision-making and electronic health record integration. Future work should investigate computational efficiency and scalability concerning wider and scattered datasets.