Objective <p>This study optimizes deep learning model performance for lung cancer classification from chest CT images by systematically evaluating preprocessing techniques, morphological lung segmentation, hybrid ensemble approaches, and transfer learning methods within a unified experimental framework.</p> Methods <p>A dataset of 262 chest CT cases (133 cancerous, 129 non-cancerous) comprising 2,710 standardized images was collected from geographically diverse medical centers in Iran and Iraq. A comprehensive preprocessing pipeline including contrast adjustment, median blur noise reduction, CLAHE enhancement, morphological lung segmentation, standardization, and normalization was implemented. The methodological framework systematically compared CNN performance with and without preprocessing, evaluated hybrid approaches combining CNN feature extraction with ensemble classifiers (SVM, Random Forest, Gradient Boosting), assessed three pre-trained models (ResNet50, VGG16, Xception), and quantified the impact of morphological lung segmentation across all architectures. Grad-CAM visualizations were employed to validate model interpretability and ensure classification decisions derived from clinically relevant anatomical features. All models underwent 10-fold cross-validation using five metrics: accuracy, precision, recall, F1-score, and ROC-AUC, with paired t-tests and effect size calculations ensuring statistical rigor.</p> Results <p>Preprocessing significantly enhanced CNN performance, improving accuracy from 79.13% to 84.45% (<i>p</i> &lt; 0.001, Cohen’s d = 2.264). Hybrid CNN-Gradient Boosting achieved 91.23% accuracy without segmentation, outperforming CNN + RF (88.94%) and CNN + SVM (86.87%). Morphological lung segmentation yielded additional improvements across all models, with CNN + GB reaching 95.28% accuracy, 94.79% precision, 93.89% recall, 94.51% F1-score, and 95.64% ROC-AUC. Pre-trained architectures demonstrated greater relative benefit from segmentation (4.52% to 4.77% improvement) compared to hybrid models (4.05% to 4.06% improvement). Grad-CAM analysis confirmed that pre-trained architectures generated focused attention on diagnostically relevant lung parenchyma, validating clinical trustworthiness.</p> Conclusion <p>This unified methodological framework demonstrates that integrating systematic preprocessing, morphological lung segmentation, and hybrid ensemble architectures significantly enhances lung cancer classification performance. The CNN-Gradient Boosting approach with morphological segmentation presents a clinically viable solution for computer-aided diagnosis, achieving high diagnostic accuracy while maintaining interpretability suitable for radiologist-assisted workflows in diverse healthcare environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimizing Deep Learning Models for Clinical Lung Cancer Detection: Comparative Analysis of Segmentation-Enhanced CNN-Ensemble Approaches in Chest CT Screening

  • Mouna Sadat Hosseini,
  • Naser Raoofi,
  • Sajjad Rezvani Boroujeni,
  • Fatemeh abedi lomer,
  • Hossein Najafzadeh

摘要

Objective

This study optimizes deep learning model performance for lung cancer classification from chest CT images by systematically evaluating preprocessing techniques, morphological lung segmentation, hybrid ensemble approaches, and transfer learning methods within a unified experimental framework.

Methods

A dataset of 262 chest CT cases (133 cancerous, 129 non-cancerous) comprising 2,710 standardized images was collected from geographically diverse medical centers in Iran and Iraq. A comprehensive preprocessing pipeline including contrast adjustment, median blur noise reduction, CLAHE enhancement, morphological lung segmentation, standardization, and normalization was implemented. The methodological framework systematically compared CNN performance with and without preprocessing, evaluated hybrid approaches combining CNN feature extraction with ensemble classifiers (SVM, Random Forest, Gradient Boosting), assessed three pre-trained models (ResNet50, VGG16, Xception), and quantified the impact of morphological lung segmentation across all architectures. Grad-CAM visualizations were employed to validate model interpretability and ensure classification decisions derived from clinically relevant anatomical features. All models underwent 10-fold cross-validation using five metrics: accuracy, precision, recall, F1-score, and ROC-AUC, with paired t-tests and effect size calculations ensuring statistical rigor.

Results

Preprocessing significantly enhanced CNN performance, improving accuracy from 79.13% to 84.45% (p < 0.001, Cohen’s d = 2.264). Hybrid CNN-Gradient Boosting achieved 91.23% accuracy without segmentation, outperforming CNN + RF (88.94%) and CNN + SVM (86.87%). Morphological lung segmentation yielded additional improvements across all models, with CNN + GB reaching 95.28% accuracy, 94.79% precision, 93.89% recall, 94.51% F1-score, and 95.64% ROC-AUC. Pre-trained architectures demonstrated greater relative benefit from segmentation (4.52% to 4.77% improvement) compared to hybrid models (4.05% to 4.06% improvement). Grad-CAM analysis confirmed that pre-trained architectures generated focused attention on diagnostically relevant lung parenchyma, validating clinical trustworthiness.

Conclusion

This unified methodological framework demonstrates that integrating systematic preprocessing, morphological lung segmentation, and hybrid ensemble architectures significantly enhances lung cancer classification performance. The CNN-Gradient Boosting approach with morphological segmentation presents a clinically viable solution for computer-aided diagnosis, achieving high diagnostic accuracy while maintaining interpretability suitable for radiologist-assisted workflows in diverse healthcare environments.