Multi-domain Histopathological Imaging Biomarkers for Machine-learning-based Classification of Lung Squamous Cell Carcinoma and Adenocarcinoma
摘要
Deep learning can classify histopathological subtypes of non-small-cell lung cancer (NSCLC) but often requires large datasets and heavy compute, limiting clinical use. We evaluate a lightweight alternative that pairs multi-domain hand-engineered features with conventional machine-learning classifiers.
MethodsWe assembled a moderate dataset of 644 NSCLC histopathology images and extracted diverse features spanning original histopathological descriptors (e.g., keratinization field, glandular differentiation), physics-based measures (entropy, fractal dimension), mathematical descriptors (Hu moments, DCT/FFT energy), aesthetics (colorfulness, prevailing color), histologist-relevant metrics (nuclear perimeter/area), statistical summaries (percentiles, dispersion), and computer-vision features (HOG, keypoints). Random-forest–based feature selection identified informative predictors used to train support vector machines (SVM), gradient boosting, logistic regression, and decision trees. Performance was assessed via accuracy, recall, F1-score, and AUC with five-fold cross-validation. Generalizability was tested on an external set of 95 NSCLC images.
ResultsMulti-domain features achieved strong performance with classical models. SVM obtained the highest cross-validated accuracy (0.85) and mean AUC (0.89). Gradient boosting was competitive (accuracy 0.84; AUC 0.87). Logistic regression outperformed decision trees overall. On external validation, the SVM trained with multi-domain features maintained robust performance (accuracy 0.84; AUC 0.85), indicating good generalizability despite the moderate training set.
ConclusionSVM classifiers trained on carefully selected multi-domain features provide stable, high-quality NSCLC subtype classification without reliance on GPUs or large datasets. This approach offers a practical, resource-efficient pathway for clinical deployment, where computational capacity and data volume are often constrained.