<p>Colorectal cancer accounts for nearly one million deaths annually and ranks among the leading causes of cancer mortality worldwide. Accurate automated analysis of histopathological tissue is central to improving diagnostic consistency and reducing pathologists' workload. This study addresses multi-class colorectal cancer classification across six tissue categories using a dataset originally constructed for segmentation, a setting that introduces genuine discriminative challenge beyond what binary or coarse-grained tasks demand. Six pretrained CNN architectures were evaluated under a transfer learning framework; high-capacity models DenseNet201, ResNet152V2, and Xception achieved baseline accuracies up to 88.7%, while lighter architectures failed to capture the morphological detail required for reliable multi-class discrimination. A Feature-Level Fusion ensemble combining these three backbones raised accuracy to 96.0% with macro-AUC of 0.998, but calibration analysis revealed residual overconfidence in the fused outputs (ECE = 0.034, Brier Score = 0.074), a limitation that accuracy figures alone do not expose. This calibration gap, rather than the accuracy ceiling itself, motivates the central contribution of this work: a rank-based ensemble, Rank_DRX, that aggregates model predictions through ordinal rank consensus instead of probability-magnitude concatenation. Applied to the same three backbones, this refinement of rank-based decision fusion achieved a comparable accuracy gain (0.968, ~ 97%) alongside a substantially larger improvement in calibration (ECE = 0.017, Brier Score = 0.045) and Matthews Correlation Coefficient of 0.964. The improvement in classification reliability is most pronounced at the diagnostically critical boundary between Low-Grade and High-Grade intraepithelial neoplasia, where inter-model rank consensus, rather than raw confidence magnitude, resolves subtle morphological ambiguity more effectively than feature-space fusion. Statistical significance was confirmed via McNemar's test, DeLong's AUC comparison, and the Wilcoxon signed-rank test (<i>p</i> &lt; 0.05). A unified interpretability framework, SHAP, LIME, Grad-CAM, and Saliency Maps applied jointly, consistently identified glandular boundaries and epithelial structures as the primary discriminative regions, though these findings remain qualitative pending formal pathologist validation. Together, these results indicate that the principal value of rank-based decision fusion lies in improving prediction reliability at least as much as classification accuracy, a property essential for a diagnostic support system to be considered clinically trustworthy.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fine-grained feature learning ensemble networks for colorectal cancer classification: bridging feature-level fusion with rank-based ensemble model

  • Rokonozzaman Ayon,
  • Md Asif Shahriar Arpon,
  • Montakir Emtiaz Ahmed Rafid,
  • Arif Faisal,
  • Abdus Sattar,
  • Sharun Akter Khushbu

摘要

Colorectal cancer accounts for nearly one million deaths annually and ranks among the leading causes of cancer mortality worldwide. Accurate automated analysis of histopathological tissue is central to improving diagnostic consistency and reducing pathologists' workload. This study addresses multi-class colorectal cancer classification across six tissue categories using a dataset originally constructed for segmentation, a setting that introduces genuine discriminative challenge beyond what binary or coarse-grained tasks demand. Six pretrained CNN architectures were evaluated under a transfer learning framework; high-capacity models DenseNet201, ResNet152V2, and Xception achieved baseline accuracies up to 88.7%, while lighter architectures failed to capture the morphological detail required for reliable multi-class discrimination. A Feature-Level Fusion ensemble combining these three backbones raised accuracy to 96.0% with macro-AUC of 0.998, but calibration analysis revealed residual overconfidence in the fused outputs (ECE = 0.034, Brier Score = 0.074), a limitation that accuracy figures alone do not expose. This calibration gap, rather than the accuracy ceiling itself, motivates the central contribution of this work: a rank-based ensemble, Rank_DRX, that aggregates model predictions through ordinal rank consensus instead of probability-magnitude concatenation. Applied to the same three backbones, this refinement of rank-based decision fusion achieved a comparable accuracy gain (0.968, ~ 97%) alongside a substantially larger improvement in calibration (ECE = 0.017, Brier Score = 0.045) and Matthews Correlation Coefficient of 0.964. The improvement in classification reliability is most pronounced at the diagnostically critical boundary between Low-Grade and High-Grade intraepithelial neoplasia, where inter-model rank consensus, rather than raw confidence magnitude, resolves subtle morphological ambiguity more effectively than feature-space fusion. Statistical significance was confirmed via McNemar's test, DeLong's AUC comparison, and the Wilcoxon signed-rank test (p < 0.05). A unified interpretability framework, SHAP, LIME, Grad-CAM, and Saliency Maps applied jointly, consistently identified glandular boundaries and epithelial structures as the primary discriminative regions, though these findings remain qualitative pending formal pathologist validation. Together, these results indicate that the principal value of rank-based decision fusion lies in improving prediction reliability at least as much as classification accuracy, a property essential for a diagnostic support system to be considered clinically trustworthy.