<p>Deep learning shows strong potential for automated skin cancer detection, but clinical adoption requires models that maintain high diagnostic accuracy while demonstrating robustness to real-world data variations. Multi-modal approaches integrating dermoscopic images, clinical images, and patient metadata achieve superior performance compared to single-modality methods. However, existing fusion strategies inadequately address the performance-robustness trade-off, often improving accuracy on lab-quality data while compromising reliability under corrupted or noisy inputs typical of real-world clinical settings. We propose DermFormer, a transformer-based multi-modal architecture addressing this limitation through entropy-weighted ensemble classification heads and a hybrid fusion mechanism that preserves uni-modal representations while capturing inter-modality relationships. Our method combines dermoscopic and clinical images with tabular metadata using hierarchical transformers and cross-attention for modality integration. The entropy-weighted ensemble dynamically adjusts modality contributions based on prediction confidence, enabling dynamic feature selection when individual modalities are corrupted. We evaluate DermFormer on the Derm7pt dataset for multi-class diagnosis and seven-point checklist classification under clean and corrupted conditions. DermFormer achieves state-of-the-art performance (diagnosis accuracy: 0.779, F-score: 0.684) while maintaining superior robustness to common corruptions including Gaussian noise, motion blur, and JPEG compression. By maintaining performance under realistic clinical conditions, this work addresses a critical adoption barrier for automated diagnostic systems, enabling reliable AI-assisted dermatology across diverse healthcare settings and supporting earlier melanoma detection at scale. Code available at: <a href="https://github.com/xraikeele/DermFormer">https://github.com/xraikeele/DermFormer</a></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DermFormer: nested multi-modal vision transformers for robust skin cancer detection

  • Matthew J. Cockayne,
  • Marco Ortolani,
  • Baidaa Al-Bander

摘要

Deep learning shows strong potential for automated skin cancer detection, but clinical adoption requires models that maintain high diagnostic accuracy while demonstrating robustness to real-world data variations. Multi-modal approaches integrating dermoscopic images, clinical images, and patient metadata achieve superior performance compared to single-modality methods. However, existing fusion strategies inadequately address the performance-robustness trade-off, often improving accuracy on lab-quality data while compromising reliability under corrupted or noisy inputs typical of real-world clinical settings. We propose DermFormer, a transformer-based multi-modal architecture addressing this limitation through entropy-weighted ensemble classification heads and a hybrid fusion mechanism that preserves uni-modal representations while capturing inter-modality relationships. Our method combines dermoscopic and clinical images with tabular metadata using hierarchical transformers and cross-attention for modality integration. The entropy-weighted ensemble dynamically adjusts modality contributions based on prediction confidence, enabling dynamic feature selection when individual modalities are corrupted. We evaluate DermFormer on the Derm7pt dataset for multi-class diagnosis and seven-point checklist classification under clean and corrupted conditions. DermFormer achieves state-of-the-art performance (diagnosis accuracy: 0.779, F-score: 0.684) while maintaining superior robustness to common corruptions including Gaussian noise, motion blur, and JPEG compression. By maintaining performance under realistic clinical conditions, this work addresses a critical adoption barrier for automated diagnostic systems, enabling reliable AI-assisted dermatology across diverse healthcare settings and supporting earlier melanoma detection at scale. Code available at: https://github.com/xraikeele/DermFormer