<p>Recent advances in deep learning have significantly improved skin cancer classification, yet concerns regarding algorithmic fairness persist because of performance disparities across skin tone groups. Existing methods often attempt to mitigate bias by suppressing sensitive attributes within images. However, they are fundamentally limited by the entanglement of lesion characteristics and skin tone in visual inputs. To address this challenge, we propose a novel contrastive learning framework that leverages explicitly constructed image-text pairs to disentangle lesion condition features from skin tone attributes. Our architecture consists of a shared text encoder and two specialized image encoders that independently align image features with the corresponding textual descriptions of lesion characteristics and skin tone. Furthermore, we measure the semantic distance between lesion conditions and skin color embeddings in both image- and text-embedding spaces and perform optimal representation alignment by matching the distances in the image space to those in the text space. We validated our method using two benchmark datasets, PAD-UFES-20 and Fitzpatrick17k, which span a wide range of skin tones. The experimental results demonstrate that our approach consistently improves both classification accuracy and fairness across multiple evaluation metrics.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FairDITA: Disentangled Image-Text Alignment for Fair Skin Cancer Diagnosis

  • Jiwon Park,
  • Seunggyu Lee,
  • Younghoon Lee

摘要

Recent advances in deep learning have significantly improved skin cancer classification, yet concerns regarding algorithmic fairness persist because of performance disparities across skin tone groups. Existing methods often attempt to mitigate bias by suppressing sensitive attributes within images. However, they are fundamentally limited by the entanglement of lesion characteristics and skin tone in visual inputs. To address this challenge, we propose a novel contrastive learning framework that leverages explicitly constructed image-text pairs to disentangle lesion condition features from skin tone attributes. Our architecture consists of a shared text encoder and two specialized image encoders that independently align image features with the corresponding textual descriptions of lesion characteristics and skin tone. Furthermore, we measure the semantic distance between lesion conditions and skin color embeddings in both image- and text-embedding spaces and perform optimal representation alignment by matching the distances in the image space to those in the text space. We validated our method using two benchmark datasets, PAD-UFES-20 and Fitzpatrick17k, which span a wide range of skin tones. The experimental results demonstrate that our approach consistently improves both classification accuracy and fairness across multiple evaluation metrics.