On Effective Multi-class Skin Lesion Classification and Skin-Tone Bias Quantification
摘要
Accurate multi-class skin lesion classification from dermatological images is critical for effective patient care, though often hindered by limited harmonized data and unaddressed skin tone biases. This work presents an analysis of multi-class classification and bias quantification using a large, harmonized image dataset compiled from 13 public sources. We developed and validated an effective deep learning model for skin tone classification from images, which then pseudo-labeled unannotated images, revealing significant imbalances favoring lighter skin tones. Our multi-class models demonstrated strong overall performance in classifying six common skin lesion types. Critically, bias analysis revealed consistently higher model performance metrics on darker skin tones compared to lighter tones, despite significant underrepresentation of darker tones in the data. This counterintuitive finding calls for dedicated research to understand these disparities and prioritize the acquisition of more balanced, human-validated datasets, essential for developing fair and reliable AI solutions for all patient populations.