<p>Skin tone affects artificial intelligence (AI) performance in dermatology. While labeling datasets for skin tone could improve algorithm generalizability for detecting dermatologic malignancies, large-scale validation of skin tone assessments is lacking. This prospective observational study assessed reliability of subjective tools (Fitzpatrick Skin Type [FST], Monk Skin Tone [MST], Pantone SkinTone Guide) and an objective colorimeter for in-person and photography-based settings to evaluate utility for labeling dermoscopic datasets. Colorimetry (gold standard for color measurement) demonstrated high precision with in-person measurements. Of subjective scales, MST demonstrated slightly tighter clustering in the color space and high repeatability for in-person and photography-based assessments (latter varied by lighting). Dermoscopic image-extracted color values correlated poorly with colorimetry values. For subjective ratings, MST more effectively captured differences in AI melanoma classification scores than FST. Findings underscore that FST is not a proxy for skin tone; an important role remains for skin tone assessment to improve AI performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating skin tone scales for dermatologic dataset labeling: a prospective-comparative study

  • Vanessa R. Weir,
  • Yingjoy Li,
  • Maura C. Gillis,
  • Nicholas R. Kurtansky,
  • Trina Salvador,
  • Allan C. Halpern,
  • Kelly C. Nelson,
  • Jenna C. Lester,
  • Veronica Rotemberg

摘要

Skin tone affects artificial intelligence (AI) performance in dermatology. While labeling datasets for skin tone could improve algorithm generalizability for detecting dermatologic malignancies, large-scale validation of skin tone assessments is lacking. This prospective observational study assessed reliability of subjective tools (Fitzpatrick Skin Type [FST], Monk Skin Tone [MST], Pantone SkinTone Guide) and an objective colorimeter for in-person and photography-based settings to evaluate utility for labeling dermoscopic datasets. Colorimetry (gold standard for color measurement) demonstrated high precision with in-person measurements. Of subjective scales, MST demonstrated slightly tighter clustering in the color space and high repeatability for in-person and photography-based assessments (latter varied by lighting). Dermoscopic image-extracted color values correlated poorly with colorimetry values. For subjective ratings, MST more effectively captured differences in AI melanoma classification scores than FST. Findings underscore that FST is not a proxy for skin tone; an important role remains for skin tone assessment to improve AI performance.