Inter-rater bias in medical image segmentation is often overlooked when automatic models, machine learning included, are considered. However, nowadays when deep neural networks (DNNs) turn to be prevalent computational tools, and since most of the training processes are supervised to some extent, its influence on the prediction quality and reliability should be examined. In this study we employed a commonly-used supervised segmentation framework to quantify the influence of the training set annotator on the inferred segmentation masks. We found out that an inter-rater bias was amplified and became more consistent when the DNN’s predicted segmentations rather than the manual annotations themselves were compared. Specifically, we used two different datasets: brain MRIs of Multiple Sclerosis (MS) patients that were annotated by two raters with different level of expertise; and Intracerebral Hemorrhage (ICH) CT scans with manual and semi-manual segmentations. The results obtained imply a worrisome clinical implication of a DNN bias induced by an inter-rater bias during training. Specifically, we found a consistent underestimate of MS-lesion loads when calculated from segmentation predictions of a DNN trained on segmentation masks provided by the less experienced rater. In the same manner, the differences in ICH volumes calculated based on outputs of identical DNNs, each trained on annotations from a different source were more consistent and larger than the differences in volumes between the manual and semi-manual annotations used for training.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The Worrisome Impact of an Inter-rater Bias on Neural Network Training

  • Or Shwartzman,
  • Harel Gazit,
  • Gal Ben-Aryeh,
  • Ilan Shalef,
  • Tammy Riklin Raviv

摘要

Inter-rater bias in medical image segmentation is often overlooked when automatic models, machine learning included, are considered. However, nowadays when deep neural networks (DNNs) turn to be prevalent computational tools, and since most of the training processes are supervised to some extent, its influence on the prediction quality and reliability should be examined. In this study we employed a commonly-used supervised segmentation framework to quantify the influence of the training set annotator on the inferred segmentation masks. We found out that an inter-rater bias was amplified and became more consistent when the DNN’s predicted segmentations rather than the manual annotations themselves were compared. Specifically, we used two different datasets: brain MRIs of Multiple Sclerosis (MS) patients that were annotated by two raters with different level of expertise; and Intracerebral Hemorrhage (ICH) CT scans with manual and semi-manual segmentations. The results obtained imply a worrisome clinical implication of a DNN bias induced by an inter-rater bias during training. Specifically, we found a consistent underestimate of MS-lesion loads when calculated from segmentation predictions of a DNN trained on segmentation masks provided by the less experienced rater. In the same manner, the differences in ICH volumes calculated based on outputs of identical DNNs, each trained on annotations from a different source were more consistent and larger than the differences in volumes between the manual and semi-manual annotations used for training.