Disagreement between annotators is often viewed as a sign of low data quality. In reality, the idea that a single underlying “ground truth” exists is often simply not true. Still pursuing it, e.g., by majority voting, can remove valuable nuances and perspectives from the data, especially for inherently subjective tasks. Recent research increasingly started to leverage disagreement between annotators by using unaggregated annotations for the training of models. More often than not, these models are still used to predict one single output. In order to truly embrace different perspectives, they should not just be considered during the training, but also when making predictions and presenting them. This chapter compares three strategies to leverage disagreement for text classification: a probability-based multi-label method, an ensemble system, and instruction tuning. The approaches were evaluated on hate speech and abusive conversation detection. To compare the performance of embracing disagreement versus only using majority label during the training, we conducted an online survey. Additionally, the survey also investigated whether potential users prefer a single prediction or a multi-label distribution based on different perspectives. The results show that in hate speech detection, the multi-label method performs best, even though it is less complex than the other two approaches. In abusive conversation detection, instruction tuning achieves the best performance given the sparse available data. The results of the survey indicate that the output for the multi-label models are considered a better representation of the texts than a single-label model.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Embracing Disagreement in Text Classification During Training and Prediction

  • Jin Xu,
  • Mariët Theune,
  • Daniel Braun

摘要

Disagreement between annotators is often viewed as a sign of low data quality. In reality, the idea that a single underlying “ground truth” exists is often simply not true. Still pursuing it, e.g., by majority voting, can remove valuable nuances and perspectives from the data, especially for inherently subjective tasks. Recent research increasingly started to leverage disagreement between annotators by using unaggregated annotations for the training of models. More often than not, these models are still used to predict one single output. In order to truly embrace different perspectives, they should not just be considered during the training, but also when making predictions and presenting them. This chapter compares three strategies to leverage disagreement for text classification: a probability-based multi-label method, an ensemble system, and instruction tuning. The approaches were evaluated on hate speech and abusive conversation detection. To compare the performance of embracing disagreement versus only using majority label during the training, we conducted an online survey. Additionally, the survey also investigated whether potential users prefer a single prediction or a multi-label distribution based on different perspectives. The results show that in hate speech detection, the multi-label method performs best, even though it is less complex than the other two approaches. In abusive conversation detection, instruction tuning achieves the best performance given the sparse available data. The results of the survey indicate that the output for the multi-label models are considered a better representation of the texts than a single-label model.