Developing language treebanks is a challenging and time-consuming process, especially for under-resourced languages with limited data. This study proposes the use of active learning approaches to automate aspects of this process, aiming to reduce both annotation duration and cost. We propose practical active annotation schemes in which experts strategically select sentences for the training set, initiating a circular process involving annotation prediction, expert correction, and model retraining. To validate the feasibility of these schemes, we applied them to 300 annotated sentences from the newly created Pomak corpus (an under-resourced language), recently published in Universal Dependencies treebanks, as the result of our efforts. Through experiments involving a simple weighted summation of annotation errors, we identified an optimal strategy. This strategy resulted in a 69% reduction in total annotation duration and an associated 81% decrease in total corresponding cost compared to a typical manual annotation, demonstrating the effectiveness of this approach.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring Active Learning Approaches in Treebank Development

  • Vasileios Arampatzakis,
  • Vivian Stamou,
  • Stella Markantonatou,
  • George Pavlidis

摘要

Developing language treebanks is a challenging and time-consuming process, especially for under-resourced languages with limited data. This study proposes the use of active learning approaches to automate aspects of this process, aiming to reduce both annotation duration and cost. We propose practical active annotation schemes in which experts strategically select sentences for the training set, initiating a circular process involving annotation prediction, expert correction, and model retraining. To validate the feasibility of these schemes, we applied them to 300 annotated sentences from the newly created Pomak corpus (an under-resourced language), recently published in Universal Dependencies treebanks, as the result of our efforts. Through experiments involving a simple weighted summation of annotation errors, we identified an optimal strategy. This strategy resulted in a 69% reduction in total annotation duration and an associated 81% decrease in total corresponding cost compared to a typical manual annotation, demonstrating the effectiveness of this approach.