Despite the promising results of large language models (LLMs) in labeling tasks, further exploration is needed to leverage them effectively for linguistic data annotation. One of the most challenging tasks in this regard is labeling discourse structures, which is highly subjective and often involves ambiguity in class description. In this paper, we address the challenge of using LLMs for hybrid annotation of the discourse structure in open-domain dialogues, relying on Eggins and Slade’s speech function theory. We conduct a comparative analysis between model-generated annotations and human annotations, exploring the potential of LLM-assisted annotation as a viable alternative to crowdsourcing.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Redefining Annotation Practices: Leveraging Large Language Models for Discourse Annotation

  • Lidiia Ostyakova,
  • Anna Mikhailova,
  • Vasily Konovalov

摘要

Despite the promising results of large language models (LLMs) in labeling tasks, further exploration is needed to leverage them effectively for linguistic data annotation. One of the most challenging tasks in this regard is labeling discourse structures, which is highly subjective and often involves ambiguity in class description. In this paper, we address the challenge of using LLMs for hybrid annotation of the discourse structure in open-domain dialogues, relying on Eggins and Slade’s speech function theory. We conduct a comparative analysis between model-generated annotations and human annotations, exploring the potential of LLM-assisted annotation as a viable alternative to crowdsourcing.