Redefining Annotation Practices: Leveraging Large Language Models for Discourse Annotation
摘要
Despite the promising results of large language models (LLMs) in labeling tasks, further exploration is needed to leverage them effectively for linguistic data annotation. One of the most challenging tasks in this regard is labeling discourse structures, which is highly subjective and often involves ambiguity in class description. In this paper, we address the challenge of using LLMs for hybrid annotation of the discourse structure in open-domain dialogues, relying on Eggins and Slade’s speech function theory. We conduct a comparative analysis between model-generated annotations and human annotations, exploring the potential of LLM-assisted annotation as a viable alternative to crowdsourcing.