错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Annotation of Text Corpora by Sentiment and Irony in a Project of Citizen Science

  • I. V. Paramonov,
  • A. Y. Poletaev

摘要

Abstract—

This paper studies the construction of a corpus of sentences annotated by general sentiment into four classes (positive, negative, neutral, and mixed), a corpus of phrasemes annotated by sentiment into three classes (positive, negative, and neutral), and a corpus of sentences annotated by the presence or absence of irony. The annotation is conducted by volunteers within the project Preparing Texts for Algorithms on the People of Science website. Based on the available knowledge of the subject area for each of the problems, guidelines for the annotators are compiled. A methodology for the statistical processing of the annotation results is also developed based on analyzing the distributions and agreement measures of the annotations of different annotators. For annotating sentences by irony and phrasemes by sentiment, the agreement measures are quite high (the full agreement rate is 0.60–0.99), while for annotating sentences by general sentiment, the agreement is low (the full agreement rate is 0.40), apparently due to the higher complexity of the problem. It is also shown that the performance of automatic algorithms for sentence sentiment analysis improves by 12–13% when using a corpus on whose sentences all annotators (3–5 people) agree compared with a corpus annotated by only one volunteer.