The Database of Constructions with Lexical Repetitions “RepLeCon” and Inter-Annotator Agreement
摘要
The chapter deals with some issues that emerge in the development of multilingual database of constructions with lexical repetitions “RepLeCon.” The “RepLeCon” database is a corpora-based resource, aimed to providing with data from cleaned-up and often sophisticated queries enriched with supplementary information. The building of such a database requires manual post-processing made by several annotators. To make the annotations maximally reliable, we use inter-annotator agreement. We took a sample of equative tautologies such as Family is family, annotated regarding (1) their rhetorical relations with the preceding and the following discourse fragments and (2) their role (i.e., the nucleus or the satellite) in these relations. The sample was tagged by two experienced raters. Then the data was examined with the use of agreement measures. We obtained the following results: First, the values of the main agreement measures (namely, Gwet’s AC, Fleiss’ Kappa, Krippendorff’s Alpha, Conger’s Kappa, Brennan, and Prediger’s AC) are in the range of [0.73; 0.98]. Thus, the raters show either an almost perfect or substantial agreement, so the annotation instruction has been successfully compiled and does not need significant revision. Second, it appeared that raters assign the particular rhetorical relations significantly less consistently than the role of tautologies in the rhetorical relation. This may be due to the effect of the number of categories in the annotation scheme on the consistency of the annotation (since there are only two possibilities of the role of constructions compared to a longer list of the rhetorical relations, annotators’ decision for the former scheme requires less effort). Third, the rhetorical relations with the following discourse fragment are annotated in a less consistent way than the rhetorical relations with the preceding fragment.