Interrater reliability of Olympic winter sports judges—a comparative analysis
摘要
In artistic sports, judges are employed to assess performance quality. Discrepancies between their scores can undermine confidence in the objectivity of these assessments and compromise spectators’ belief in the integrity of competitions. Previous research suggests that such discrepancies occur to a greater extent and more frequently in some sports and less frequently in others, implicating differences in scoring systems as a potential cause. However, methodological limitations in assessing interrater reliability, along with variations in judging expertise, hinder meaningful comparisons across sports. This study aims to assess and compare the interrater reliability of Olympic Figure Skating, Ski Jumping, and Snowboarding judges, which are governed by distinct scoring systems.
MethodsInterrater reliability of Figure Skating, Ski Jumping and Snowboarding judges was evaluated using the intraclass correlation coefficient (ICC). Statistical comparisons were conducted across these sports based on ICC values. Descriptive analyses were also performed to examine the extent to which judges utilized the available scoring scales.
ResultsWhile judges in all three sports demonstrated high overall interrater reliability, Snowboarding judges exhibited exceptional reliability, indicating nearly perfect agreement and consistency. In contrast, greater variability was observed among judges in Figure Skating and Ski Jumping. Descriptive analyses revealed that Snowboarding judges utilized the full scoring range, whereas judges in Figure Skating and Ski Jumping displayed significant limitations in their use of the scoring scale.
ConclusionThese findings suggest that the interrater reliability of sport judges is heavily influenced by the structure of the scoring systems employed. Contrary to present beliefs, more detailed and prescriptive scoring rules paradoxically result in less reliable assessments by judges.