In interactive information retrieval (IIR), the inherent variability in user behavior and the complexities of user-system interactions pose significant challenges to reproducibility, which remain largely underexplored. To address these challenges, we propose a three-level model for evaluating the reproducibility of IIR experiments through the similarity of experimental findings, underlying measurements, and user behavior across original and reproduction studies. For each level, we introduce specific criteria, offering a structured framework for assessing reproducibility in IIR research. We demonstrate the framework’s utility in a case study where we simulate ideal reproductions by repeatedly dividing the sessions from a user experiment into two groups, treating one as the original and the other as the reproduction, while maintaining consistent experimental conditions. This approach enables us to focus on IIR-specific effects within our methodology. Our analysis indicates that the high variability in user behavior can hinder successful reproductions, raising concerns about the reliability of experimental results. Borderline significant results, in particular, are frequently not reproducible, despite an optimal reproduction setup. Furthermore, we find that cognitive abilities significantly influence search behavior, emphasizing the need to account for these characteristics in the design of future tests.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Towards Reproducibility of Interactive Retrieval Experiments: Framework and Case Study

  • Jana Isabelle Friese,
  • Norbert Fuhr

摘要

In interactive information retrieval (IIR), the inherent variability in user behavior and the complexities of user-system interactions pose significant challenges to reproducibility, which remain largely underexplored. To address these challenges, we propose a three-level model for evaluating the reproducibility of IIR experiments through the similarity of experimental findings, underlying measurements, and user behavior across original and reproduction studies. For each level, we introduce specific criteria, offering a structured framework for assessing reproducibility in IIR research. We demonstrate the framework’s utility in a case study where we simulate ideal reproductions by repeatedly dividing the sessions from a user experiment into two groups, treating one as the original and the other as the reproduction, while maintaining consistent experimental conditions. This approach enables us to focus on IIR-specific effects within our methodology. Our analysis indicates that the high variability in user behavior can hinder successful reproductions, raising concerns about the reliability of experimental results. Borderline significant results, in particular, are frequently not reproducible, despite an optimal reproduction setup. Furthermore, we find that cognitive abilities significantly influence search behavior, emphasizing the need to account for these characteristics in the design of future tests.