<p>This paper informs users of data collected in international large-scale assessments (ILSA), by presenting arguments underlining the importance of considering two design features employed in these studies. We examine a common misconception stating that the uncertainty arising from the assessment design is negligible compared with that arising from the sampling design. This misconception can lead to the erroneous conclusion that there is always a relatively low risk of ignoring the uncertainty arising from the assessment design when reporting estimates of population parameters. We use the design effect framework to assess the impact that the sampling and the assessment design have on the estimation. We first evaluate the loss in efficiency in the estimation of a population parameter attributable to each of the designs. We then examine whether knowledge about the effect of one design feature can justify any belief about the effect of the other design feature. We repeat this examination across different parameters characterizing the achievement distribution in a population. We provide empirical results using data collected for PIRLS 2016. Our empirical results can be summarized in two general findings. First, when estimating mean achievement, the effect of the sampling design is often substantially larger than that of the assessment design. This finding might explain the misconception we try to address. However, we show that this is not true in all instances, and the magnitude of the difference between both design effects is context dependent and hence not generalizable. Second, differences in design effects become less predictable when estimating other parameters, e.g. the proportion of students reaching a certain threshold in the achievement scale (i.e., benchmarks), or an association estimated using linear regression. This contribution underlines that accounting for all sources of uncertainty in the estimation is of paramount importance to obtain credible inferences. We conclude that it is difficult to justify a priori the belief that the effect of the sampling design in the estimation is always greater than that of the assessment design.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating uncertainty: the impact of the sampling and assessment design on statistical inference in the context of ILSA

  • Diego Cortes,
  • Dirk Hastedt,
  • Sabine Meinck

摘要

This paper informs users of data collected in international large-scale assessments (ILSA), by presenting arguments underlining the importance of considering two design features employed in these studies. We examine a common misconception stating that the uncertainty arising from the assessment design is negligible compared with that arising from the sampling design. This misconception can lead to the erroneous conclusion that there is always a relatively low risk of ignoring the uncertainty arising from the assessment design when reporting estimates of population parameters. We use the design effect framework to assess the impact that the sampling and the assessment design have on the estimation. We first evaluate the loss in efficiency in the estimation of a population parameter attributable to each of the designs. We then examine whether knowledge about the effect of one design feature can justify any belief about the effect of the other design feature. We repeat this examination across different parameters characterizing the achievement distribution in a population. We provide empirical results using data collected for PIRLS 2016. Our empirical results can be summarized in two general findings. First, when estimating mean achievement, the effect of the sampling design is often substantially larger than that of the assessment design. This finding might explain the misconception we try to address. However, we show that this is not true in all instances, and the magnitude of the difference between both design effects is context dependent and hence not generalizable. Second, differences in design effects become less predictable when estimating other parameters, e.g. the proportion of students reaching a certain threshold in the achievement scale (i.e., benchmarks), or an association estimated using linear regression. This contribution underlines that accounting for all sources of uncertainty in the estimation is of paramount importance to obtain credible inferences. We conclude that it is difficult to justify a priori the belief that the effect of the sampling design in the estimation is always greater than that of the assessment design.