Exploring Attenuation of Reliability in Categorical Subscore Reporting
摘要
Research on subscores has consistently advocated discontinuing reporting when they lack sufficient psychometric properties, yet many educational agencies and operational testing programs have not changed. This may be due to several real-world complications such as user demand, competitors, or contractual obligations. Given these challenges, some test providers have continued to report subscores but in a categorical format to mitigate misinterpretation of small differences likely due to error. However, there also exists robust literature on how continuous scores grouped into categories can be less reliable than the scores from which they were constructed. Using a resampling design based on real data, a variation on the Lord–Wingersky recursion described by Feinberg and von Davier (J Educ Behav Stat 45(5):515–533, 2020) was applied to compare two different approaches of discretizing subscore into categories, relative-to-self and relative-to-average. Results support categorical subscores as a promising alternative when the reliability of a continuous subscore may be too imprecise to report a numeric score. Implications for practice, operational utility, and the extent to which categorical subscore reporting represents an appropriate compromise are discussed.