Item performance in context: differential item functioning between Pilot and Formal administration of the Norwegian language test
摘要
The Norwegian language test (No: Norskprøven) administrated by Norwegian Directorate for Higher Education and Skills (Skills Norway) is a high-stakes assessment, the results of which are used by test takers in various ways. However, the item parameters used in multistage testing in Norskprøven are calibrated from a low-stakes situation, the Pilot test. Potential item parameter shift from the Pilot test to the Formal test might be a concern of practitioners since it reduces the test’s reliability and validity. In this study, differential item functioning (DIF) was examined between the Pilot and Formal reading comprehension tests in Norskprøven, using a log-likelihood ratio test method. Purification method was conducted to clean the invariant items and to improve the precision of DIF item detections. The results revealed 10 DIF items with a large effect size. DIF items also showed a tendency to vary more in item discrimination than in item difficulty. Lower discrimination parameters in the Pilot test indicated more random error and might have resulted from the low-stakes situation. Regarding the item features, item format and count of words, there was no clear evidence of being related to DIF. However, more items among the invariant items were piloted across two stages in multistage testing design, in contrast to the DIF items piloted in only one stage. Therefore, test administration and calibration design seem to be more related to the shift in individual item performance rather than to the item features.