Design of CALT
摘要
In summary, when passage-based items were treated as a polytomous item, the problem of violation of local independence assumption could be circumvented. The CBLT was supported by sufficient unidimensionality, and the distinctness of the construct tapped by the listening and reading sections informed the decision of separate IRT calibrations of the two sections. The 2PLM was chosen to model the dichotomous items out of theoretical and practical considerations. Compared to the GRM, the GPCM was found to be a preferred model for the polytomous items not only in terms of overall model fit but also item-level model fit. The items that showed misfit to the model, that failed to meet the requirements for item banking concerning the discrimination and difficulty parameters, and that exhibited significant and practical gender DIF, were removed from the item pool. The final item pool was divided into 4 sub-pools based on task types, so as to be used as 4 sub-tests in later CALT. The scale information for the first sub-pool of short conversation was fairly peaked, indicating that it provided the largest amount of information for test takers in the medium range of the distribution and less information for test takers at the extreme ends of the distribution. Nonetheless, the scale information for the other three sub-pools were comparatively flat, suggesting that test takers along the ability continuum can be measured in an approximately equally quick and precise way. Since the CALT was designed to measure students’ English proficiency as a norm-referenced test, the rectangular distributions of these three sub-pools could be considered ideal and desirable. Moreover, the four sub-pools, though slightly positively skewed, were almost parallel to the ability distribution of the target population of test takers, implying that nonconvergent ability estimates in CALT could in general be avoided.