Leveraging Large Reasoning Models for Test Equating Without Anchor Items: A Simulation Study with O1 and DeepSeek-R1
摘要
The emergence of advanced reasoning models like O1 and DeepSeek-R1 marks a significant advancement in LLM capabilities, enabling applications in complex tasks such as test equating in educational measurement. This study explored test equating without anchor items by leveraging four large language models—GPT-4o, O1-mini, O1-preview, and DeepSeek-R1—to simulate participant data for linking two distinct mathematics tests (X and Y). The two-parameter item response theory (IRT) model was used for equating. Results showed that advanced reasoning models outperformed general-purpose LLMs, with DeepSeek-R1 demonstrating smaller errors than O1-preview, and both outperforming GPT-4o and O1-mini. Increasing the size of linking groups reduced errors, and prompt characteristics such as gender roles influenced error magnitude. This study highlights the potential of advanced reasoning LLMs in enhancing test equating processes.