The emergence of advanced reasoning models like O1 and DeepSeek-R1 marks a significant advancement in LLM capabilities, enabling applications in complex tasks such as test equating in educational measurement. This study explored test equating without anchor items by leveraging four large language models—GPT-4o, O1-mini, O1-preview, and DeepSeek-R1—to simulate participant data for linking two distinct mathematics tests (X and Y). The two-parameter item response theory (IRT) model was used for equating. Results showed that advanced reasoning models outperformed general-purpose LLMs, with DeepSeek-R1 demonstrating smaller errors than O1-preview, and both outperforming GPT-4o and O1-mini. Increasing the size of linking groups reduced errors, and prompt characteristics such as gender roles influenced error magnitude. This study highlights the potential of advanced reasoning LLMs in enhancing test equating processes.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leveraging Large Reasoning Models for Test Equating Without Anchor Items: A Simulation Study with O1 and DeepSeek-R1

  • Junlei Du,
  • Qinhua Zheng,
  • Shuang Li

摘要

The emergence of advanced reasoning models like O1 and DeepSeek-R1 marks a significant advancement in LLM capabilities, enabling applications in complex tasks such as test equating in educational measurement. This study explored test equating without anchor items by leveraging four large language models—GPT-4o, O1-mini, O1-preview, and DeepSeek-R1—to simulate participant data for linking two distinct mathematics tests (X and Y). The two-parameter item response theory (IRT) model was used for equating. Results showed that advanced reasoning models outperformed general-purpose LLMs, with DeepSeek-R1 demonstrating smaller errors than O1-preview, and both outperforming GPT-4o and O1-mini. Increasing the size of linking groups reduced errors, and prompt characteristics such as gender roles influenced error magnitude. This study highlights the potential of advanced reasoning LLMs in enhancing test equating processes.