<p>Most current LLM applications in social science are measurement tasks, but they are rarely treated systematically within an explicit measurement framework. The field has largely used LLMs as drop-in replacements for human coders at a single stage of the conventional measurement pipeline, without asking whether a pipeline built around modular human judgment and downstream statistical modeling is still the right architecture. This paper takes a perspective from measurement theory, mapping existing applications onto the classical measurement pipeline and showing that most studies simply swap an LLM in for a human coder at one stage, without asking how a fundamentally different kind of instrument might affect the pipeline itself. We then develop a taxonomy of LLM-based measurement strategies organized by degree of integration into the pipeline, and identify areas where evaluating LLM-based measures requires attention that differs fundamentally from conventional practice. We treat LLM outputs not as automatic measurements but as candidate measures of a distinctive kind. The aim is not to provide a single workflow but to identify the issues that must be seriously considered when using LLMs for measurement.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Beyond the Conventional Pipeline: Large Language Models for Social Science Measurement

  • Le Bao

摘要

Most current LLM applications in social science are measurement tasks, but they are rarely treated systematically within an explicit measurement framework. The field has largely used LLMs as drop-in replacements for human coders at a single stage of the conventional measurement pipeline, without asking whether a pipeline built around modular human judgment and downstream statistical modeling is still the right architecture. This paper takes a perspective from measurement theory, mapping existing applications onto the classical measurement pipeline and showing that most studies simply swap an LLM in for a human coder at one stage, without asking how a fundamentally different kind of instrument might affect the pipeline itself. We then develop a taxonomy of LLM-based measurement strategies organized by degree of integration into the pipeline, and identify areas where evaluating LLM-based measures requires attention that differs fundamentally from conventional practice. We treat LLM outputs not as automatic measurements but as candidate measures of a distinctive kind. The aim is not to provide a single workflow but to identify the issues that must be seriously considered when using LLMs for measurement.