Beyond the Conventional Pipeline: Large Language Models for Social Science Measurement
摘要
Most current LLM applications in social science are measurement tasks, but they are rarely treated systematically within an explicit measurement framework. The field has largely used LLMs as drop-in replacements for human coders at a single stage of the conventional measurement pipeline, without asking whether a pipeline built around modular human judgment and downstream statistical modeling is still the right architecture. This paper takes a perspective from measurement theory, mapping existing applications onto the classical measurement pipeline and showing that most studies simply swap an LLM in for a human coder at one stage, without asking how a fundamentally different kind of instrument might affect the pipeline itself. We then develop a taxonomy of LLM-based measurement strategies organized by degree of integration into the pipeline, and identify areas where evaluating LLM-based measures requires attention that differs fundamentally from conventional practice. We treat LLM outputs not as automatic measurements but as candidate measures of a distinctive kind. The aim is not to provide a single workflow but to identify the issues that must be seriously considered when using LLMs for measurement.