The fidelity of large language models in social science: a human-ai parallel experimental test of middle-range theory
摘要
Middle-range theory, as conceptualised by Merton, is distinguished by its empirical testability and its capacity to bridge abstract theory and empirical social inquiry. In this paper, we ask whether large language models (LLMs) can be used as surrogate human actors in the testing and reformulation of middle-range theories. Focusing on agenda-setting theory, we investigate how agenda priority, issue detailedness and geographical proximity shape public attention and policy attitudes in the context of climate change. Drawing on climate change public discourses on BiliBili, a major Chinese video-sharing social media platform, we design parallel experiments with human participants, LLM-based virtual agents and LLM-based virtual agents embedded with human personas. By comparing the results across three experiments, we explore how LLMs reproduce or diverge from the dynamics anticipated by agenda-setting theory. The study engages with current theoretical debates on the boundaries of LLMs for social theory construction, addressing their potential and limitations not only as tools for simulation but also as platforms of theoretical innovation. Our findings reveal that while LLMs demonstrate high directional fidelity, they struggle with simulating the magnitude and structural complexity of human behaviours. We thus provide a methodological framework for the use of LLMs as heuristic simulators rather than empirical substitutes of human experiments. We argue that the combination of middle-range theory and generative models yields new possibilities for computational social science while also raising fundamental questions about the nature of empirical validation in the age of generative AI.