Whose Politics Do LLMs Represent? Uncovering Political Bias in LLMs’ Latent Space: A Case Study in Australia
摘要
Large Language Models (LLMs) are increasingly used in political and civic contexts; however, it remains unclear whose perspectives they represent. To examine this question in the latent space of LLMs, we follow the methodology of [4] and train linear probing classifiers on Australian political opinion statements from Roll Call records, the Manifesto Project, and the V-Party dataset to identify party-related representations across 13 models. We then extract party-aligned value vectors and compute persona–party scaling factors using synthetic sociodemographic profiles derived from the Australian Election Study (AES 2022). Results indicate higher probe accuracy for major parties (Liberal and Labor, \(\approx 0.6\) ) than for minor ones (Greens and Nationals, \(\approx 0.5\) ). Across models, age, education, and political position show the most consistent associations with party-aligned representations. We also find that model-generated voting distributions are flatter and exhibit higher entropy than the AES 2022 patterns. Furthermore, the divergence from the AES 2022 dataset is inconsistent and shifts between base and fine-tuned models.