Saxony-Anhalt is the Worst: Bias Towards German Federal States in Large Language Models
摘要
Recent research demonstrates geographic biases in various Large Language Models that reflects common human biases, which are presumably present in the training data. We hypothesize that these biases also exist on smaller scales. Within Germany, there is still a strong divide between the former states of the German Democratic Republic (the “East”) and those of the Federal Republic of Germany (the “West”) in many respects as well as other perceived geographic disparities. We evaluate the responses of ChatGPT-3.5, ChatGPT-4, and LeoLM for various ratings and estimations by state. Those include objectively measurable values as well as subjective assessments of residents’ characteristics. Experiments are conducted in English and German. We show that there is a very visible bias in both the subjective and the objective ratings, and analyze various effects. In particular, we demonstrate that Eastern states are consistently rated lower (or worse, depending on task), whereas Southern states frequently rate higher. We also discuss models’ behaviors when prompted with these tasks.