Assessing the Sociocultural Alignment of Large Language Models: An Empirical Study of Chinese-Speaking Populations
摘要
With the recent advances taking place in producing increasingly sophisticated large language models (LLMs), there is an increasing need for LLMs to align with the sociocultural characteristics of potential end users. However, the extent to which LLMs presently capture these characteristics is not fully known, and methodologies for improving LLMs in this regard is an active area of research. In this study, we empirically evaluate some the most widely available LLMs (Phi3, Llama3, Mixtral, Mistral, Gemma, Qwen, and GPT-3.5 Turbo) on how well they represent the perspectives of individuals from different Chinese-speaking regions across Asia, finding that notable differences exist in the performance of these models across these regions.