Culturally-Grounded Multi-dimensional Alignment of LLMs with Chinese Social Values
摘要
Despite Large Language Models (LLMs) have demonstrated remarkable capabilities, they also present unintended biases and harmful behaviors driven by encoded values, emphasizing the urgent need to align LLMs’ outputs with human intentions. However, current studies on values alignment face two key limitations: lacking alignment diversity and lacking multi-dimensional alignment strategies. In this paper, we propose a novel framework called ValueAlignment, aiming to expand research on LLMs’ alignment with culture-specific values. As a case study, we focus on Chinese Social Values, which reflect the values embedded in a unique community and offer a different perspective from the predominant western-centric framework. We first construct C-Voice, a large-scale bilingual benchmark for training and evaluating Chinese Social Values in LLMs. Building upon C-Voice, we introduce a multi-dimensional alignment method that integrates semantic clustering with Direct Preference Optimization (DPO) to enhance balanced values alignment. Extensive experiments on three representative LLMs demonstrate the effectiveness of our framework, offering both a practical solution and a theoretical foundation for values alignment. We will release the benchmark and code later.