错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CC-Eval: Benchmarking Cross-Lingual Value Alignment and Chinese Context Tasks in Large Language Models

  • Qiushi Dong,
  • Jiayin Qi

摘要

Existing evaluation frameworks do not adequately measure whether a single large language model exhibits consistent value alignment across Chinese and English contexts. We introduce CC-Eval, a benchmark that combines a bilingual parallel value-alignment subset with a Chinese-context task subset spanning six Chinese cultural domains. CC-Eval contains 2,217 instances, including 1,020 Chinese–English prompt pairs and 1,197 Chinese-context instances, and we evaluate eight representative models with open-ended generation under a unified LLM-as-a-Judge protocol. The benchmark yields two main empirical findings: models more often receive Western-aligned judgments in English than Chinese-aligned judgments in Chinese, and performance on Chinese-context tasks varies substantially across domains. Modern Chinese internet slang remains the most challenging domain for most models in our evaluation, highlighting a persistent gap in culturally grounded Chinese-context understanding.