<p>Chain-of-thought (CoT) prompting, as a simple yet effective reasoning enhancement strategy, has significantly improved the performance of large language models (LLMs) on complex reasoning tasks. In particular, zero-shot CoT can substantially enhance LLMs’ reasoning abilities across various domains using only a concise prompt template (e.g., “Let’s think step by step”). However, existing zero-shot CoT methods still face two key challenges: First, their reasoning capabilities exhibit significant limitations in multilingual generalization, making them less adaptable to multilingual task scenarios and constraining their feasibility in global applications. Second, although zero-shot CoT does not require human annotations, it lacks task-specific guidance signals, often leading to unstable reasoning quality. Moreover, manually designing optimized prompt templates is both costly and difficult to scale across diverse task scenarios. To address these issues, this paper proposes RSO-CoT, a cross-lingual reasoning enhancement framework that integrates rectification mechanism and self-optimization strategy to improve the multilingual reasoning capabilities of zero-shot CoT. RSO-CoT consists of three key modules: (1) aligned reasoning prompting, which guides LLMs to generate initially aligned reasoning paths and answers in multilingual settings. (2) Verification and rectification, which evaluates reasoning results using alternative verification methods and dynamically corrects errors based on feedback. (3) Self-optimization prompting, which filters high-quality examples through self-consistency and dynamically constructs optimal prompt templates to generate the final answer. Experimental results indicate that RSO-CoT performs well on multiple cross-lingual benchmarks, enhancing LLMs’ multilingual reasoning capabilities compared to existing zero-shot CoT methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving zero-shot chain-of-thought reasoning across languages with rectification and self-optimization prompting

  • Hongwei Chen,
  • Jiajun Wang,
  • Wei Wang,
  • Yang Xu

摘要

Chain-of-thought (CoT) prompting, as a simple yet effective reasoning enhancement strategy, has significantly improved the performance of large language models (LLMs) on complex reasoning tasks. In particular, zero-shot CoT can substantially enhance LLMs’ reasoning abilities across various domains using only a concise prompt template (e.g., “Let’s think step by step”). However, existing zero-shot CoT methods still face two key challenges: First, their reasoning capabilities exhibit significant limitations in multilingual generalization, making them less adaptable to multilingual task scenarios and constraining their feasibility in global applications. Second, although zero-shot CoT does not require human annotations, it lacks task-specific guidance signals, often leading to unstable reasoning quality. Moreover, manually designing optimized prompt templates is both costly and difficult to scale across diverse task scenarios. To address these issues, this paper proposes RSO-CoT, a cross-lingual reasoning enhancement framework that integrates rectification mechanism and self-optimization strategy to improve the multilingual reasoning capabilities of zero-shot CoT. RSO-CoT consists of three key modules: (1) aligned reasoning prompting, which guides LLMs to generate initially aligned reasoning paths and answers in multilingual settings. (2) Verification and rectification, which evaluates reasoning results using alternative verification methods and dynamically corrects errors based on feedback. (3) Self-optimization prompting, which filters high-quality examples through self-consistency and dynamically constructs optimal prompt templates to generate the final answer. Experimental results indicate that RSO-CoT performs well on multiple cross-lingual benchmarks, enhancing LLMs’ multilingual reasoning capabilities compared to existing zero-shot CoT methods.