The large language models (LLMs) have shown their superiority in solving various natural language processing tasks. However, owing to the lack of knowledge of specific domains, LLMs often perform poorly for applications of expert systems. In this paper, we introduce a new prompt engineering technique, called Reinforcement Chain of Thought (R-CoT), which incorporates the critical feedback and interactive steps to improve the Chain of Thought (CoT). By leveraging the automation advantages of large language models (LLMs), this technique reduces the need for extensive manual data annotation, and can be effectively applied in expert systems for specific domains. We have conducted experiments for the GSM8K math problem-solving dataset, the CSQA commonsense understanding dataset, and tasks from the Super- NaturalInstructions V2 benchmark using ChatGPT. Experimental results show that the accuracy of R-CoT achieves an accuracy of 80.7% on the CSQA dataset and 45.6% on the GSM8K dataset, which outperforms related methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

R-CoT: Reinforcement Chain of Thought Prompting for Task Specific Training

  • Hsu-Chih Chiu,
  • Iuan-Kai Fang,
  • Che-Rung Lee

摘要

The large language models (LLMs) have shown their superiority in solving various natural language processing tasks. However, owing to the lack of knowledge of specific domains, LLMs often perform poorly for applications of expert systems. In this paper, we introduce a new prompt engineering technique, called Reinforcement Chain of Thought (R-CoT), which incorporates the critical feedback and interactive steps to improve the Chain of Thought (CoT). By leveraging the automation advantages of large language models (LLMs), this technique reduces the need for extensive manual data annotation, and can be effectively applied in expert systems for specific domains. We have conducted experiments for the GSM8K math problem-solving dataset, the CSQA commonsense understanding dataset, and tasks from the Super- NaturalInstructions V2 benchmark using ChatGPT. Experimental results show that the accuracy of R-CoT achieves an accuracy of 80.7% on the CSQA dataset and 45.6% on the GSM8K dataset, which outperforms related methods.