This paper explores the implications for the problem-solving performance of large language models (LLM) when utilizing augmented intelligence by infusing human reasoning into the generated intermediate steps before a model outputs a result. We propose a framework that includes steps for injecting human reasoning and feedback into the prompting steps of large language models, using three different categories of human feedback, namely Substitution, N-Shot Learning, and Conversational, aiming to improve the model’s problem-solving ability. In order to test the framework, we conducted a user study with participants who edited the intermediate steps to align with their reasoning. The results of revised prompts are compared against their base performance on the ARC, DROP, and WinoGrande datasets. Our findings reveal that the injection of human reasoning steps can boost a large language model’s problem-solving accuracy regarding the test benchmarks, particularly regarding the N-Shot Learning approach. We aim to shed light on the different applications of infusing human feedback into the output of pre-trained models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Infusing Human Feedback into Intermediate Prompting Steps of Large Language Models

  • Wangfan Li,
  • Claire Gendron,
  • Carlos Toxtli

摘要

This paper explores the implications for the problem-solving performance of large language models (LLM) when utilizing augmented intelligence by infusing human reasoning into the generated intermediate steps before a model outputs a result. We propose a framework that includes steps for injecting human reasoning and feedback into the prompting steps of large language models, using three different categories of human feedback, namely Substitution, N-Shot Learning, and Conversational, aiming to improve the model’s problem-solving ability. In order to test the framework, we conducted a user study with participants who edited the intermediate steps to align with their reasoning. The results of revised prompts are compared against their base performance on the ARC, DROP, and WinoGrande datasets. Our findings reveal that the injection of human reasoning steps can boost a large language model’s problem-solving accuracy regarding the test benchmarks, particularly regarding the N-Shot Learning approach. We aim to shed light on the different applications of infusing human feedback into the output of pre-trained models.