Large Language Models (LLMs) have shown remarkable performance on a variety of natural language tasks, but eliciting their abilities for more consistent predictions is still highly reliant on researchers’ well-crafted prompts, which often require costly computational consumption as well as onerous trial-and-error efforts. For generative tasks such as open-domain question answering (QA), LLMs face a particularly great challenge. To alleviate the issue in QA tasks, we propose a scheme that can both automatically optimize and efficiently leverage candidate prompts to obtain consistent and accurate predicted answers. Superior self-consistency of LLMs’ predictions requires diverse high-quality reasoning paths, thus multiple candidate prompts in the iterative optimization process of automatic prompt engineering (APE) are fully leveraged to drive LLMs to generate multiple reasoning paths leading to improved self-consistency. The evaluation performance of the candidate prompt on the sampled dataset is used as a reference (similar to the weights of the base learners in ensemble learning) to determine its contribution to the final answer. Experimental results demonstrate that our scheme yields more reliable and diverse predicted answers and outperforms conventional self-consistency baseline models in several typical QA benchmark tests.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving Self-consistency for Open-Domain Question Answering via Automatic Prompt Engineering and Ensemble Learning

  • Jie Liu,
  • Xue Han,
  • Chao Deng,
  • Junlan Feng

摘要

Large Language Models (LLMs) have shown remarkable performance on a variety of natural language tasks, but eliciting their abilities for more consistent predictions is still highly reliant on researchers’ well-crafted prompts, which often require costly computational consumption as well as onerous trial-and-error efforts. For generative tasks such as open-domain question answering (QA), LLMs face a particularly great challenge. To alleviate the issue in QA tasks, we propose a scheme that can both automatically optimize and efficiently leverage candidate prompts to obtain consistent and accurate predicted answers. Superior self-consistency of LLMs’ predictions requires diverse high-quality reasoning paths, thus multiple candidate prompts in the iterative optimization process of automatic prompt engineering (APE) are fully leveraged to drive LLMs to generate multiple reasoning paths leading to improved self-consistency. The evaluation performance of the candidate prompt on the sampled dataset is used as a reference (similar to the weights of the base learners in ensemble learning) to determine its contribution to the final answer. Experimental results demonstrate that our scheme yields more reliable and diverse predicted answers and outperforms conventional self-consistency baseline models in several typical QA benchmark tests.