Large language models (LLMs) have shown remarkable capabilities in various natural language processing tasks, leading to their widespread deployment as intelligent assistants in decision-making contexts. However, the increasing complexity of these models raises concerns about their reliability, particularly regarding bias and hallucination. In this work, we evaluate the robustness of LLMs to perturbed variations of the original inquiry in decision-making tasks. We show that contrary to previous studies, perturbations can mitigate bias and hallucination in some LLMs over other models. It’s found that Claude 3 is more effective for the tasks represented in most datasets, whereas models like GPT3.5 exhibit varying levels of adequacy, performing comparably in some cases but falling significantly behind in others. These insights are crucial for understanding the practical implications of deploying LLM-based assistants as effective decision-support tools in real-world applications, emphasising the need for rigorous testing and validation to ensure reliability and effectiveness. This study contributes to the growing body of research on LLM evaluation and provides insights for developing more robust and trustworthy AI assistants in critical decision-making contexts.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating the Effects of Prompt Perturbation on Bias and Hallucination in Large Language Models

  • Mamehgol Yousefi,
  • Ahmad Shahi,
  • Mos Sharifi,
  • Alvaro Romera,
  • Simon Hoermann,
  • Tham Piumsomboon

摘要

Large language models (LLMs) have shown remarkable capabilities in various natural language processing tasks, leading to their widespread deployment as intelligent assistants in decision-making contexts. However, the increasing complexity of these models raises concerns about their reliability, particularly regarding bias and hallucination. In this work, we evaluate the robustness of LLMs to perturbed variations of the original inquiry in decision-making tasks. We show that contrary to previous studies, perturbations can mitigate bias and hallucination in some LLMs over other models. It’s found that Claude 3 is more effective for the tasks represented in most datasets, whereas models like GPT3.5 exhibit varying levels of adequacy, performing comparably in some cases but falling significantly behind in others. These insights are crucial for understanding the practical implications of deploying LLM-based assistants as effective decision-support tools in real-world applications, emphasising the need for rigorous testing and validation to ensure reliability and effectiveness. This study contributes to the growing body of research on LLM evaluation and provides insights for developing more robust and trustworthy AI assistants in critical decision-making contexts.