错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CoURAGE: A Framework to Evaluate RAG Systems

  • Divyanshi Galla,
  • Shaz Hoda,
  • Meiwei Zhang,
  • Wenzhe Quan,
  • Tommy Dong Yang,
  • Joseph Voyles

摘要

In the rapidly evolving domain of Generative AI(GenAI), evaluating models’ effectiveness for a business use case remains a significant challenge, particularly due to the diverse array of available metrics, the absence of a standardized framework for their application and varied challenges in use cases. This paper proposes a structured framework de signed to assist practitioners, including new adopters, in selecting appropriate metrics for the evaluation of GenAI models, specifically within question answering (QA) systems. The framework focuses on considerations such as data availability, the nature of the dataset, and the necessity for Large Language Models (LLMs) calls for evaluation. By categorizing metrics into quantitative and qualitative types, and distinguishing between scenarios that require golden labels, this framework seeks to streamline the evaluation process.