As administrative language tends to be formal and exempt from double meanings or figurative expressions, it is a particular domain in which to explore the performance of Language Models. This paper presents a study on the feasibility of creating administrative texts-based RAG systems to serve as chatbots, analyzing the performance for this task of several Small and Large Language Models and defining ways of evaluating whether they hallucinate or not and whether they provide the user useful information or not. Conventional metrics depending on ground truth labels, such as cosine similarity or those from the ROUGE family, are explored, as well as new approaches to using other metrics not so popular in text evaluation, such as Euclidean and Manhattan distances. Moreover, all those objective metrics are compared with a subjective Likert scale to assess their performance at solving real users’ problems and to find relations between subjective perceptions and objectively measured metrics for each of the RAG systems proposed. The results show that SLM models (such as NeuralChat) can perform as well as an LLM if RAG programming provides them with an appropriate context.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating Performance and Trustworthiness of RAG Systems for Generating Administrative Text

  • Hugo Sánchez-Navalón,
  • Carlos Monserrat,
  • Dario Garigliotti,
  • Cèsar Ferri

摘要

As administrative language tends to be formal and exempt from double meanings or figurative expressions, it is a particular domain in which to explore the performance of Language Models. This paper presents a study on the feasibility of creating administrative texts-based RAG systems to serve as chatbots, analyzing the performance for this task of several Small and Large Language Models and defining ways of evaluating whether they hallucinate or not and whether they provide the user useful information or not. Conventional metrics depending on ground truth labels, such as cosine similarity or those from the ROUGE family, are explored, as well as new approaches to using other metrics not so popular in text evaluation, such as Euclidean and Manhattan distances. Moreover, all those objective metrics are compared with a subjective Likert scale to assess their performance at solving real users’ problems and to find relations between subjective perceptions and objectively measured metrics for each of the RAG systems proposed. The results show that SLM models (such as NeuralChat) can perform as well as an LLM if RAG programming provides them with an appropriate context.