Trustworthiness Evaluation of Large Language Models
摘要
Large language models (LLMs) have gained considerable attention for their excellent natural language processing capabilities. Nonetheless, these LLMs present many challenges, particularly in the realm of trustworthiness. Therefore, ensuring the trustworthiness of LLMs emerges as an important topic. This chapter presents the TrustLLM framework (Sun et al., Trustllm: Trustworthiness in large language models. International Conference on Machine Learning (2024)), a comprehensive study of trustworthiness in LLMs, including principles for different dimensions of trustworthiness, established benchmark, evaluation, and analysis of trustworthiness for mainstream LLMs. Specifically, we first introduce a set of principles for trustworthy LLMs that span eight dimensions. Based on these principles, we further establish a benchmark across six dimensions including truthfulness, safety, fairness, robustness, privacy, and machine ethics. Based on the evaluation of 16 mainstream LLMs in TrustLLM (Sun et al., Trustllm: Trustworthiness in large language models. International Conference on Machine Learning (2024)), consisting of over 30 datasets, this chapter summarizes the main findings.