With the rapid development of large language models (LLMs), there is growing concern about their safety issues. LLMs may generate offensive and discriminatory content, reflect incorrect public interest, and be used for toxic purposes such as fraud and the dissemination of misleading information. Therefore, evaluating the safety of LLMs has become a necessary task to promote their widespread application. However, the safety evaluation for Chinese Human-AI Conversations of LLMs still encounters significant challenges. These include limited categories or questions in the question pool and the suboptimal quality of questions, which collectively result in less comprehensive safety assessments. Furthermore, the evaluation of replies generated by LLMs is hindered by the absence of automated and effective toxicity detection methods tailored to Chinese. Current practices predominantly rely on human judgment or sophisticated LLMs like GPT-4, which introduces inefficiencies and potential biases. To address such challenges, we first construct a question pool with sufficient categories and quantities, consisting of 4 major categories and a total of 29 subcategories. Furthermore, we construct an automated safety evaluation framework LLMSentinel based on a toxicity detection model for Chinese replies, further enhancing the safety evaluation of Chinese Human-AI Conversations of LLMs. Experiments conducted on seven models, including three open-source models and four closed-source models, revealed that each LLM performed poorly in at least one subcategory. Furthermore, all LLMs demonstrated poor performance in the categories of “Infringement of Others’ Right to Reputation” and “Spreading False and Harmful Information”, indicating that these models still have unresolved safety issues.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LLMSentinel: An Automatic and Efficient Toxicity Evaluation Scheme for Chinese Human-AI Conversations with Large Language Models

  • Kai Zhang,
  • Yonghui Huang,
  • Bo Wang,
  • Shujie Yang,
  • Yi Sun,
  • Yuchen Wang,
  • Chao Wang,
  • Zan Zhou

摘要

With the rapid development of large language models (LLMs), there is growing concern about their safety issues. LLMs may generate offensive and discriminatory content, reflect incorrect public interest, and be used for toxic purposes such as fraud and the dissemination of misleading information. Therefore, evaluating the safety of LLMs has become a necessary task to promote their widespread application. However, the safety evaluation for Chinese Human-AI Conversations of LLMs still encounters significant challenges. These include limited categories or questions in the question pool and the suboptimal quality of questions, which collectively result in less comprehensive safety assessments. Furthermore, the evaluation of replies generated by LLMs is hindered by the absence of automated and effective toxicity detection methods tailored to Chinese. Current practices predominantly rely on human judgment or sophisticated LLMs like GPT-4, which introduces inefficiencies and potential biases. To address such challenges, we first construct a question pool with sufficient categories and quantities, consisting of 4 major categories and a total of 29 subcategories. Furthermore, we construct an automated safety evaluation framework LLMSentinel based on a toxicity detection model for Chinese replies, further enhancing the safety evaluation of Chinese Human-AI Conversations of LLMs. Experiments conducted on seven models, including three open-source models and four closed-source models, revealed that each LLM performed poorly in at least one subcategory. Furthermore, all LLMs demonstrated poor performance in the categories of “Infringement of Others’ Right to Reputation” and “Spreading False and Harmful Information”, indicating that these models still have unresolved safety issues.