LLMSentinel: An Automatic and Efficient Toxicity Evaluation Scheme for Chinese Human-AI Conversations with Large Language Models
摘要
With the rapid development of large language models (LLMs), there is growing concern about their safety issues. LLMs may generate offensive and discriminatory content, reflect incorrect public interest, and be used for toxic purposes such as fraud and the dissemination of misleading information. Therefore, evaluating the safety of LLMs has become a necessary task to promote their widespread application. However, the safety evaluation for Chinese Human-AI Conversations of LLMs still encounters significant challenges. These include limited categories or questions in the question pool and the suboptimal quality of questions, which collectively result in less comprehensive safety assessments. Furthermore, the evaluation of replies generated by LLMs is hindered by the absence of automated and effective toxicity detection methods tailored to Chinese. Current practices predominantly rely on human judgment or sophisticated LLMs like GPT-4, which introduces inefficiencies and potential biases. To address such challenges, we first construct a question pool with sufficient categories and quantities, consisting of 4 major categories and a total of 29 subcategories. Furthermore, we construct an automated safety evaluation framework LLMSentinel based on a toxicity detection model for Chinese replies, further enhancing the safety evaluation of Chinese Human-AI Conversations of LLMs. Experiments conducted on seven models, including three open-source models and four closed-source models, revealed that each LLM performed poorly in at least one subcategory. Furthermore, all LLMs demonstrated poor performance in the categories of “Infringement of Others’ Right to Reputation” and “Spreading False and Harmful Information”, indicating that these models still have unresolved safety issues.