Research of Multidimensional Adversarial Examples in LLMs for Recognizing Ethics and Security Issues
摘要
With the extensive use of LLMs in research and practical applications, it has become more and more important to evaluate them effectively, and by studying the evaluation methods can help to better understand the LLMs, guard against unknowns, avoid risks, and provide a basis for their better and faster iterative upgrading. In this research, the ability of recognizing ethics and security issues in text is investigated through a multi-dimensional adversarial example evaluation method, using ERNIE Bot (V2.2.3) as an example. The ESIIP of ERNIE Bot (V2.2.3) is evaluated by slightly perturbing the input data through multidimensional adversarial examples to induce the model to make false predictions. In this research, the evaluation objectives are classified into ethics and security issues such as discrimination and prejudice detection, values analysis, and ethical conflict identification, and security issues such as false information detection, privacy violation detection, and network security detection. Multiple representative datasets, using different attack strategies to perturb them slightly, the research formulated a rigorous evaluation criterion developed for the model’s responses, and comprehensively analyzed the scores of all the LLMs; the research drew the corresponding metrics conclusions and complexity conclusions, and compared the performance with other LLMs models in recognizing the ethics and security issues in the text. The results show that ERNIE Bot (V2.2.3) performs well in ESIIP, not reaching the perfect level; it also shows the reliability and feasibility of the research method.