Private Information Leakage in LLMs
摘要
Large Language Models (LLMs) can memorize training data and, if specifically prompted, reproduce or leak information on their training data. Information leakage has been observed for all types of machine-learning models. However, this threat exists at a much larger scale for LLMs because of their various applications as generative AI. This chapter relates the threat of information leakage with other adversarial threats, provides an overview of the current state of research on the mechanisms involved in memorization in LLMs, and discusses adversarial attacks aiming to extract memorized information from LLMs.