<p>Recently, retrieval-augmented generation (RAG) systems have attracted attention for addressing issues like hallucinations and reliance on outdated knowledge in large language models (LLMs). Privacy studies have revealed that RAG systems are vulnerable to membership leakage in determining whether a specific target sample is included in the RAG knowledge base. Existing membership inference attack (MIA) methods for RAGs primarily rely on similarity scores between the system’s responses and the true answers. These methods assume that a higher similarity score indicates the sample is more likely to have been used by the RAG system to enhance its response, suggesting it is a member of the knowledge base. However, this study uncovers an important insight: the similarity metric does not directly represent the membership status, instead measures the response difficulty of the sample. To address this, we propose a novel membership inference attack for RAG systems, called difficulty-calibrated membership inference attack (DC-MIA). It first classifies high-similarity samples as members, and then calibrates the membership scores of samples with comparable raw similarity scores using a likelihood ratio test. Experimental results demonstrate that our approach significantly improves the performance of membership inference attacks on RAG systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

RAG-leaks: difficulty-calibrated membership inference attacks on retrieval-augmented generation

  • Guangshuo Wang,
  • Jiajun He,
  • Hao Li,
  • Min Zhang,
  • Dengguo Feng

摘要

Recently, retrieval-augmented generation (RAG) systems have attracted attention for addressing issues like hallucinations and reliance on outdated knowledge in large language models (LLMs). Privacy studies have revealed that RAG systems are vulnerable to membership leakage in determining whether a specific target sample is included in the RAG knowledge base. Existing membership inference attack (MIA) methods for RAGs primarily rely on similarity scores between the system’s responses and the true answers. These methods assume that a higher similarity score indicates the sample is more likely to have been used by the RAG system to enhance its response, suggesting it is a member of the knowledge base. However, this study uncovers an important insight: the similarity metric does not directly represent the membership status, instead measures the response difficulty of the sample. To address this, we propose a novel membership inference attack for RAG systems, called difficulty-calibrated membership inference attack (DC-MIA). It first classifies high-similarity samples as members, and then calibrates the membership scores of samples with comparable raw similarity scores using a likelihood ratio test. Experimental results demonstrate that our approach significantly improves the performance of membership inference attacks on RAG systems.