Assessing Maryland Watermark Robustness Against Paraphrasing Attacks
摘要
The rise of large language models (LLMs) and the potential for misuse necessitate the development of reliable methods to identify and authenticate machine-generated content. The Maryland Watermark, proposed by Kirchenbauer et al. (Robust distortion-free watermarks for language models, [1]), is a notable technique that embeds identifiable signatures into text generated by LLMs. This paper investigates the robustness of the Maryland Watermark, particularly its vulnerability to watermark removal through paraphrasing attacks. I introduce a novel, low-cost watermark removal technique and assess the impact of various paraphrasing strategies on detection accuracy. Our findings reveal that the detection accuracy of watermarked documents drops dramatically from 100% to 1.6% following paraphrasing. Additionally, I analyze the cost and effectiveness of different paraphrasing approaches, concluding that while recursive paraphrasing is unreliable, sentence-level paraphrasing is a feasible method for watermark removal in academic contexts. The study underscores the need for more robust watermarking methods to withstand sophisticated removal techniques.