Recent advancements in Natural Language Processing (NLP) have significantly enhanced the capabilities of large language models (LLMs) in generating complex texts such as narratives and dialogues. However, there is difficulty distinguishing between human-generated and machine-generated texts, coupled with the potential for malicious misuse (e.g., generating fake news and phishing emails). Text watermarking has emerged as a promising solution to mitigate these risks by embedding imperceptible, hidden information within texts that are detectable through specialized algorithms. Although there exist various methods devoted to embedding watermarks into texts generated by large language models, they face challenges in real-time environments where LLMs function as black boxes without direct decoding manipulation. To address these limitations, we introduce an innovative model-free watermarking technique that uses post-processing to embed watermarks without altering the decoding process. This technique leverages a comprehensive static table of n-gram synonym pairs, selected based on semantic similarity, to ensure the watermarks' traceability without compromising the semantic quality of the text. Preliminary results demonstrate that our approach significantly enhances the efficiency of watermarking in black-box model settings. It maintains fast processing speeds while ensuring the detectability and integrity of the watermarks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

N-gram Sliding Window Watermarking

  • Qifeng Su,
  • Zhicong Wu,
  • Xiaodong Shi

摘要

Recent advancements in Natural Language Processing (NLP) have significantly enhanced the capabilities of large language models (LLMs) in generating complex texts such as narratives and dialogues. However, there is difficulty distinguishing between human-generated and machine-generated texts, coupled with the potential for malicious misuse (e.g., generating fake news and phishing emails). Text watermarking has emerged as a promising solution to mitigate these risks by embedding imperceptible, hidden information within texts that are detectable through specialized algorithms. Although there exist various methods devoted to embedding watermarks into texts generated by large language models, they face challenges in real-time environments where LLMs function as black boxes without direct decoding manipulation. To address these limitations, we introduce an innovative model-free watermarking technique that uses post-processing to embed watermarks without altering the decoding process. This technique leverages a comprehensive static table of n-gram synonym pairs, selected based on semantic similarity, to ensure the watermarks' traceability without compromising the semantic quality of the text. Preliminary results demonstrate that our approach significantly enhances the efficiency of watermarking in black-box model settings. It maintains fast processing speeds while ensuring the detectability and integrity of the watermarks.