Large language models (LLMs) have significantly enhanced human productivity but also created opportunities for malicious actors to disseminate fake news for influence or steal generated text for profit. Watermark offers a viable solution to address these challenges. Current LLM watermark research predominantly focuses on detecting generated text (i.e., zero-bit watermark). However, tracing model users to address the root cause requires embedding the multi-bit watermark into generated text. In this work, we propose a novel multi-bit watermark based on robust encoding and green-zone refinement. We propose a robust message encoding mechanism, while employing red-green lists and green-zone refinement to prepare for watermark detection and message extraction. Crucially, our method treats all original message bits, checksum, and error-correcting codes uniformly during the process of watermark embedding, allocating an equal number of tokens to each bit to ensure equal embedding opportunities. This balanced distribution reduces bias and optimizes token utilization across the generated text, thereby enhancing the accuracy and robustness of the watermark. Additionally, our extraction process requires no access to the original model, significantly improving efficiency. Experimental results demonstrate that our method effectively resolves the trade-off between accuracy and extraction efficiency in the multi-bit watermark for LLMs, while exhibiting strong robustness in adversarial attack scenarios.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Reliable Multi-bit Watermark for Large Language Models via Robust Encoding and Green-Zone Refinement

  • Wenqin Jin,
  • Weihai Li

摘要

Large language models (LLMs) have significantly enhanced human productivity but also created opportunities for malicious actors to disseminate fake news for influence or steal generated text for profit. Watermark offers a viable solution to address these challenges. Current LLM watermark research predominantly focuses on detecting generated text (i.e., zero-bit watermark). However, tracing model users to address the root cause requires embedding the multi-bit watermark into generated text. In this work, we propose a novel multi-bit watermark based on robust encoding and green-zone refinement. We propose a robust message encoding mechanism, while employing red-green lists and green-zone refinement to prepare for watermark detection and message extraction. Crucially, our method treats all original message bits, checksum, and error-correcting codes uniformly during the process of watermark embedding, allocating an equal number of tokens to each bit to ensure equal embedding opportunities. This balanced distribution reduces bias and optimizes token utilization across the generated text, thereby enhancing the accuracy and robustness of the watermark. Additionally, our extraction process requires no access to the original model, significantly improving efficiency. Experimental results demonstrate that our method effectively resolves the trade-off between accuracy and extraction efficiency in the multi-bit watermark for LLMs, while exhibiting strong robustness in adversarial attack scenarios.