Since the advent of ChatGPT, the popularity of large language models has made distinguishing between model-generated and human-created text a significant challenge. Embedding watermarks during text generation is a crucial method for identifying model-generated content. However, existing approaches predominantly focus on 0-bit watermarks, limiting their capacity to embed more substantive information like authority information, version information and so on. The development of n-bit watermarks remains in its nascent stage. Most n-bit watermarking methods have constraints in text quality and robustness, often struggling to achieve high accuracy over limited text sequences. Hence, we propose an n-bit watermark(NLWM) based on language model proxy. To address the issue of reduction in text quality observed in current methods, we partition the model’s vocabulary using a fixed random seed and introduce a novel strategy to reweight the original probability distribution. Simultaneously, in order to enhance watermark robustness and extraction accuracy, we incorporate a BCH error-correcting code mechanism to our method. Since NLWM does not need to access the language model when extracting watermarks, it has high extraction efficiency. In open text generation experiments, we compared our method with baseline models. The results demonstrate that NLWM not only enhances watermark robustness and text quality but also improves the efficiency and accuracy of watermark extraction, highlighting the practical value of this method.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

NLWM: A Robust, Efficient and High-Quality Watermark for Large Language Models

  • Mengting Song,
  • Ziyuan Li,
  • Kai Liu,
  • Min Peng,
  • Gang Tian

摘要

Since the advent of ChatGPT, the popularity of large language models has made distinguishing between model-generated and human-created text a significant challenge. Embedding watermarks during text generation is a crucial method for identifying model-generated content. However, existing approaches predominantly focus on 0-bit watermarks, limiting their capacity to embed more substantive information like authority information, version information and so on. The development of n-bit watermarks remains in its nascent stage. Most n-bit watermarking methods have constraints in text quality and robustness, often struggling to achieve high accuracy over limited text sequences. Hence, we propose an n-bit watermark(NLWM) based on language model proxy. To address the issue of reduction in text quality observed in current methods, we partition the model’s vocabulary using a fixed random seed and introduce a novel strategy to reweight the original probability distribution. Simultaneously, in order to enhance watermark robustness and extraction accuracy, we incorporate a BCH error-correcting code mechanism to our method. Since NLWM does not need to access the language model when extracting watermarks, it has high extraction efficiency. In open text generation experiments, we compared our method with baseline models. The results demonstrate that NLWM not only enhances watermark robustness and text quality but also improves the efficiency and accuracy of watermark extraction, highlighting the practical value of this method.