<p>The rapid development of large language models (LLMs) has significantly improved the quality and diversity of AI-generated content(AIGC). LLM-Generated text detection plays an important role in preventing the harmful misuse of large language models. Existing approaches primarily analyze texts individually, overlooking the structural relationships between them. This limitation restricts their ability to generalize across diverse LLMs, as they fail to capture the shared statistical patterns inherent in generated texts. To address this, an unsupervised-based structural information for LLM-generated text detection (SILTD) method is proposed. The key insight is that texts from different LLMs exhibit latent similarities in their generative statistical space, which can be modeled to improve cross-model generalization. First, we construct a multi-relational text graph based on the similarity of text features, which aims to model the intricate similarities and correlations between texts. Second, we propose a novel unsupervised graph clustering method. The multi-relational graph is transformed into an encoding tree, which is then optimized based on a two-dimensional structure entropy minimization algorithm to achieve hierarchical clustering of texts. Structural entropy minimization enables achieving high-quality clusters, by measuring the uncertainty of random walks within the graph. Finally, we introduce a new method that measures text similarity and computes the intensity of text aggregation within each cluster, to perform in-cluster label inference. Extensive experiments show that, compared to baseline methods, our approach is more effective and generalizable in detecting six popular LLMs across five datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SILTD: Structural Information for LLM-Generated Text Detection

  • Jing Yang,
  • Shi Wang,
  • Kangli Zi,
  • Yanshun Sun,
  • Yuwei Huang,
  • Tianyu Luo

摘要

The rapid development of large language models (LLMs) has significantly improved the quality and diversity of AI-generated content(AIGC). LLM-Generated text detection plays an important role in preventing the harmful misuse of large language models. Existing approaches primarily analyze texts individually, overlooking the structural relationships between them. This limitation restricts their ability to generalize across diverse LLMs, as they fail to capture the shared statistical patterns inherent in generated texts. To address this, an unsupervised-based structural information for LLM-generated text detection (SILTD) method is proposed. The key insight is that texts from different LLMs exhibit latent similarities in their generative statistical space, which can be modeled to improve cross-model generalization. First, we construct a multi-relational text graph based on the similarity of text features, which aims to model the intricate similarities and correlations between texts. Second, we propose a novel unsupervised graph clustering method. The multi-relational graph is transformed into an encoding tree, which is then optimized based on a two-dimensional structure entropy minimization algorithm to achieve hierarchical clustering of texts. Structural entropy minimization enables achieving high-quality clusters, by measuring the uncertainty of random walks within the graph. Finally, we introduce a new method that measures text similarity and computes the intensity of text aggregation within each cluster, to perform in-cluster label inference. Extensive experiments show that, compared to baseline methods, our approach is more effective and generalizable in detecting six popular LLMs across five datasets.