Semi-supervised learning enhances training efficiency by incorporating unlabeled data into the learning process, thereby reducing dependence on costly manual annotations. However, teacher models trained on limited labeled samples often generate pseudo-labels that deviate from the ground truth, introducing noise that amplifies deterministic bias and hinders model performance. To address this challenge, we reformulate label prediction as a progressive refinement process starting from an initial random guess, and propose LDiT (Label Diffusion Transformer) for pseudo-label noise adaptation. By modeling label uncertainty through a diffusion process, LDiT enables more robust learning under noisy supervision. In addition, to effectively model long-range interactions in textual data, we adopt a Transformer-based latent denoising architecture with self-attention mechanisms. Empirical results across three standard semi-supervised text classification datasets underscore the performance and adaptability of LDiT in varied scenarios.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LDiT: Pseudo-label Noise Adaptation via Label Diffusion Transformer

  • Jiawang Peng,
  • Pengfei Duan,
  • Mingru Huang,
  • Shengwu Xiong

摘要

Semi-supervised learning enhances training efficiency by incorporating unlabeled data into the learning process, thereby reducing dependence on costly manual annotations. However, teacher models trained on limited labeled samples often generate pseudo-labels that deviate from the ground truth, introducing noise that amplifies deterministic bias and hinders model performance. To address this challenge, we reformulate label prediction as a progressive refinement process starting from an initial random guess, and propose LDiT (Label Diffusion Transformer) for pseudo-label noise adaptation. By modeling label uncertainty through a diffusion process, LDiT enables more robust learning under noisy supervision. In addition, to effectively model long-range interactions in textual data, we adopt a Transformer-based latent denoising architecture with self-attention mechanisms. Empirical results across three standard semi-supervised text classification datasets underscore the performance and adaptability of LDiT in varied scenarios.