Existing supervised facial attribute recognition (FAR) methods that rely on large labeled datasets can pose a challenge in real-world scenarios. In the case of limited labeled data, the current methods that introduce auxiliary tasks with a large number of parameters are not conducive to the embedded applications of FAR. To overcome these challenges, this paper develops an adjustable gating prompt Transformer that can handle the limited labeled FAR task with a small number of training parameters. Specifically, we employ an effective image-guided prompt tuning, where the image-related prompt sequence is first generated by feeding image tokens into an image-guided prompt generation network (IPG-Net). Then, the prompt sequence can learn facial image information and guide the frozen pre-trained Transformer to fine-tune the model. In addition, dynamically adjustable gating is applied to the prompt sequence to adaptively adjust the contribution of the prompts from different encoder layers, which enhances the interaction between the different encoder layers and retains effective feature information during the iterative process. Experimental results on the CelebA and LFWA datasets demonstrate that our method outperforms competitive methods with a very small amount of training parameters when only limited labeled data are used.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adjustable Gating Prompt Transformer for Facial Attribute Recognition with Limited Labeled Data

  • Qinxian Ye,
  • Si Chen,
  • Da-Han Wang,
  • Nanfeng Jiang,
  • Yanfei Su,
  • Yan Yan

摘要

Existing supervised facial attribute recognition (FAR) methods that rely on large labeled datasets can pose a challenge in real-world scenarios. In the case of limited labeled data, the current methods that introduce auxiliary tasks with a large number of parameters are not conducive to the embedded applications of FAR. To overcome these challenges, this paper develops an adjustable gating prompt Transformer that can handle the limited labeled FAR task with a small number of training parameters. Specifically, we employ an effective image-guided prompt tuning, where the image-related prompt sequence is first generated by feeding image tokens into an image-guided prompt generation network (IPG-Net). Then, the prompt sequence can learn facial image information and guide the frozen pre-trained Transformer to fine-tune the model. In addition, dynamically adjustable gating is applied to the prompt sequence to adaptively adjust the contribution of the prompts from different encoder layers, which enhances the interaction between the different encoder layers and retains effective feature information during the iterative process. Experimental results on the CelebA and LFWA datasets demonstrate that our method outperforms competitive methods with a very small amount of training parameters when only limited labeled data are used.