Facial action unit (AU) recognition is a challenging task, due to the subtlety of each AU and the correlations among AUs in global face. However, the learning of local-global features has not been thoroughly exploited in most of the existing methods. In this paper, we propose a novel micro-action-aware transformer to integrate local and global feature extractions, which effectively captures subtle AU details while maintaining the global relational modeling capacity of transformers. Besides, we jointly train facial AU recognition and facial landmark detection, in which the two correlated tasks contribute to each other and further facilitate the learning of local-global AU-related feature. Extensive experiments demonstrate that our approach achieves comparable performance to the state-of-the-art AU recognition methods on the challenging BP4D and GFT benchmarks, and works well for landmark detection. Particularly, our approach achieves average F1 score results of 63.3% and 55.8% on BP4D and GFT datasets, respectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Facial Action Unit Recognition with Micro-Action-Aware Transformer

  • Yichen Yuan,
  • Yifan Cheng,
  • Zhiwen Shao,
  • Qianwen Dang,
  • Rui Chen,
  • Mingjian Fu,
  • Shengtian Jiang,
  • Chunyu Li,
  • Lizhuang Ma

摘要

Facial action unit (AU) recognition is a challenging task, due to the subtlety of each AU and the correlations among AUs in global face. However, the learning of local-global features has not been thoroughly exploited in most of the existing methods. In this paper, we propose a novel micro-action-aware transformer to integrate local and global feature extractions, which effectively captures subtle AU details while maintaining the global relational modeling capacity of transformers. Besides, we jointly train facial AU recognition and facial landmark detection, in which the two correlated tasks contribute to each other and further facilitate the learning of local-global AU-related feature. Extensive experiments demonstrate that our approach achieves comparable performance to the state-of-the-art AU recognition methods on the challenging BP4D and GFT benchmarks, and works well for landmark detection. Particularly, our approach achieves average F1 score results of 63.3% and 55.8% on BP4D and GFT datasets, respectively.