Remote photoplethysmography (rPPG) is a promising technology that consists of contactless measuring of cardiac activity from facial videos. However, current approaches are limited by data scarcity and environmental noise robustness. Most recent approaches utilize convolutional networks with limited temporal modeling capability or ignore long temporal context. Purely supervised rPPG methods are also severely limited by scarce data availability. In this work, we propose PhySU-Net, the first long temporal context rPPG transformer network and a novel self-supervised pre-training strategy that exploits unlabeled data to improve our model. Our strategy leverages traditional methods and image masking to provide pseudo-labels for physiologically relevant self-supervised pre-training. Our model is tested on three public benchmark datasets (OBF, VIPL-HR and MMSE-HR) and shows state-of-the-art performance in supervised training. Furthermore, we demonstrate that our self-supervised pre-training strategy further improves our model’s performance by leveraging representations learned from unlabeled data. Our code is available at: https://github.com/marukosan93/PhySU-Net .

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

PhySU-Net: Long Temporal Context Transformer for rPPG with Self-supervised Pre-training

  • Marko Savic,
  • Guoying Zhao

摘要

Remote photoplethysmography (rPPG) is a promising technology that consists of contactless measuring of cardiac activity from facial videos. However, current approaches are limited by data scarcity and environmental noise robustness. Most recent approaches utilize convolutional networks with limited temporal modeling capability or ignore long temporal context. Purely supervised rPPG methods are also severely limited by scarce data availability. In this work, we propose PhySU-Net, the first long temporal context rPPG transformer network and a novel self-supervised pre-training strategy that exploits unlabeled data to improve our model. Our strategy leverages traditional methods and image masking to provide pseudo-labels for physiologically relevant self-supervised pre-training. Our model is tested on three public benchmark datasets (OBF, VIPL-HR and MMSE-HR) and shows state-of-the-art performance in supervised training. Furthermore, we demonstrate that our self-supervised pre-training strategy further improves our model’s performance by leveraging representations learned from unlabeled data. Our code is available at: https://github.com/marukosan93/PhySU-Net .