Event cameras, with their high temporal resolution, low latency, and immunity to motion blur, have become valuable in dynamic environments for tasks like facial expression and action recognition. However, deploying multimodal systems that incorporate both RGB and event data or other sensor combinations presents challenges, including synchronization requirements, increased power consumption, and the added complexity of maintaining multiple sensors. This paper introduces a novel approach to facial action unit detection, employing the Privileged Information (PI) paradigm to improve single-modality models. By training separate event-based and RGB classifiers to “hallucinate” features from an absent modality, our method emulates a multimodal system even under single-modality constraints. Experiments on the FACEMORPHIC dataset show that our PI-based framework substantially outperforms single-modality baselines, achieving performance close to that of fully multimodal models. This work establishes a basis for adaptable and resilient event-camera-based systems that function effectively without consistent multimodal input, making them well-suited for deployment in complex, resource-constrained environments.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Feature Hallucination via Privileged Information for Neuromorphic Face Analysis

  • Lorenzo Berlincioni,
  • Luca Cultrera,
  • Gabriele Magrini,
  • Federico Becattini,
  • Pietro Pala,
  • Alberto Del Bimbo

摘要

Event cameras, with their high temporal resolution, low latency, and immunity to motion blur, have become valuable in dynamic environments for tasks like facial expression and action recognition. However, deploying multimodal systems that incorporate both RGB and event data or other sensor combinations presents challenges, including synchronization requirements, increased power consumption, and the added complexity of maintaining multiple sensors. This paper introduces a novel approach to facial action unit detection, employing the Privileged Information (PI) paradigm to improve single-modality models. By training separate event-based and RGB classifiers to “hallucinate” features from an absent modality, our method emulates a multimodal system even under single-modality constraints. Experiments on the FACEMORPHIC dataset show that our PI-based framework substantially outperforms single-modality baselines, achieving performance close to that of fully multimodal models. This work establishes a basis for adaptable and resilient event-camera-based systems that function effectively without consistent multimodal input, making them well-suited for deployment in complex, resource-constrained environments.