NeRFs have shown remarkable results in modeling the 4D dynamics and appearance of human faces. However, they require per-identity optimization. A crucial step towards building foundation models for humans would be to learn a unified representation for multiple subjects. In this work, we introduce MI-NeRF (multi-identity NeRF), a single network that models complex non-rigid facial motion for multiple identities, using only monocular videos. The core premise in our method is to learn the non-linear interactions between identity and non-identity specific information with a multiplicative module. We present an extensive study of different variants of our proposed module and their technical derivations. We demonstrate results for both facial expression transfer and talking face video synthesis. By training on multiple videos simultaneously, MI-NeRF not only reduces the total training time compared to standard single-identity NeRFs, but also demonstrates robustness in synthesizing novel expressions for any input identity. Our method can be further personalized for a target identity given only a short video. Project page: https://aggelinacha.github.io/MI-NeRF/ .

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MI-NeRF: Learning a Single NeRF for Multiple Identities

  • Aggelina Chatziagapi,
  • Grigorios G. Chrysos,
  • Dimitris Samaras

摘要

NeRFs have shown remarkable results in modeling the 4D dynamics and appearance of human faces. However, they require per-identity optimization. A crucial step towards building foundation models for humans would be to learn a unified representation for multiple subjects. In this work, we introduce MI-NeRF (multi-identity NeRF), a single network that models complex non-rigid facial motion for multiple identities, using only monocular videos. The core premise in our method is to learn the non-linear interactions between identity and non-identity specific information with a multiplicative module. We present an extensive study of different variants of our proposed module and their technical derivations. We demonstrate results for both facial expression transfer and talking face video synthesis. By training on multiple videos simultaneously, MI-NeRF not only reduces the total training time compared to standard single-identity NeRFs, but also demonstrates robustness in synthesizing novel expressions for any input identity. Our method can be further personalized for a target identity given only a short video. Project page: https://aggelinacha.github.io/MI-NeRF/ .