This paper introduces a novel and robust approach for the performance capture of multiple humans engaged in direct physical interactions with very sparse RGB camera setups. Unlike existing methods that only perform well under specific conditions, such as when humans are relatively distant from each other, when a scene is surrounded by a large array of cameras, or when precise segmentation is available, our method operates without any of these requirements. We introduce a novel layered network architecture to represent the foreground and background together, as well as a tailored compositional volumetric rendering technique and objective functions, along with a new sampling method. These innovations enable the accurate reconstruction of humans engaged in direct physical interactions using only images and roughly estimated SMPL models. Our work demonstrates that our method is able not only to extract high-quality geometry of interacting people but also to provide segmentation and free viewpoint video, outperforming competitors that work in similar setups. Also we show the ability to improve the quality of the roughly estimated SMPL models. We have conducted experiments on a variety of scenes using the HI4D and CMU Panoptic datasets. The code and examples are available at https://github.com/mv2mp/MV2MP .

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MV2MP: Segmentation Free Performance Capture of Humans in Direct Physical Contact from Sparse Multi-Cam Setups

  • Sergei Eliseev,
  • Leonid Shtanko,
  • Rasim Akhunzianov,
  • Yaroslav Romanenko,
  • Anatoly Starostin

摘要

This paper introduces a novel and robust approach for the performance capture of multiple humans engaged in direct physical interactions with very sparse RGB camera setups. Unlike existing methods that only perform well under specific conditions, such as when humans are relatively distant from each other, when a scene is surrounded by a large array of cameras, or when precise segmentation is available, our method operates without any of these requirements. We introduce a novel layered network architecture to represent the foreground and background together, as well as a tailored compositional volumetric rendering technique and objective functions, along with a new sampling method. These innovations enable the accurate reconstruction of humans engaged in direct physical interactions using only images and roughly estimated SMPL models. Our work demonstrates that our method is able not only to extract high-quality geometry of interacting people but also to provide segmentation and free viewpoint video, outperforming competitors that work in similar setups. Also we show the ability to improve the quality of the roughly estimated SMPL models. We have conducted experiments on a variety of scenes using the HI4D and CMU Panoptic datasets. The code and examples are available at https://github.com/mv2mp/MV2MP .