<p>With the ubiquity of cameras as mainstream sensing infrastructure, privacy concerns have escalated. This paper introduces CMPIR, a novel cross-modal transformation framework that employs low-resolution infrared sensors to facilitate user pose image reconstruction without privacy leakage. Technically, CMPIR designs a cross-modal pose image generation model based on conditional generative adversarial networks, integrating the pose semantics of infrared heatmaps with the style attributes of RGB images to generate detailed virtual images suitable for various privacy applications. Meanwhile, to address the challenge of severe semantic information loss in infrared heatmaps, we design a novel pixel-fusion-based data augmentation algorithm, enhancing the ability to extract semantic pose information from infrared heatmaps. We collect <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="42486_2024_184_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="64" /> </InlineMediaObject> <EquationSource Format="TEX">\(\sim 80{,}000\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mo>∼</mo> <mn>80</mn> <mo>,</mo> <mn>000</mn> </mrow> </math></EquationSource> </InlineEquation> frames of daily activity data for experimental verification. Experimental outcomes indicate an average Fréchet Inception Distance (FID) score of 40.55 between the synthesized images and the authentic images, markedly lower than the FID scores for images from disparate activities <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="42486_2024_184_Article_IEq2.gif" Format="GIF" Height="19" Rendition="HTML" Resolution="72" Type="Linedraw" Width="61" /> </InlineMediaObject> <EquationSource Format="TEX">\((\ge 300).\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mo stretchy="false">(</mo> <mo>≥</mo> <mn>300</mn> <mo stretchy="false">)</mo> <mo>.</mo> </mrow> </math></EquationSource> </InlineEquation> These results underscore the superior efficacy of CMPIR in human pose image synthesis, offering a promising solution for privacy-preserving applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CMPIR: cross-modal pose image reconstruction via style-semantic fusion

  • Ruili Shi,
  • Shuai Wang,
  • Zhao-Dong Xu,
  • Shuai Wang,
  • Xiaolei Zhou,
  • Yueqi Su

摘要

With the ubiquity of cameras as mainstream sensing infrastructure, privacy concerns have escalated. This paper introduces CMPIR, a novel cross-modal transformation framework that employs low-resolution infrared sensors to facilitate user pose image reconstruction without privacy leakage. Technically, CMPIR designs a cross-modal pose image generation model based on conditional generative adversarial networks, integrating the pose semantics of infrared heatmaps with the style attributes of RGB images to generate detailed virtual images suitable for various privacy applications. Meanwhile, to address the challenge of severe semantic information loss in infrared heatmaps, we design a novel pixel-fusion-based data augmentation algorithm, enhancing the ability to extract semantic pose information from infrared heatmaps. We collect \(\sim 80{,}000\) 80 , 000 frames of daily activity data for experimental verification. Experimental outcomes indicate an average Fréchet Inception Distance (FID) score of 40.55 between the synthesized images and the authentic images, markedly lower than the FID scores for images from disparate activities \((\ge 300).\) ( 300 ) . These results underscore the superior efficacy of CMPIR in human pose image synthesis, offering a promising solution for privacy-preserving applications.