With the groundbreaking development of Artificial Intelligence (AI) technology, the volume of video and image consumed by machines has surpassed that consumed by humans. Consequently, video coding for machines technology has grown rapidly, with feature coding as a prominent technique demonstrating exceptional compression and task performance. This technology has developed rapidly and has entered the stages of chip integration and industrialization. Feature coding for machines can only reconstruct feature tensors, not images. However, in typical machine vision application scenarios such as smart cities, industrial quality inspection, intelligent transportation, and automated broadcasting, there still persists a demand for video and image review. These needs are essential for retrospective analysis of abnormal events, confirmation of incidents, and evidentiary purposes. This work attempts to reconstruct images from feature tensors extracted for machine vision tasks to meet the needs of human visual observation. We propose a lightweight and plug-in feature-to-image reconstruction method for feature coding for machines, with low complexity neural network blocks. The proposed method achieves an average Peak Signal-to-Noise Ratio (PSNR) of up to 28.92 dB, meeting the requirements of human visual perception.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FIR: A Plug-in Feature-to-Image Reconstruction Method for Feature Coding for Machines

  • Yuan Zhang,
  • Junda Xue,
  • Huifen Wang,
  • Yunlong Li,
  • Lu Yu

摘要

With the groundbreaking development of Artificial Intelligence (AI) technology, the volume of video and image consumed by machines has surpassed that consumed by humans. Consequently, video coding for machines technology has grown rapidly, with feature coding as a prominent technique demonstrating exceptional compression and task performance. This technology has developed rapidly and has entered the stages of chip integration and industrialization. Feature coding for machines can only reconstruct feature tensors, not images. However, in typical machine vision application scenarios such as smart cities, industrial quality inspection, intelligent transportation, and automated broadcasting, there still persists a demand for video and image review. These needs are essential for retrospective analysis of abnormal events, confirmation of incidents, and evidentiary purposes. This work attempts to reconstruct images from feature tensors extracted for machine vision tasks to meet the needs of human visual observation. We propose a lightweight and plug-in feature-to-image reconstruction method for feature coding for machines, with low complexity neural network blocks. The proposed method achieves an average Peak Signal-to-Noise Ratio (PSNR) of up to 28.92 dB, meeting the requirements of human visual perception.