With the development of the Internet, the volume of image and video compression technologies has increased. Traditional compression methods primarily optimize for the Human Visual System (HVS) to maximize visual quality at a certain bitrate. Nevertheless, with the widespread application of machine vision tasks, traditional compression methods have failed to effectively support machine analysis while meeting human visual demands, resulting in increased computational costs during decoding and reduced feature extraction performance. Therefore, this chapter reviews human and machine perception-friendly image and video compression methods, focusing on the application of deep learning in optimizing human perceptual quality and enhancing machine analysis efficiency. Specific topics include the design of loss functions for human perception, semantic coding, Region-of-Interest (ROI) compression, Just-Noticeable-Distortion (JND) modeling, and compression networks specifically designed for machine vision tasks (such as LSMNet). Additionally, this paper introduces the latest developments in Video Coding for Machines (VCM), covering multi-scale feature compression, codebook Hyperprior modelhyperprior models, and competitive learning strategies for content-specific filters. This chapter concludes by evaluating the merits and shortcomings of contemporary methods and identifies future research directions, such as large-scale model integration and cross-model compression, to further refine compression techniques tailored for human and machine perception.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Human and Machine Perception Oriented Image and Video Coding

  • Wei Gao

摘要

With the development of the Internet, the volume of image and video compression technologies has increased. Traditional compression methods primarily optimize for the Human Visual System (HVS) to maximize visual quality at a certain bitrate. Nevertheless, with the widespread application of machine vision tasks, traditional compression methods have failed to effectively support machine analysis while meeting human visual demands, resulting in increased computational costs during decoding and reduced feature extraction performance. Therefore, this chapter reviews human and machine perception-friendly image and video compression methods, focusing on the application of deep learning in optimizing human perceptual quality and enhancing machine analysis efficiency. Specific topics include the design of loss functions for human perception, semantic coding, Region-of-Interest (ROI) compression, Just-Noticeable-Distortion (JND) modeling, and compression networks specifically designed for machine vision tasks (such as LSMNet). Additionally, this paper introduces the latest developments in Video Coding for Machines (VCM), covering multi-scale feature compression, codebook Hyperprior modelhyperprior models, and competitive learning strategies for content-specific filters. This chapter concludes by evaluating the merits and shortcomings of contemporary methods and identifies future research directions, such as large-scale model integration and cross-model compression, to further refine compression techniques tailored for human and machine perception.