Feature Compression with Spatial Reduction and Hyperprior Enhancement for Collaborative Intelligences
摘要
Machine vision in mobile applications are widely deployed by distributing the computational workload between the mobile device and the edge server. The efficiency of the collaborative intelligence (CI) can be further improved by compression of the features extracted from the devices. In this work, we propose a learned feature compression scheme for collaborative intelligence, comprising a spatial reduction encoder, a spatial reconstruction decoder, a hyperprior enhanced reconstruct feature module, a deep entropy model, a split visual model. The channel-aware feature selection in encoder is conditionally used based on the bandwidth. To ensure the critical semantic information is preserved, we extend the rate-distortion optimization (RDO) to rate-accuracy optimization (RAO) by introducing the machine-specific distortion. Experimental results show that the proposed scheme achieve a remarkable bitrate savings of up to 51.65% while maintaining better analysis performance to Versatile Video Coding(VVC). Additionally, our scheme is notably lightweight compared to relevant models.