Semantic-Aware HEVC Video Compression: Leveraging Transformer Networks for Enhanced QoS in Real-Time Transmission
摘要
In recent years, efficient transmission of compressed video has become the cornerstone of modern multimedia systems, especially in bandwidth and energy-constrained environments. HEVC has become the standard for video compression. However, more is needed to meet modern multimedia systems requirements. This work proposes a semantic-aware HEVC video compression method that uses Vision Transformers (ViTs) for semantic detection and long Short-Term Memory Models (LSTM) for bandwidth prediction. This approach ensures that important regions like faces and text are preserved with better quality and less important areas are encoded with fewer resources. Experimental results of the proposed framework show significant improvement in PSNR and SSIM compared to state-of-the-art methods, 3 dB in PSNR and 0.04 in SSIM. Also, the proposed framework reduces energy consumption by 20% making it suitable for energy-constrained environments like mobile devices and remote surveillance. A comparison with existing methods highlights the effectiveness of our approach and opens the door for more intelligent and sustainable video transmission frameworks.