<p>This study explores the effects of lossy image and video compression on the accuracy and robustness of deep learning-based person detection models, with particular emphasis on real-world applications such as search and rescue missions during natural disasters. In these critical scenarios, unmanned aerial vehicles (UAVs) and drones are often deployed to capture aerial imagery under operational constraints that include limited onboard processing power, restricted storage capacity, and constrained transmission bandwidth. These limitations frequently necessitate the use of image compression techniques, which introduce visual artifacts that can degrade detection performance. To analyze the resilience of detection algorithms under such conditions, we evaluate and compare the performance of two established deep learning models-RescueNet and YOLOv8-against a novel Proposed Framework that integrates Principal Component Analysis (PCA) for dimensionality reduction and a dual-branch architecture combining YOLOv8 and ResNet-50 for enhanced feature extraction. The evaluation is conducted using datasets compressed with three widely used codecs: HEVC, H.264 (both in intra-frame and inter-frame configurations), and JPEG2000, each tested at four different quantization parameters (QP22, QP28, QP32, QP40), simulating a range of compression levels from mild to aggressive. Experimental results demonstrate that the Proposed Framework consistently outperforms the baseline models across all tested scenarios. On the original, uncompressed dataset, it achieves a top accuracy of 98.99%, outperforming YOLOv8 (96.90%) and RescueNet (82.41%). When exposed to severe compression conditions-such as QP40 using the HEVC Intra codec-the Proposed Framework maintains an accuracy of 64.68%, while YOLOv8 drops slightly to 64.09% and RescueNet drastically declines to 7.95%. These findings highlight the Proposed Framework’s robustness to visual degradation, affirming its potential for deployment in UAV-based surveillance and rescue systems operating under adverse imaging conditions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Development of a computational deep learning model for detecting people in aerial images and videos degraded by compression artifacts

  • Fernando Rodrigues Trindade Ferreira,
  • Loena Marins do Couto

摘要

This study explores the effects of lossy image and video compression on the accuracy and robustness of deep learning-based person detection models, with particular emphasis on real-world applications such as search and rescue missions during natural disasters. In these critical scenarios, unmanned aerial vehicles (UAVs) and drones are often deployed to capture aerial imagery under operational constraints that include limited onboard processing power, restricted storage capacity, and constrained transmission bandwidth. These limitations frequently necessitate the use of image compression techniques, which introduce visual artifacts that can degrade detection performance. To analyze the resilience of detection algorithms under such conditions, we evaluate and compare the performance of two established deep learning models-RescueNet and YOLOv8-against a novel Proposed Framework that integrates Principal Component Analysis (PCA) for dimensionality reduction and a dual-branch architecture combining YOLOv8 and ResNet-50 for enhanced feature extraction. The evaluation is conducted using datasets compressed with three widely used codecs: HEVC, H.264 (both in intra-frame and inter-frame configurations), and JPEG2000, each tested at four different quantization parameters (QP22, QP28, QP32, QP40), simulating a range of compression levels from mild to aggressive. Experimental results demonstrate that the Proposed Framework consistently outperforms the baseline models across all tested scenarios. On the original, uncompressed dataset, it achieves a top accuracy of 98.99%, outperforming YOLOv8 (96.90%) and RescueNet (82.41%). When exposed to severe compression conditions-such as QP40 using the HEVC Intra codec-the Proposed Framework maintains an accuracy of 64.68%, while YOLOv8 drops slightly to 64.09% and RescueNet drastically declines to 7.95%. These findings highlight the Proposed Framework’s robustness to visual degradation, affirming its potential for deployment in UAV-based surveillance and rescue systems operating under adverse imaging conditions.