<p>In order to establish a unified performance evaluation benchmark for large-scale object detection algorithms, solve the problem of fragmented evaluation system, and provide objective basis for algorithm selection and optimization through system experiments. By integrating the MS COCO and KITTI datasets, ablation experiments were designed to compare the performance of the two algorithms, with F1 score, convergence effect, and detection timeliness as the core indicators. The experiments revealed that the convergence speeds and detection accuracy of the two methods on different datasets were basically the same, and the F1 scores were all higher than 0.9, up to 0.931. The ablation experiments further revealed the critical impact of the module design on the performance. Ablation experiments revealed the significant impact of module design. Notably, supplementing the grid attention feature fusion module substantially improved detection accuracy (DA) for specific engineering vehicles, raising DA to 0.615, 0.793, and 0.692 for the three vehicle types tested (representing an improvement of up to 29% over the baseline). Increasing DL model depth by 50% yielded a modest accuracy gain (from 0.921 to 0.929) but incurred a significant detection speed penalty (from 284 FPS to 202 FPS). The study provides reproducible and unified evaluation standards for academia and industry, resolves the uncertainty of algorithm selection and optimization, and lays the foundation for the convergence and extension of multimodal target detection technology.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Unified benchmark construction and algorithm performance evaluation for large-scale object detection

  • Jiayan Wang,
  • Shirong Zou,
  • Yixuan Wang,
  • Wendong Huang,
  • Jiajun Song

摘要

In order to establish a unified performance evaluation benchmark for large-scale object detection algorithms, solve the problem of fragmented evaluation system, and provide objective basis for algorithm selection and optimization through system experiments. By integrating the MS COCO and KITTI datasets, ablation experiments were designed to compare the performance of the two algorithms, with F1 score, convergence effect, and detection timeliness as the core indicators. The experiments revealed that the convergence speeds and detection accuracy of the two methods on different datasets were basically the same, and the F1 scores were all higher than 0.9, up to 0.931. The ablation experiments further revealed the critical impact of the module design on the performance. Ablation experiments revealed the significant impact of module design. Notably, supplementing the grid attention feature fusion module substantially improved detection accuracy (DA) for specific engineering vehicles, raising DA to 0.615, 0.793, and 0.692 for the three vehicle types tested (representing an improvement of up to 29% over the baseline). Increasing DL model depth by 50% yielded a modest accuracy gain (from 0.921 to 0.929) but incurred a significant detection speed penalty (from 284 FPS to 202 FPS). The study provides reproducible and unified evaluation standards for academia and industry, resolves the uncertainty of algorithm selection and optimization, and lays the foundation for the convergence and extension of multimodal target detection technology.