With the active spread of video recording systems, there is a need to introduce automated systems using neural network-based solutions for object detection and classification. This task is often complicated by the fact that there are requirements for processing frames with optimal accuracy at high FPS on low-power equipment. To solve the problems with these requirements, there are currently a certain number of solutions, of which the most effective approaches using the YOLO family of computer vision architectures are highlighted. The main problem in their application is the large number of existing solutions based on this architecture. Also, a significant complication is the differences in the achieved performance and accuracy indicators for different models. In the course of this scientific work, architectures such as YOLOv4 and YOLOX were selected to compare efficiency, a brief overview of their features was conducted. The purpose of the comparison is to compare the speed and accuracy of frame processing for the most popular architectures of the YOLO family. As part of this study, the best machine learning model was determined for use in object recognition tasks on video.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Comparative Research of the Effectiveness of YOLOv4 and YOLOX Architectures in Object Recognition Tasks on Video

  • Dmitry Gura,
  • Vladislav Dovgal,
  • Roman Dyachenko,
  • Arseniy Kolomytsev,
  • Ivan Budagov

摘要

With the active spread of video recording systems, there is a need to introduce automated systems using neural network-based solutions for object detection and classification. This task is often complicated by the fact that there are requirements for processing frames with optimal accuracy at high FPS on low-power equipment. To solve the problems with these requirements, there are currently a certain number of solutions, of which the most effective approaches using the YOLO family of computer vision architectures are highlighted. The main problem in their application is the large number of existing solutions based on this architecture. Also, a significant complication is the differences in the achieved performance and accuracy indicators for different models. In the course of this scientific work, architectures such as YOLOv4 and YOLOX were selected to compare efficiency, a brief overview of their features was conducted. The purpose of the comparison is to compare the speed and accuracy of frame processing for the most popular architectures of the YOLO family. As part of this study, the best machine learning model was determined for use in object recognition tasks on video.