Computer vision technology has notably been an important area of research, and in this research, we have integrated voice recognition technology with its application in real-time. Object detection and classification is the most popular computer vision task that involves the identification and localization of objects within images or videos. It has a crucial contribution in various sectors, such the industry, healthcare, transportation, and military. Our objective is to implement YOLOv7, a popular object detection framework, to count specific object categories including Car, Bus, Motorcycle, and Person using human voice as input. Predesigned voice recognition algorithm was employed to convert the human voice to text. By capitalizing on the high accuracy and processing capabilities of YOLOv7, we have developed an efficient counting system that can analyze images and provide object counts for each category of interest. We have evaluated the performance of the system on diverse images taken from various angles, considering challenging scenarios such as occlusions and varying scales. The outcomes demonstrate the effectiveness and reliability of YOLOv7 in precisely counting instances of the specified object classes, even during nighttime and in crowded roads and highways from far distances. This methodology has the potential to be equipped for a drone to provide detecting and counting in an aerial photography view.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Object Detection and Counting Using Voice Recognition for Real-Time Human-Machine Teaming Tasks

  • Mokhles M. Abdulghani,
  • Wilbur L. Walters,
  • Khalid H. Abed

摘要

Computer vision technology has notably been an important area of research, and in this research, we have integrated voice recognition technology with its application in real-time. Object detection and classification is the most popular computer vision task that involves the identification and localization of objects within images or videos. It has a crucial contribution in various sectors, such the industry, healthcare, transportation, and military. Our objective is to implement YOLOv7, a popular object detection framework, to count specific object categories including Car, Bus, Motorcycle, and Person using human voice as input. Predesigned voice recognition algorithm was employed to convert the human voice to text. By capitalizing on the high accuracy and processing capabilities of YOLOv7, we have developed an efficient counting system that can analyze images and provide object counts for each category of interest. We have evaluated the performance of the system on diverse images taken from various angles, considering challenging scenarios such as occlusions and varying scales. The outcomes demonstrate the effectiveness and reliability of YOLOv7 in precisely counting instances of the specified object classes, even during nighttime and in crowded roads and highways from far distances. This methodology has the potential to be equipped for a drone to provide detecting and counting in an aerial photography view.