Through creating AI systems that can interpret and grasp video content in a way that is similar to that of a human, human-centric deep video understanding seeks to close the gap between computer vision and human perception. Deep learning-based methods currently in use overlook the complexities of social relationships, human behavior, and emotional intelligence in favor of object detection, action recognition, and scene interpretation. Their inability to fully comprehend audiovisual content is hampered by this restriction. Human-centric deep video comprehension is an effort to bridge the gap between computer vision and human perception by developing AI systems that can comprehend and interpret video content in a manner akin to that of a human. When it comes to object detection, action recognition, and scene interpretation, deep learning-based techniques now in use ignore the intricacies of social interactions, human behavior, and emotional intelligence. This limitation makes it difficult for them to properly understand audiovisual content. Our method has applications in video analytics, social robotics, and human–computer interaction, among other areas. We can enhance AI systems’ capacity to communicate with people, identify social signs, and offer more precise insights into human behavior by creating systems that can comprehend video information from a human perspective. The proposed method is Vision Transformers (ViTs) that is human-centric deep video detection. The method emphasizes on social interactions, human behavior, and emotional intelligence. The result achieve by methodology is 93% accuracy. (In order to develop AI systems that are more like humans, this abstract suggests a novel approach to deep video understanding that focuses on the subtleties of social interactions and human behavior.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Human-Centric Video Analysis in Industrial Environments

  • Hayder Mohammedqasim,
  • Roa’a Mohammedqasem,
  • Bilal A. Ozturk,
  • Habib Rahman Hamedy,
  • Ali bin Asghar

摘要

Through creating AI systems that can interpret and grasp video content in a way that is similar to that of a human, human-centric deep video understanding seeks to close the gap between computer vision and human perception. Deep learning-based methods currently in use overlook the complexities of social relationships, human behavior, and emotional intelligence in favor of object detection, action recognition, and scene interpretation. Their inability to fully comprehend audiovisual content is hampered by this restriction. Human-centric deep video comprehension is an effort to bridge the gap between computer vision and human perception by developing AI systems that can comprehend and interpret video content in a manner akin to that of a human. When it comes to object detection, action recognition, and scene interpretation, deep learning-based techniques now in use ignore the intricacies of social interactions, human behavior, and emotional intelligence. This limitation makes it difficult for them to properly understand audiovisual content. Our method has applications in video analytics, social robotics, and human–computer interaction, among other areas. We can enhance AI systems’ capacity to communicate with people, identify social signs, and offer more precise insights into human behavior by creating systems that can comprehend video information from a human perspective. The proposed method is Vision Transformers (ViTs) that is human-centric deep video detection. The method emphasizes on social interactions, human behavior, and emotional intelligence. The result achieve by methodology is 93% accuracy. (In order to develop AI systems that are more like humans, this abstract suggests a novel approach to deep video understanding that focuses on the subtleties of social interactions and human behavior.