<p>Video classification has several essential applications in diversified areas like human action recognition, image captioning, traffic management, etc. Spatiotemporal feature extraction plays a vital role in video classification. However, extracting spatiotemporal features from the video has several difficulties due to its multidimensional nature, encompassing spatial, temporal, and contextual information that demands complex processing and storage. To handle this complex data nature, spatiotemporal deep learning architecture is required. This paper reviews essential architecture for spatio-temporal feature extraction based on 3D convolutional neural network (3DCNN). Furthermore, we also discuss prominent variants of 3DCNN and explore its diverse applications. A comprehensive analysis of evaluation metrics for multiclass-label video classification, multilabel video classification, and spatiotemporal localization are reviewed to provide valuable insights into both model performance and resource utilization. Despite the spatio-temporal architecture's success, this paper emphasizes significant research gaps, including high parameter computation, long-term dependencies, viewpoint variance, spatial fluctuations, model generalization, and more. Additionally, this paper highlights effective methods to tackle these research gaps. This review also summarized highlights for further innovations in spatiotemporal architecture based on model compression techniques, pruning for the deployment on resource constrained devices.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Analyzing spatiotemporal architectures in video classification: A deep dive into 3DCNNs

  • Dhatri Pandya,
  • Keyur Rana

摘要

Video classification has several essential applications in diversified areas like human action recognition, image captioning, traffic management, etc. Spatiotemporal feature extraction plays a vital role in video classification. However, extracting spatiotemporal features from the video has several difficulties due to its multidimensional nature, encompassing spatial, temporal, and contextual information that demands complex processing and storage. To handle this complex data nature, spatiotemporal deep learning architecture is required. This paper reviews essential architecture for spatio-temporal feature extraction based on 3D convolutional neural network (3DCNN). Furthermore, we also discuss prominent variants of 3DCNN and explore its diverse applications. A comprehensive analysis of evaluation metrics for multiclass-label video classification, multilabel video classification, and spatiotemporal localization are reviewed to provide valuable insights into both model performance and resource utilization. Despite the spatio-temporal architecture's success, this paper emphasizes significant research gaps, including high parameter computation, long-term dependencies, viewpoint variance, spatial fluctuations, model generalization, and more. Additionally, this paper highlights effective methods to tackle these research gaps. This review also summarized highlights for further innovations in spatiotemporal architecture based on model compression techniques, pruning for the deployment on resource constrained devices.