Video classification is used in fields such as surveillance, entertainment, and autonomous driving. Video classification models pose challenges for deployment on resource-constrained devices, and this study focuses on optimizing deep neural networks (DNNs) for computation and memory-efficient video classification. DNN models are optimized using an approach called stripe-wise-pruning (SWP). SWP is a parameter removal method that selectively removes stripes within the filters in the convolutional layers of DNNs based on their importance. In order to maximize computational and memory efficiency, the SparPen optimizer is used with SWP. Key findings of this study indicate that SWP significantly reduces the computational cost and memory utilization of DNN-based video classification models without compromising their accuracy. Experiments are conducted on the UCF50 dataset, comprising 50 diverse human action classes, using the VGG16 architecture. It has provided 90.16% reduction in FLOPs with minimum accuracy drop.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optimization of Human Activity Recognition DNN Model Using Filter Pruning Technique

  • Abhishek Patil,
  • S. M. Meena,
  • Uday Kulkarni,
  • Sunil V. Gurlahosur

摘要

Video classification is used in fields such as surveillance, entertainment, and autonomous driving. Video classification models pose challenges for deployment on resource-constrained devices, and this study focuses on optimizing deep neural networks (DNNs) for computation and memory-efficient video classification. DNN models are optimized using an approach called stripe-wise-pruning (SWP). SWP is a parameter removal method that selectively removes stripes within the filters in the convolutional layers of DNNs based on their importance. In order to maximize computational and memory efficiency, the SparPen optimizer is used with SWP. Key findings of this study indicate that SWP significantly reduces the computational cost and memory utilization of DNN-based video classification models without compromising their accuracy. Experiments are conducted on the UCF50 dataset, comprising 50 diverse human action classes, using the VGG16 architecture. It has provided 90.16% reduction in FLOPs with minimum accuracy drop.