A machine’s ability to perceive the world as seen by humans still holds lots of challenges in terms of perspective and contextual understanding. HAR—Human action recognition has a variety of applications in real world like surveillance, health care, manufacturing, transportation, etc., which in turn will give a way to smarter societies. The main aim of this paper is to outline and tabulate the various feature engineering, learning methods and datasets used in research of HAR in vision-based perspective. This survey expands in four folds: (1) chronological view of HAR in research and application; (2) various feature engineering processes and how it affects the learning models; (3) focus on various learning methods used in the process of action recognition and action classification from raw video; and (4) exploring the benchmark datasets available for action or scene recognition from videos. This work also includes the comparative study of two different models in recognizing the activities of eight different classifications of actions from the benchmark COIN dataset. The results obtained from SVM versus 3DCNN-ResNet are tabulated, and 3DCNN-ResNet is found to be having 86% accuracy as average, whereas SVM has the accuracy of 87% as average for eight different actions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Comprehensive Study of Feature Extraction, Learning Methods and Datasets for HAR and Comparative Study of SVM Versus CNN Approach in Classifying Action from COIN Dataset

  • M. Shanmughapriya,
  • S. Gunasundari

摘要

A machine’s ability to perceive the world as seen by humans still holds lots of challenges in terms of perspective and contextual understanding. HAR—Human action recognition has a variety of applications in real world like surveillance, health care, manufacturing, transportation, etc., which in turn will give a way to smarter societies. The main aim of this paper is to outline and tabulate the various feature engineering, learning methods and datasets used in research of HAR in vision-based perspective. This survey expands in four folds: (1) chronological view of HAR in research and application; (2) various feature engineering processes and how it affects the learning models; (3) focus on various learning methods used in the process of action recognition and action classification from raw video; and (4) exploring the benchmark datasets available for action or scene recognition from videos. This work also includes the comparative study of two different models in recognizing the activities of eight different classifications of actions from the benchmark COIN dataset. The results obtained from SVM versus 3DCNN-ResNet are tabulated, and 3DCNN-ResNet is found to be having 86% accuracy as average, whereas SVM has the accuracy of 87% as average for eight different actions.