Deep learning techniques have large scopes in terms of applications and resource requirements, hence calling for a deep analysis of their behavior. Many studies focused on the application part of these models and there is a relatively lesser discussion about the unit-wise behavior. A multi-class dataset was taken and three different models were trained where their performances of training processes are described in terms of runtime characteristics like learning rate, times, etc. An apriori analysis was conducted where the runtime characteristics are investigated for their similarity with pre-defined knowledge of their architectures, as opposed to the posterior analysis which particularly focuses on observations. The work also serves as a framework for defining the logistics of the model training processes and fine-tunes the parameters devised in future for the maximum accuracy of the model. The results were further converted into graphical representations for enhanced delivery of insights. The study is also novel where, as opposed to previous works, in which performances are only tested, the model is evaluated in a design-specific manner dealing with only architectures and the parameters defining them.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Apriori Analysis of Deep Learning Models on a Multi-class Data Set

  • Sai Sathwik Kosuru,
  • Y. C. A. Padmanabha Reddy,
  • Perala Venkata Akanksha,
  • Kotagiri Chaitanya,
  • Snehaja Parsi,
  • Madhu Babu Chunduri

摘要

Deep learning techniques have large scopes in terms of applications and resource requirements, hence calling for a deep analysis of their behavior. Many studies focused on the application part of these models and there is a relatively lesser discussion about the unit-wise behavior. A multi-class dataset was taken and three different models were trained where their performances of training processes are described in terms of runtime characteristics like learning rate, times, etc. An apriori analysis was conducted where the runtime characteristics are investigated for their similarity with pre-defined knowledge of their architectures, as opposed to the posterior analysis which particularly focuses on observations. The work also serves as a framework for defining the logistics of the model training processes and fine-tunes the parameters devised in future for the maximum accuracy of the model. The results were further converted into graphical representations for enhanced delivery of insights. The study is also novel where, as opposed to previous works, in which performances are only tested, the model is evaluated in a design-specific manner dealing with only architectures and the parameters defining them.