Analysis of Computational Costs in Classification Datasets Using Neural Network Quantization Techniques
摘要
Neural network quantization is a technique that has been developed to reduce the computational resources required for deep learning models while maintaining high accuracy levels. There has been recent research on quantification techniques for classification datasets. In this research paper, the performance of neural network quantization on a range of classification datasets, including MNIST, CIFAR-10, and IMDB is investigated. Several quantization techniques have been investigated for reducing the size of models and improving inference speed, such as post-training quantization and weight quantization. The experiments demonstrate that neural network quantization can significantly reduce the computational resources required for classification models while maintaining high levels of accuracy. The results show that different quantization techniques perform differently on different datasets, suggesting that further research is needed to identify the most effective techniques for different types of images. Mobile devices and edge devices, e.g., are resource-constrained environments where deep learning models can be deployed. Overall, this research contributes to the understanding of neural network quantization and its potential for reducing the computational requirements of deep learning models.