The recent progress in Machine Learning (Géron in Hands-on machine learning with Scikit-Learn, Keras, and TensorFlow. O’Reilly Media, 2022) and particularly Deep Learning (Goodfellow et al. in Deep learning. Cambridge, Massachusetts, MIT, London, 2016) models exposed the limitations of traditional computer architectures. Modern algorithms demonstrate highly increased computational demands and data requirements that most existing architectures cannot handle efficiently. These demands result in training speed, inference latency, and power consumption bottlenecks, which is why advanced methods of computer architecture optimization are required to enable the development of ML/DL-dedicated efficient hardware platforms (I. o. E. a. E. Engineers in 25th IEEE international symposium on high performance computer architecture, p. 734, 2019). The optimization of computer architecture for applications of ML/DL becomes critical, due to the tremendous demand for efficient execution of complex computations by Neural Networks (Goodfellow et al. in Deep learning. Cambridge, Massachusetts, MIT, London, 2016). This paper reviewed the numerous approaches and methods utilized to optimize computer architecture for ML/DL workloads. The following sections contain substantial discussion concerning the hardware-level optimizations, enhancements of traditional software frameworks and their unique versions, and innovative explorations of architectures. In particular, we considered hardware including specialized accelerators, which can improve the performance and efficiency of a computation system using various techniques, specifically describing accelerators like CPUs (multicore) (Hennessy and Patterson in Computer architecture—a quantitative approach. Morgan Kaufman, 2017), GPUs (Wen-Mei in GPU computing gems, Emerald Edition, 2015), and TPUs (Jouppi et al. in In-datacenter performance analysis of a tensor processing unit, 2017), parallelism in multicore architectures, data movement in hardware systems, especially techniques such as caching and sparsity, compression, and quantization, other special techniques and configurations, such as using specialized data formats, and measurement sparsity. Moreover, this paper provided a comprehensive analysis of current trends in Software Frameworks, Data Movement optimization strategies (Olson et al. in Modeling data movement performance on heterogeneous architectures. Waltham, MA, USA, 2021), Sparsity, Quantization, and Compression methods, using ML for Architecture exploration, and, DVFS (Hennessy and Patterson in Computer architecture—a quantitative approach. Morgan Kaufman, 2017), which provides strategies for maximizing hardware utilization and power consumption. The objective of implementing these optimization techniques is to largely minimize the current gap between the computational needs of ML/DL algorithms and the current hardware’s capability. This will lead to significant improvements in training times, enable real-time inference for various applications, and ultimately unlock the full potential of cutting-edge machine learning algorithms.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Computer Architecture Optimization Techniques for AI Workloads

  • Shefqet Meda,
  • Ervin Domazet

摘要

The recent progress in Machine Learning (Géron in Hands-on machine learning with Scikit-Learn, Keras, and TensorFlow. O’Reilly Media, 2022) and particularly Deep Learning (Goodfellow et al. in Deep learning. Cambridge, Massachusetts, MIT, London, 2016) models exposed the limitations of traditional computer architectures. Modern algorithms demonstrate highly increased computational demands and data requirements that most existing architectures cannot handle efficiently. These demands result in training speed, inference latency, and power consumption bottlenecks, which is why advanced methods of computer architecture optimization are required to enable the development of ML/DL-dedicated efficient hardware platforms (I. o. E. a. E. Engineers in 25th IEEE international symposium on high performance computer architecture, p. 734, 2019). The optimization of computer architecture for applications of ML/DL becomes critical, due to the tremendous demand for efficient execution of complex computations by Neural Networks (Goodfellow et al. in Deep learning. Cambridge, Massachusetts, MIT, London, 2016). This paper reviewed the numerous approaches and methods utilized to optimize computer architecture for ML/DL workloads. The following sections contain substantial discussion concerning the hardware-level optimizations, enhancements of traditional software frameworks and their unique versions, and innovative explorations of architectures. In particular, we considered hardware including specialized accelerators, which can improve the performance and efficiency of a computation system using various techniques, specifically describing accelerators like CPUs (multicore) (Hennessy and Patterson in Computer architecture—a quantitative approach. Morgan Kaufman, 2017), GPUs (Wen-Mei in GPU computing gems, Emerald Edition, 2015), and TPUs (Jouppi et al. in In-datacenter performance analysis of a tensor processing unit, 2017), parallelism in multicore architectures, data movement in hardware systems, especially techniques such as caching and sparsity, compression, and quantization, other special techniques and configurations, such as using specialized data formats, and measurement sparsity. Moreover, this paper provided a comprehensive analysis of current trends in Software Frameworks, Data Movement optimization strategies (Olson et al. in Modeling data movement performance on heterogeneous architectures. Waltham, MA, USA, 2021), Sparsity, Quantization, and Compression methods, using ML for Architecture exploration, and, DVFS (Hennessy and Patterson in Computer architecture—a quantitative approach. Morgan Kaufman, 2017), which provides strategies for maximizing hardware utilization and power consumption. The objective of implementing these optimization techniques is to largely minimize the current gap between the computational needs of ML/DL algorithms and the current hardware’s capability. This will lead to significant improvements in training times, enable real-time inference for various applications, and ultimately unlock the full potential of cutting-edge machine learning algorithms.