<p>Several pruning methods prune a neural network at initialization. These methods carefully determine the importance of each weight, retaining only the important ones and pruning the others. However, subsequent studies have shown that random pruning can perform similarly, provided that the number of remaining weights within each layer is the same as the number of important weights within the corresponding layer. How can this random pruning, which disregards weight importance, especially within layers, still achieve comparable performance? In this work, we shed light on this question. Specifically, we demonstrate that by simply setting the number of weights to retain within layers in this manner, a large number of important weights tend to remain after pruning, and the importance of the remaining weights tends to be high. These statistical benefits are shown by comparing this random pruning method, which retains weights in a layer equal to the number of important weights in that layer, to another random pruning method that uses a uniform pruning ratio across all layers. With randomness applied, where important weights cannot be selectively distinguished from unimportant ones, the superiority of the former over the latter should be clarified. Theoretical proofs are provided, as well as empirical results from various architectures and datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Prior knowledge of layer-specific pruning numbers guarantees effective random pruning at initialization

  • Minju Jung,
  • Sunghyun Baek,
  • Yunho Jeon,
  • Junmo Kim

摘要

Several pruning methods prune a neural network at initialization. These methods carefully determine the importance of each weight, retaining only the important ones and pruning the others. However, subsequent studies have shown that random pruning can perform similarly, provided that the number of remaining weights within each layer is the same as the number of important weights within the corresponding layer. How can this random pruning, which disregards weight importance, especially within layers, still achieve comparable performance? In this work, we shed light on this question. Specifically, we demonstrate that by simply setting the number of weights to retain within layers in this manner, a large number of important weights tend to remain after pruning, and the importance of the remaining weights tends to be high. These statistical benefits are shown by comparing this random pruning method, which retains weights in a layer equal to the number of important weights in that layer, to another random pruning method that uses a uniform pruning ratio across all layers. With randomness applied, where important weights cannot be selectively distinguished from unimportant ones, the superiority of the former over the latter should be clarified. Theoretical proofs are provided, as well as empirical results from various architectures and datasets.