Scaling Laws of Deep Learning Neural Networks: Information Loss
摘要
This chapter investigates the scaling laws of deep ReLU networks through the lens of spectral bias–the phenomenon where neural networks preferentially learn lower-frequency components. Understanding this behavior is critical for bridging the gap between empirical scaling observations and rigorous theoretical foundations of generalization in over-parameterized models. The chapter is structured to transition from fundamental observations to complex multi-layer analysis. It first develops a recursive framework to characterize the evolution of Neural Tangent Kernel (NTK) eigenvalues, proving that they follow a power-law decay as depth increases. Subsequently, the text derives scaling laws for training loss dynamics and extends the analysis to finite-width networks. By establishing quantitative relationships between generalization error, dataset size, and network width, this work provides theoretical guidance for optimizing architectures and data scaling strategies. The structure concludes with detailed mathematical proofs involving spherical harmonics and Rademacher complexity to substantiate the derived scaling relationships.