<p>This paper introduces a novel exponential decay step size for warm restart stochastic gradient descent (SGD), incorporating a logarithmic term that leads to larger step size values compared to the conventional exponential step size proposed in Li et&#xa0;al. (<CitationRef CitationID="CR1">2021</CitationRef>). We provide a rigorous theoretical analysis of the proposed step size decay in the context of non-convex optimization, examining its convergence properties under both standard assumptions and the Polyak-Łojasiewicz (PL) condition. To demonstrate its practical effectiveness, we conduct comprehensive experiments across a range of machine learning tasks, including image classification, language modeling, and object detection. In image classification, the proposed step size is evaluated on benchmark datasets, i.e., FashionMNIST, CIFAR10, and CIFAR100. The results demonstrate that the proposed step size outperformed the traditional exponential decay method, achieving improvements of <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="500_2025_10879_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="43" /> </InlineMediaObject> <EquationSource Format="TEX">\(0.74\%\)</EquationSource> </InlineEquation> on CIFAR-10 and <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="500_2025_10879_Article_IEq2.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="43" /> </InlineMediaObject> <EquationSource Format="TEX">\(2.06\%\)</EquationSource> </InlineEquation> on CIFAR-100 in test accuracy. For language modeling, experiments on the Penn Treebank dataset show that our approach achieves the lowest validation perplexity, indicating superior model optimization. In object detection, using the PASCALVOC dataset and YOLOv5s model, our method attains a mean average precision (mAP@0.5) of 0.842, further confirming its versatility and robustness across tasks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Advancing SGD performance with a new exponential-logarithmic decay step size

  • Mahsa Soheil Shamaee,
  • Sajad Fathi Hafshejani,
  • Zeinab Saeidian

摘要

This paper introduces a novel exponential decay step size for warm restart stochastic gradient descent (SGD), incorporating a logarithmic term that leads to larger step size values compared to the conventional exponential step size proposed in Li et al. (2021). We provide a rigorous theoretical analysis of the proposed step size decay in the context of non-convex optimization, examining its convergence properties under both standard assumptions and the Polyak-Łojasiewicz (PL) condition. To demonstrate its practical effectiveness, we conduct comprehensive experiments across a range of machine learning tasks, including image classification, language modeling, and object detection. In image classification, the proposed step size is evaluated on benchmark datasets, i.e., FashionMNIST, CIFAR10, and CIFAR100. The results demonstrate that the proposed step size outperformed the traditional exponential decay method, achieving improvements of \(0.74\%\) on CIFAR-10 and \(2.06\%\) on CIFAR-100 in test accuracy. For language modeling, experiments on the Penn Treebank dataset show that our approach achieves the lowest validation perplexity, indicating superior model optimization. In object detection, using the PASCALVOC dataset and YOLOv5s model, our method attains a mean average precision (mAP@0.5) of 0.842, further confirming its versatility and robustness across tasks.