<p>In this paper, we train various Natural Language Processing (NLP) tasks using the Input Normalized Stochastic Gradient Descent (INSGD) optimizer. We fine-tune the Bidirectional Encoder Representations from Transformer (BERT) model on the General Language Understanding Evaluation (GLUE) benchmark with INSGD optimizer. Adaptive Moment Estimation (Adam) and Stochastic Gradient Descent (SGD) optimizers are used as performance comparison baselines. INSGD optimizer leverages SGD optimizer with adaptive <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11760_2025_4209_Article_IEq1.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\(L_{1}\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi>L</mi> <mn>1</mn> </msub> </math></EquationSource> </InlineEquation> and <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11760_2025_4209_Article_IEq2.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\(L_{2}\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi>L</mi> <mn>2</mn> </msub> </math></EquationSource> </InlineEquation> based learning rate normalizations using layer inputs, drawing inspiration from Normalized Least Mean Square (NLMS) algorithm for improved weight updates. We assess the performance of the experiments using the GLUE score on the validation data. Our experiments demonstrate that INSGD achieves higher GLUE scores compared to SGD and Adam across multiple datasets in GLUE benchmark. INSGD surpasses SGD in RTE, MRPC and CoLA datasets, Adam in MNLI-mismatched dataset, and both SGD and Adam in MNLI-matched, QQP, QNLI and SST-2 datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Input normalized stochastic gradient descent for language tasks

  • Esra Türeyen,
  • Salih Furkan Atıcı,
  • Ahmet Enis Çetin,
  • Ömer Morgül

摘要

In this paper, we train various Natural Language Processing (NLP) tasks using the Input Normalized Stochastic Gradient Descent (INSGD) optimizer. We fine-tune the Bidirectional Encoder Representations from Transformer (BERT) model on the General Language Understanding Evaluation (GLUE) benchmark with INSGD optimizer. Adaptive Moment Estimation (Adam) and Stochastic Gradient Descent (SGD) optimizers are used as performance comparison baselines. INSGD optimizer leverages SGD optimizer with adaptive \(L_{1}\) L 1 and \(L_{2}\) L 2 based learning rate normalizations using layer inputs, drawing inspiration from Normalized Least Mean Square (NLMS) algorithm for improved weight updates. We assess the performance of the experiments using the GLUE score on the validation data. Our experiments demonstrate that INSGD achieves higher GLUE scores compared to SGD and Adam across multiple datasets in GLUE benchmark. INSGD surpasses SGD in RTE, MRPC and CoLA datasets, Adam in MNLI-mismatched dataset, and both SGD and Adam in MNLI-matched, QQP, QNLI and SST-2 datasets.