Input normalized stochastic gradient descent for language tasks
摘要
In this paper, we train various Natural Language Processing (NLP) tasks using the Input Normalized Stochastic Gradient Descent (INSGD) optimizer. We fine-tune the Bidirectional Encoder Representations from Transformer (BERT) model on the General Language Understanding Evaluation (GLUE) benchmark with INSGD optimizer. Adaptive Moment Estimation (Adam) and Stochastic Gradient Descent (SGD) optimizers are used as performance comparison baselines. INSGD optimizer leverages SGD optimizer with adaptive