Parameter-efficient fine-tuning (PEFT) has recently garnered significant attention, due to the enormous size of LLMs. Among various PEFT methods, low-rank adaptation (LoRA) demonstrates comparable performance to full fine-tuning, despite having significantly fewer trainable parameters. In this work, we first generalize LoRA from a low-rank linear adaptation/mapping to low-dimensional, non-linear adaptation/mapping, which we name “low-dimensional adaptation” (LoDA). We also propose LoDA+, which further improves the expressiveness of the non-linear adaptation, while still using nearly the same number of tunable parameters as LoRA. Both LoDA and LoDA+ include LoRA as a special case. To improve computational efficiency at the inference phase, we further propose R-LoDA(+) and S-LoDA(+), by replacing the pre-trained weight matrix with its low-rank or sparse approximation, which is frozen during fine-tuning. Empirical evaluations on natural language generation tasks demonstrate that variants of LoDA outperform LoRA and other baselines.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LoDA: Low-Dimensional Adaptation of Large Language Models

  • Jing Liu,
  • Toshiaki Koike-Akino,
  • Pu Wang,
  • Matthew Brand,
  • Kieran Parsons,
  • Ye Wang

摘要

Parameter-efficient fine-tuning (PEFT) has recently garnered significant attention, due to the enormous size of LLMs. Among various PEFT methods, low-rank adaptation (LoRA) demonstrates comparable performance to full fine-tuning, despite having significantly fewer trainable parameters. In this work, we first generalize LoRA from a low-rank linear adaptation/mapping to low-dimensional, non-linear adaptation/mapping, which we name “low-dimensional adaptation” (LoDA). We also propose LoDA+, which further improves the expressiveness of the non-linear adaptation, while still using nearly the same number of tunable parameters as LoRA. Both LoDA and LoDA+ include LoRA as a special case. To improve computational efficiency at the inference phase, we further propose R-LoDA(+) and S-LoDA(+), by replacing the pre-trained weight matrix with its low-rank or sparse approximation, which is frozen during fine-tuning. Empirical evaluations on natural language generation tasks demonstrate that variants of LoDA outperform LoRA and other baselines.